RAG changed how we build AI apps. But agentic RAG and agent memory? That's where things get really interesting. I broke down the full evolution, where we are now, and where I think this is all heading.
I spent a good chunk of last month trying to get my head around agent memory. The documentation is everywhere, the terminology is inconsistent, and every new paper seems to introduce three more nouns you've never heard before. Episodic memory. Procedural memory. Semantic memory. At some point I just opened a blank doc and started drawing boxes. What actually helped was stepping back and asking a simpler question: how did we get here? Because agent memory didn't appear from nowhere. It evolved, pretty naturally, from something most of us already know: RAG. So that's what this is; a walkthrough of that evolution, from vanilla RAG to agentic RAG to full agent memory, with actual Spring AI 2.0 code to make it concrete. I'm writing this partly to consolidate my own thinking, and partly because I think the standard "short-term vs long-term memory" framing that most blog posts lead with is the wrong place to start. It drops you into the deep end of a taxonomy before you understand why any of it exists. RAG: The Read-Only Starting Point When RAG (Retrieval-Augmented Generation) showed up in the Lewis et al. paper back in 2020, the core insight was pretty simple. LLMs are stateless, their training data has a cutoff, and they can't know things that weren't baked into their weights. So why not give them access to an external knowledge source at query time? The basic workflow has two stages. First, an offline ingestion step where you embed your documents and store them in a vector store. Then at query time, you embed the user's question, do a similarity search, retrieve the top-k results, and stuff them into the prompt. Spring AI 2.0 makes this genuinely clean. The VectorStore abstraction works across PGVector, Redis, Weaviate, Chroma, and a handful of others, so you're not locked into one backend. Here's what a basic ingestion and query setup looks like: // Stage 1: Offline ingestion @Service public class DocumentIngestionService { private final VectorStore vectorStore; public DocumentIngestionService(VectorStore vectorStore) { this.vectorStore = vectorStore; } public void ingest(List<Document> documents) { vectorStore.add(documents); } } // Stage 2: Query time using QuestionAnswerAdvisor @Service public class RagService { private final ChatClient chatClient; public RagService(ChatClient.Builder builder, VectorStore vectorStore) { this.chatClient = builder .defaultAdvisors(new QuestionAnswerAdvisor(vectorStore)) .build(); } public String answer(String question) { return chatClient.prompt() .user(question) .call() .content(); } } The QuestionAnswerAdvisor handles the retrieval and prompt augmentation for you. Clean, predictable, and honestly good enough for a lot of use cases. I wired up something like this when our customer success team needed to query an internal policy knowledge base. For straightforward factual questions it worked fine. Hallucinations dropped noticeably compared to prompting the model cold. But it has a real ceiling. The retrieval always fires, whether the information is actually needed or not. There's one monolithic knowledge source for everything. And if the retrieved chunks happen to be irrelevant to what the user actually asked, the model just works with what it got. It doesn't go back and try again. One shot, one answer. Not ideal for anything complex. The other thing I kept bumping into: there was no way for the system to get smarter over time. Every query started from scratch. The pipeline had no memory of what worked, what didn't, or what users actually cared about. Stateless all the way down. Agentic RAG: When Retrieval Becomes a Tool The shift to agentic RAG is conceptually clean once you see it. Instead of retrieval being a hardcoded step in a pipeline, it becomes a tool the agent can choose to call. Or not call. That one change opens up a lot. Now the agent can decide: do I actually need external information here? If yes, which source should I query? Is the result I got relevant, or should I try a different query? And it can loop. Retry, rephrase, pull from multiple sources in sequence. Spring AI 2.0 has solid support for this through its @Tool annotation and the ChatClient fluent API. You define the search capability as a tool, hand it to the model, and let the agent loop handle the rest: // Define the search tool @Component public class DocumentSearchTool { private final VectorStore vectorStore; public DocumentSearchTool(VectorStore vectorStore) { this.vectorStore = vectorStore; } @Tool(description = "Search the knowledge base for information relevant to the query") public String search(String query) { SearchRequest request = SearchRequest.query(query).withTopK(5); List<Document> results = vectorStore.similaritySearch(request); return results.stream() .map(Document::getContent) .collect(Collectors.joining("\n---\n")); } } // Agentic RAG loop @Service public class AgenticRagService { private final ChatClient chatClient; public AgenticRagService(ChatClient.Builder builder, DocumentSearchTool searchTool) { this.chatClient = builder .defaultTools(searchTool) .build(); } public String answer(String question) { return chatClient.prompt() .user(question) .call() .content(); } } Spring AI handles the tool call loop internally. The model decides when to invoke search , inspects the result, and keeps generating until it has enough to respond. You don't have to wire the loop yourself. When our team switched to this approach on an internal research assistant (we were pulling from a few different internal wikis and a product changelog), the difference was immediately obvious. The agent would sometimes search twice with different phrasings, or skip the search entirely for simple clarifying questions. The failure modes changed too. Instead of confidently wrong answers built on irrelevant chunks, we'd see the agent say it couldn't find the information. Easier to debug. Easier to trust. But here's the thing, and this is easy to miss: both vanilla RAG and agentic RAG share the same fundamental constraint. The external knowledge source is read-only during inference. Data gets written offline, during ingestion, and the agent only ever reads from it at runtime. Neither system can learn from what happens during a conversation. That's the gap agent memory fills. Agent Memory: Finally, Read-Write The step from agentic RAG to agent memory is smaller than it sounds. You keep everything from agentic RAG, the tool-based retrieval, the agent loop, the ability to decide when and where to search. You just add write access to the memory store. The agent can now store things, not just retrieve them. Spring AI 2.0 introduced a proper ChatMemory abstraction for in-conversation memory, and you can extend the same tool pattern to give the agent write access to a persistent store between sessions. The interesting design question here is actually what that persistent store should be. The naive version puts everything in a vector store. Embed the memory, write it in, retrieve by similarity later. That works, and it's fine for prototyping. But in production you often want something more structured. Preferences and user profile data fit naturally in MongoDB or PostgreSQL. You get proper querying, indexing, update semantics, and you can audit what the agent wrote. A vector store is great for semantic recall, but it's a poor fit for structured facts you might need to update, delete, or query by field. Here's what a more realistic MemoryWriteTool looks like when you're writing to MongoDB alongside a vector store for semantic recall: // Memory record stored in MongoDB public class UserMemory { @Id private Stri...