Deep Dive on Spring AI 2.0

Spring AI 2.0 Is a Big Deal. Here's What I Actually Think After Using It.

I've been building Spring-based backends for about 12 years now, and I'll be honest: when the Spring team first announced Spring AI, my reaction was something like "great, another opinionated abstraction over OpenAI's API." I was wrong. Not completely wrong, but wrong enough that I owe the project a proper look. Spring AI 2.0 landed a few months back and I've spent real time with it across two projects, including a document Q&A service we built for an internal knowledge base. This is what I found. Why 2.0 Feels Different from What Came Before The 0.x and 1.x releases were genuinely useful, but they felt experimental. The API surface changed constantly. I remember migrating from 0.8 to 1.0 on our recommendation engine prototype and spending half a day fixing import paths and renamed interfaces. Not fun. 2.0 is stable in a way the earlier versions weren't. The team committed to a real API contract, and you can feel it. Things that felt bolted on before, like the chat memory abstraction and tool calling, now feel like first-class citizens. It's a more coherent library. One practical note: 2.0 requires Spring Boot 3.3+, so make sure your project is on Java 21 before you start. We had to bump one service from Java 17, which took about an afternoon but was completely worth it. The ChatClient API The old ChatClient was fine. The new one is genuinely nicer to work with. The fluent builder API makes common things feel obvious, and the defaults are sensible. Here's what a basic setup looks like with the OpenAI starter: @Configuration public class AiConfig { @Bean public ChatClient chatClient(ChatClient.Builder builder) { return builder .defaultSystem("You are a helpful assistant for a software documentation platform.") .build(); } } And then using it in a service: @Service public class DocumentAssistant { private final ChatClient chatClient; public DocumentAssistant(ChatClient chatClient) { this.chatClient = chatClient; } public String ask(String question) { return chatClient.prompt() .user(question) .call() .content(); } } That's it. No boilerplate mess, no manually building Message lists, no wrestling with request/response mapping. The defaultSystem on the builder means you're not repeating the system prompt on every call. Small thing, but it adds up fast when you're making that call from fifteen different places in a service. I do still keep the raw OpenAiChatModel around for cases where I need fine-grained control over temperature or top-p on a per-request basis. But for 90% of my use cases, ChatClient is enough. Structured Output with toEntity() This is the feature I wish I'd found on day one instead of week three. The toEntity() method on the call response lets you map the model's output directly to a Java record or class, with Spring AI handling the JSON schema generation and deserialization automatically. Say you want the model to extract structured information from a block of unstructured text. Define a record: public record DocumentSummary( String title, String author, List<String> keyTopics, String sentiment ) {} Then call it like this: public DocumentSummary extractSummary(String rawText) { return chatClient.prompt() .user(u -> u.text(""" Extract the document metadata from the following text. Return title, author, key topics, and overall sentiment. Text: {text} """) .param("text", rawText)) .call() .entity(DocumentSummary.class); } That's it. No manual JSON parsing, no ObjectMapper calls, no fragile string splitting. Spring AI generates the JSON schema from your record definition, sends it to the model as a response format constraint, and deserializes the response back into your type. When I first got this working on the knowledge base project, I actually went back and deleted about 60 lines of parsing code I'd written the week before. It works with generics too. If you want a list of entities back: public List<DocumentSummary> extractMultiple(String rawText) { return chatClient.prompt() .user(u -> u.text("Extract all documents mentioned in: {text}") .param("text", rawText)) .call() .entity(new ParameterizedTypeReference<List<DocumentSummary>>() {}); } The ParameterizedTypeReference approach is the same pattern Spring's RestTemplate and WebClient use, so it should feel familiar. Worth knowing: this feature depends on the model supporting structured/JSON output mode. OpenAI's GPT-4o and GPT-4o-mini both handle it well. Claude does too. Older models can be flaky about it. Prompt Templates One more thing in the ChatClient API that I use constantly: prompt templates with named parameters. Instead of string concatenation (which gets ugly fast), you can do this: public String generateReport(String topic, String audience) { return chatClient.prompt() .user(u -> u.text(""" Write a technical summary about {topic} for a {audience} audience. Be concise, use examples, and keep it under 300 words. """) .param("topic", topic) .param("audience", audience)) .call() .content(); } Clean, readable, and easy to move the template out to an external file later if needed. Tool Calling, Which Is Honestly the Feature I Was Most Excited About Spring AI 2.0 ships with a proper tool calling abstraction that works across providers. You define a Java method, annotate it, and the model can invoke it when it decides to. The framework handles the serialization and the back-and-forth with the model automatically. Here's a simplified version of something I actually used on the knowledge base project: @Component public class DocumentTools { private final DocumentRepository documentRepository; public DocumentTools(DocumentRepository documentRepository) { this.documentRepository = documentRepository; } @Tool(description = "Search internal documentation by keyword. Returns matching document titles and summaries.") public List<DocumentSummary> searchDocuments(String query) { return documentRepository.searchByKeyword(query); } @Tool(description = "Fetch the full content of a document by its ID.") public String getDocumentContent(@ToolParam(description = "The unique document identifier") String documentId) { return documentRepository.findContentById(documentId) .orElseThrow(() -> new IllegalArgumentException("Document not found: " + documentId)); } } Then when building the client: @Bean public ChatClient chatClient(ChatClient.Builder builder, DocumentTools tools) { return builder .defaultSystem("You help employees find answers in internal documentation.") .defaultTools(tools) .build(); } The model decides when to call searchDocuments and what to pass it. It works. Like, it actually works in a way that feels a bit magical the first time you see it in a log. The model calls the tool, gets the results back, and synthesizes a response, all without me writing a single line of tool-dispatch logic. One thing I'd warn people about: be specific in your @Tool description. I had a vague description early on and the model called the search tool for questions that didn't need a search at all. More precise descriptions fixed it. Took me embarrassingly long to figure out that the description is basically a prompt fragment. Lesson learned. You can also register tools dynamically per call rather than baking them into the client at build time, which is useful when the available tools depend on the current user's permissions: public String askWithDynamicTools(String question, List<Object> userTools) { return chatClient.prompt() .user(question) .tools(userTools.toArray()) .call() .content(); } RAG...