Using Anthropic's Agent Skills in Spring AI (And Why I'm Actually Excited About This)
When I heard that Spring AI had added support for Anthropic's agent skills (what Anthropic calls "tool use" in their API, though the Spring AI abstraction wraps it a bit differently), I was genuinely curious whether this solved the problem cleanly or just moved the mess somewhere else. Short answer: it mostly solves it. Here's what I found. What Agent Skills Actually Are Anthropic introduced tool use in Claude a while back, but the "agent skills" framing is a bit more specific. The idea is that you define structured capabilities, essentially typed function signatures, and Claude commits to invoking them with well-formed arguments instead of free-text prose. So rather than Claude saying "here's the weather in Paris, it's 18 degrees and partly cloudy" in a paragraph you have to parse, it calls a get_weather skill with {"city": "Paris"} and you get structured data back. Spring AI 1.0.x wraps this in a way that fits the existing ChatClient and FunctionCallback model. If you've already used Spring AI's function calling support with OpenAI, the mental model transfers directly. That's either reassuring or boring, depending on your mood. Setting Up the Dependency I'm running Spring Boot 3.2.5 with Spring AI 1.0.0. The Anthropic starter is straightforward: <dependency> <groupId>org.springframework.ai</groupId> <artifactId>spring-ai-anthropic-spring-boot-starter</artifactId> </dependency> And in application.properties : spring.ai.anthropic.api-key=${ANTHROPIC_API_KEY} spring.ai.anthropic.chat.options.model=claude-3-5-sonnet-20241022 spring.ai.anthropic.chat.options.max-tokens=1024 One thing that tripped me up initially: if you forget max-tokens , the Anthropic client throws an error at runtime rather than using a default. Every other provider I've used falls back to something sensible. Anthropic doesn't. Not a big deal, just annoying the first time. Defining a Skill This is where it gets interesting. You define a skill as a Spring bean that implements java.util.function.Function . Spring AI picks it up, wraps it in a FunctionCallback , and tells Claude about it when building the request. Here's a simple example I used when rebuilding our internal incident triage tool. The idea was to let Claude query a fake ticket store and return structured incident data: @JsonClassDescription("Request to look up an incident by ID") public record IncidentRequest( @JsonProperty(required = true) @JsonPropertyDescription("The unique incident identifier, e.g. INC-1042") String incidentId ) {} @JsonClassDescription("Details of a found incident") public record IncidentResponse( String incidentId, String severity, String status, String assignedTeam, String summary ) {} And the function itself: @Component("lookupIncident") @Description("Look up an incident by its ID and return current status and severity") public class IncidentLookupFunction implements Function<IncidentRequest, IncidentResponse> { @Override public IncidentResponse apply(IncidentRequest request) { // In the real version this hits our PagerDuty wrapper // For demo purposes: return new IncidentResponse( request.incidentId(), "SEV-2", "investigating", "Platform Engineering", "Elevated error rate on checkout service, root cause unknown" ); } } The @Description annotation is what Claude uses to understand when to call this skill. Getting that description right matters a lot. I tried "look up incident details" at first and Claude kept calling it for things it shouldn't. "Look up an incident by its ID and return current status and severity" worked much better. Garbage in, garbage out, basically. Wiring It Into ChatClient @Service public class IncidentAssistant { private final ChatClient chatClient; public IncidentAssistant(ChatClient.Builder builder) { this.chatClient = builder .defaultFunctions("lookupIncident") .build(); } public String triage(String userMessage) { return chatClient.prompt() .user(userMessage) .call() .content(); } } That's really it. When you call triage("What's the current status of INC-1042?") , Spring AI sends the conversation to Claude along with the skill definition. Claude decides to invoke lookupIncident , Spring AI calls your function with the parsed arguments, sends the result back to Claude, and Claude returns a final response incorporating the structured data. The round-trip is automatic. You don't manually parse the tool call or re-invoke the model. Spring AI handles all the multi-turn stuff internally. Getting Structured Output From the Final Response This is where I spent the most time. There are two layers of structure here and I conflated them at first, which caused a lot of confusion. Layer one: the skill invocation arguments. These are always structured, because that's the whole point of tool use. Claude sends back a typed JSON payload that maps to your IncidentRequest record. Spring AI handles deserialization. Layer two: Claude's final text response after it gets the function result. By default, this is still prose. If you want the final response to be structured JSON too (say, for a downstream service), you need to combine agent skills with Spring AI's BeanOutputConverter or the entity() method on the response. public TriageReport triageStructured(String userMessage) { return chatClient.prompt() .user(userMessage) .call() .entity(TriageReport.class); } Where TriageReport is a simple record: public record TriageReport( String incidentId, String recommendedAction, String urgencyLevel, List<String> affectedServices ) {} Spring AI appends format instructions to the prompt automatically, telling Claude to return its response as JSON matching the schema. Combined with the skill call, you end up with Claude fetching real data and then packaging its analysis into a typed structure. That's the combo that actually solved my original problem. Okay, well, it's not quite magic. If Claude decides to wrap the JSON in a markdown code fence, the parser still throws. I added a small fallback that strips code fences before deserialization. Annoying, but a two-line fix. Multi-Skill Agents You're not limited to one function. For the incident tool I ended up registering three skills: lookupIncident : fetches current incident state listRecentIncidents : returns the last N incidents for a given service, useful for "have we seen this before?" queries getRunbook : looks up a runbook by service name and error type, and this one honestly took the longest to get the description right this.chatClient = builder .defaultFunctions("lookupIncident", "listRecentIncidents", "getRunbook") .build(); Claude figures out on its own which skills to call and in what order. For a query like "Is this a recurring issue on the checkout service, and what does the runbook say?", it called listRecentIncidents first, then getRunbook , then synthesized both results. I didn't have to orchestrate that sequence explicitly. That feels like a genuinely useful property. Not just a parlor trick. What I'd Watch Out For A few things I'd flag from actual usage, not from reading the docs: Token consumption adds up fast. Each skill definition goes into the system context, and with three moderately detailed skills plus conversation history, I was burning through the context window faster than expected. With claude-3-5-sonnet it's not a huge deal budget-wise, but keep an eye on it if you're running high volume. Error handling from your function matters. If your Function throws an unchecked exception, Spring AI propagates it up and the whole request fails. I wrap my function logic in a try-catch and return an error-state response object instead. Claude handles "I couldn't find that incident" in a response object much more gracefully than a...