Token Frugality in Generative AI Applications
I've spent a good chunk of the past year thinking seriously about token consumption, not in an academic way but in the "this is now competing with our EC2 budget" way. What I've learned is that token frugality isn't really about being cheap. It's about being intentional. Every token you send or receive is a choice you're making, and most of the time that choice gets made unconsciously, buried inside a default setting or a copy-pasted system prompt nobody's touched in six months. Here's what I actually do now. Start With Measurement, Not Guessing You can't optimize what you can't see. Before touching a single prompt, I instrument everything. Spring AI (the Spring Boot integration for LLM calls, now past its 1.0 GA release) exposes usage metadata directly on the ChatResponse object. On the billing service rewrite last fall, we wrapped every LLM call in a small service layer that pulled promptTokens and generationTokens out of the response and published them as Micrometer metrics. Those fed straight into our Datadog dashboards. Took maybe half a day to wire up properly, and within a week we had a clear picture of which endpoints were burning the most money. It was the document summarization flow, not the chat interface. The chat was actually pretty lean. @Service public class TrackedChatService { private final ChatClient chatClient; private final MeterRegistry meterRegistry; public TrackedChatService(ChatClient.Builder builder, MeterRegistry meterRegistry) { this.chatClient = builder.build(); this.meterRegistry = meterRegistry; } public String chat(String userMessage, String model) { ChatResponse response = chatClient.prompt() .user(userMessage) .options(OpenAiChatOptions.builder().model(model).build()) .call() .chatResponse(); Usage usage = response.getMetadata().getUsage(); meterRegistry.counter("llm.tokens.prompt", "model", model).increment(usage.getPromptTokens()); meterRegistry.counter("llm.tokens.completion", "model", model).increment(usage.getGenerationTokens()); return response.getResult().getOutput().getContent(); } } Simple, not fancy but once you have this data flowing, decisions come from reality instead of gut feel, and that changes everything about how you prioritize fixes. Prompt Engineering Is Actually Token Engineering I used to think of prompt engineering as the art of getting better outputs. And it is, but it's also the art of not wasting tokens to get there. The biggest offender I see in production systems is verbose system prompts. I've reviewed codebases where the system prompt was 800 tokens of boilerplate explaining the assistant's "persona", its values, its communication style, and three paragraphs reminding it to be helpful. That stuff gets sent on every single call. At scale, it's brutal. My current practice: trim system prompts to the minimum functional instruction set. No fluff. No "you are a helpful, knowledgeable, and friendly assistant who always tries to..." Just tell it what to do and what constraints apply. I've gotten system prompts from 600 tokens down to under 150 without any measurable drop in output quality. Worth it, every time. In Spring AI, system prompts typically live in .st (StringTemplate) files under resources/prompts/ . That's actually a good thing, because it forces you to treat them as first-class artifacts rather than inline string literals scattered across service classes. When they're in one place, you actually notice when they're bloated. @Service public class SummarizationService { private final ChatClient chatClient; // system-prompt.st loaded from resources/prompts/system-prompt.st @Value("classpath:prompts/system-prompt.st") private Resource systemPromptResource; public SummarizationService(ChatClient.Builder builder) { this.chatClient = builder.build(); } public String summarize(String documentText) { return chatClient.prompt() .system(systemPromptResource) .user(u -> u.text("Summarize the following in 3 sentences: {doc}") .param("doc", documentText)) .call() .content(); } } The other thing I do is use few-shot examples sparingly. Yes, they help but three examples where one would do is three times the prompt tokens for that section. On a high-traffic endpoint, that math adds up fast. And please, stop sending HTML or raw markdown to the model when plain text will do. On one project, we were passing in rendered HTML snippets for the model to analyze. A quick pass through Jsoup to strip tags cut prompt size by about 35% on average. Nobody noticed a quality difference. Not a single complaint. Caching: The Most Underused Tool in the Box Not every LLM call needs to hit the API. This sounds obvious, but I'm still surprised how rarely I see caching implemented in production GenAI apps. There are two kinds worth thinking about. Exact caching is straightforward: if you've seen this exact prompt before, return the cached response. Works well for classification tasks, template-based generation, or any flow where the same question gets asked repeatedly. Spring Boot's @Cacheable with a Redis backend is all you need. We use a SHA-256 hash of the serialized message list as the cache key, and it slots in cleanly without changing the rest of the service layer at all. @Service public class CachedChatService { private final ChatClient chatClient; public CachedChatService(ChatClient.Builder builder) { this.chatClient = builder.build(); } @Cacheable(value = "llm-responses", key = "#root.target.hashMessages(#messages)") public String cachedCall(List<Message> messages) { return chatClient.prompt() .messages(messages) .call() .content(); } public String hashMessages(List<Message> messages) { try { String payload = messages.stream() .map(m -> m.getMessageType() + ":" + m.getContent()) .collect(Collectors.joining("|")); MessageDigest digest = MessageDigest.getInstance("SHA-256"); byte[] hash = digest.digest(payload.getBytes(StandardCharsets.UTF_8)); return HexFormat.of().formatHex(hash); } catch (NoSuchAlgorithmException e) { throw new RuntimeException(e); } } } The Redis config in application.yml is about four lines. Spring Boot autoconfigures the rest. Semantic caching is more interesting and honestly a bit trickier to get right. The idea is that "What's the refund policy?" and "How do I get a refund?" are similar enough that you might serve the same cached response. This isn't something Spring AI handles natively yet, so we've wired it up manually using Spring AI's EmbeddingClient to generate query vectors and a Redis vector index (via Redis Stack) to find similar past queries. I've seen this approach cut API call volume by 20-30% in FAQ-style chatbots. One caveat: semantic caching can go sideways if your similarity threshold is too loose. We had a case where two slightly different questions got the same cached answer, and one of those answers was just wrong for the context. Tighten the threshold and actually test it. Easier said than done, but worth the effort. RAG Context Is Not "More Is Better" If you're building with RAG (Retrieval-Augmented Generation), the naive approach is to retrieve the top-K chunks and dump them all into the context window. K defaults to 5 in most setups, including Spring AI's QuestionAnswerAdvisor . Easy mistake. Those 5 chunks might be 4,000 tokens by themselves. If your query only actually needed 1 chunk to answer correctly, you've just burned 3,000 tokens on noise. And sometimes extra context actively hurts quality, not just cost. The...