The New Senior Developer Job Description

The senior developer role is quietly splitting in two, and most of us haven't caught up yet. Half your job is still writing code. The other half? Architecting how AI writes it with you. I broke down what this shift actually means, read on if you're figuring out where you stand.

The bar moved, and nobody sent a memo Here's what bugs me about this shift: nobody officially updated the senior developer job description. Companies are still posting the same boilerplate ("5+ years, strong CS fundamentals, excellent communication skills") while quietly expecting candidates to also know how to design AI-augmented workflows. It's an unspoken requirement now, and unspoken requirements are the worst kind, because you don't find out you're missing one until you've already lost the job. I don't think this is entirely fair but I also don't think it's avoidable. When Copilot, Cursor, and whatever model wrapper drops next month can genuinely cut a feature's implementation time in half, teams start expecting the senior person in the room to know how to set that up properly, not just use it themselves. There's a real difference between "I use ChatGPT to write boilerplate" and "I designed the review gate that decides which AI output actually reaches production." The first is a skill. The second is architecture. What "AI systems architect" actually means day to day This isn't about prompt engineering, or at least not only that. When I say architecture, I mean things like: Deciding where in the pipeline an AI call happens. Pre-commit hooks catch things early but slow down local dev and annoy people. PR-time review is the sweet spot for most teams. CI-stage review catches what slipped through PR comments but adds latency to merges. Post-merge review is basically a safety net, useful for catching drift but too late to block anything. Example: on the invoice tool I mentioned earlier, we ran lightweight checks pre-commit (formatting, obvious null checks), full model review at PR time, and a nightly batch job that re-scanned merged validation logic against a golden test set to catch anything the PR reviewer missed. Setting up fallback behavior. What happens when the model API times out or returns garbage? On a project last month, our review bot silently failed for two days because nobody configured a timeout, and PRs just sat there with no comment at all. Nobody noticed until a junior dev asked why the "AI review" checkbox never turned green. The fix wasn't complicated. We added a hard timeout, a retry with backoff, and a Slack alert if the bot failed three times in a row. The lesson was more about the blind spot than the code. Building observability into the thing. Acceptance rate on suggestions, tracked weekly, so you know if the model is actually useful or just noisy. False positive rate on flagged issues, because a bot that cries wolf gets ignored within a month. Trace-level logging so you can go back and see exactly what context the model had when it made a bad call, instead of guessing. Managing context windows and retrieval so the model actually knows about your codebase instead of hallucinating a method that doesn't exist. Retrieval over your own repo, or at minimum a solid vector index, isn't optional anymore for anything beyond toy projects. On a shared services codebase, we index changed files plus their direct callers and callees, not just the diff itself, because reviewing a diff in isolation misses breaking changes downstream. Knowing when not to use AI. Genuinely underrated skill. Security-sensitive auth logic. I still want a human writing that from scratch, full stop. Anything touching PII or compliance-flagged data paths. Migration scripts that touch prod data, where a confident-sounding wrong answer can cost you a weekend and an incident report. None of that is "prompting." It's systems design with a probabilistic component bolted on, which is a stranger problem than it sounds. And as I mentioned above, that includes the systems evaluating you for the job, not just the ones you're building on it. A quick example of what changed for me personally When I first started using Copilot back in 2023, it felt like autocomplete with better taste. Type a method signature, get a reasonable body back, accept or reject, move on. Low stakes. Nothing architectural about it. Fast forward to a project I worked on last year, an internal tool that ingests vendor invoices and routes them through validation before hitting our ERP system. Four services touched that pipeline, all Spring Boot. I ended up building a review layer on top of Spring AI 1.0 that ran two different models depending on the diff: a smaller, faster model for straightforward CRUD changes, and a heavier one with more context for anything touching the validation logic itself. Here's the routing logic, simplified: public record DiffStats(int filesChanged, int linesChanged, boolean touchesValidation) {} @Service public class ModelSelector { public String selectModel(DiffStats diffStats) { if (diffStats.filesChanged() <= 2 && !diffStats.touchesValidation()) { return "gpt-4o-mini"; } if (diffStats.touchesValidation() || diffStats.linesChanged() > 300) { return "gpt-4.1"; } return "gpt-4o-mini"; } } Dead simple on paper. But getting the thresholds right took two weeks of watching acceptance rates and adjusting. And the fallback logic, what happens when gpt-4.1 is slow or rate-limited, mattered just as much as the happy path. We ended up queueing and retrying with exponential backoff instead of silently downgrading to the smaller model, because a downgraded review on validation logic is worse than a slow one. Spring Retry made this almost embarrassingly clean: @Service public class ReviewService { private final ChatClient chatClient; public ReviewService(ChatClient chatClient) { this.chatClient = chatClient; } @Retryable( retryFor = RateLimitException.class, maxAttempts = 3, backoff = @Backoff(delay = 2000, multiplier = 2) ) public ChatResponse callWithFallback(String model, Prompt prompt) { return chatClient.prompt(prompt) .options(ChatOptions.builder().model(model).build()) .call() .chatResponse(); } @Recover public ChatResponse recover(RateLimitException e, String model, Prompt prompt) { throw new ReviewPipelineException( "Model " + model + " unavailable after retries", e); } } To make the scenario concrete, here's what a single PR actually walked through in that pipeline: A developer pushes a 40-line change to the invoice validation service, touching the tax calculation logic. ModelSelector flags touchesValidation as true, so it routes to gpt-4.1 instead of the cheaper model. The call goes out with a retrieval-augmented prompt that includes the changed file, its two callers, and the relevant unit tests, not just the raw diff. gpt-4.1 returns a rate limit error on the first attempt because three other PRs merged in the same five-minute window. Spring Retry waits two seconds, retries, gets rate limited again, waits four seconds, and succeeds on the third attempt. The review comment posts to the PR: a flagged edge case around rounding on partial-quantity invoices, correctly caught. The developer accepts the suggestion, fixes the rounding logic, and that acceptance gets logged to our feedback table. A week later, that logged acceptance becomes part of the golden set we use to eval future model versions before we swap them into the pipeline. That's maybe thirty lines of actual code behind all of that. But deciding when to retry, when to escalate, and when to just let a human take over instead, that's the part nobody's teaching in a bootcamp, and it's the part that got the other candidate the job over the person I interviewed. The tools you're expected to know now I won't pretend there's a settled stack here, because there isn't. Things change every few months. But as of right now, if you're a senior candidate, I'd expect you to at least have opinions on: Cursor vs. Windsurf vs. plain Copilot vs. IntelliJ's AI Assistant for in-editor assistance. I personally still lean on...