AI-Native Engineering Culture

I got a rare look inside how Anthropic actually structures their engineering teams, and honestly, it changed how I think about building with AI. Not the polished PR version. The real cultural principles driving one of the most talked-about AI labs in the world.

I've been thinking a lot about what "AI-native" actually means for engineering teams. Not the marketing version. The real, structural version: how do you actually reorganize a team when AI can write your code, run your tests, and coordinate your deployments? Most answers I've seen are either hand-wavy ("just use AI tools more!") or they go too far the other direction and basically suggest laying off half the team. Neither of those is satisfying. Then I found this blog from Gregor Ojstersek and came across some detailed insights from Katelyn Lesse, Head of Platform Engineering at Anthropic, I read them pretty carefully. Anthropic is in a uniquely weird position: they're building AI products while simultaneously using AI to build those products. That recursiveness gives their engineering culture a kind of pressure-tested quality you don't see at companies just now bolting Copilot onto a Jira board. Here's what stood out to me. "AI-Native" Doesn't Mean What You Think It Means A lot of teams use AI tools but aren't AI-native. There's a real difference, and I think it matters more than most people admit. Using AI tools means you have Cursor open in one tab, Claude in another, and you copy-paste between them when it's useful. It's additive. It's still fundamentally the old workflow with a fancy autocomplete slapped on top. Being AI-native means AI is woven into the process so completely that removing it would break how the team functions, not just slow it down. The workflow was designed with AI in it from the start, not retrofitted after the fact. That's the framing Katelyn describes at Anthropic. AI is not a productivity plugin. It's part of the system design. Whether that's collaborative pair-coding with Claude Code or autonomous agents running in the background handling routine tasks, the team's process assumes AI is there. Full stop. Honestly, most teams I've worked on are still in the "tools" category. We added GitHub Copilot to our standard onboarding checklist sometime around mid-2023, and it definitely helped, but our underlying process didn't change. Standups, PR reviews, sprint planning: all the same shape as before. Faster in spots, maybe. Different in structure, not really. The Team Structure Hasn't Changed - The Output Has. This surprised me. I expected Anthropic to have some exotic org structure with weird job titles and no managers or something. Turns out they still use the 2-pizza team model: roughly 5-8 engineers, an engineering manager, a PM, and a designer. Pretty standard cross-functional setup. But here's the interesting part. In a traditional team, you had one tech lead coordinating the technical direction while everyone else focused mostly on feature work and completing tickets. At Anthropic, basically every engineer is expected to step into a tech lead role when needed. The structure looks the same on paper. The expectations don't. The reason is throughput. A team that used to run 1-2 projects at a time can now credibly run 4-5 in parallel. Same headcount, way more surface area. Each engineer operates more independently, makes more architectural calls, and drives their slice of the work end-to-end. I've seen a version of this. During the notifications service extraction we ran last year (pulling it out of a Django monolith that had gotten genuinely painful to work in), we had two engineers carrying what would've previously needed four people coordinating tightly. AI tooling was part of that. But the bigger shift was that both engineers were comfortable owning the whole vertical, not just their assigned tickets. And yeah, it worked. Mostly. More PMs, Not Fewer This is the take I think is most counterintuitive, and I'd push back on anyone who dismisses it without thinking it through. Some companies have aggressively cut PM headcount as engineering velocity increases. The logic: if engineers can ship faster, maybe they can also decide faster. OpenAI reportedly runs something like a 30:1 engineer-to-PM ratio, which is wild when you think about what that actually means day-to-day. Anthropic's view is the opposite. They think PM demand is going up, not down. The reasoning is pretty clean. When engineering stops being the bottleneck, decision-making becomes the bottleneck. You can build anything faster now. The hard question is what to build. That's a PM problem, and no amount of faster code generation solves it. I find this persuasive. Not because I'm a PM advocate (full disclosure: I've had plenty of frustrating PM interactions over the years, and I'm not going to pretend otherwise), but because the logic tracks with what I've seen happen when you give an engineering team unconstrained velocity. They build stuff. Often the wrong stuff. Fast. Anthropic also has a role called Prod Ops (product operations) that I found genuinely interesting. It's basically a function focused on improving the decision-making infrastructure: how teams gather customer insight, how they prioritize, how they run the feedback loops that feed into what gets built next. Not product management exactly, but sitting right next to it. Worth noting: I don't think this model works for every company. A five-person startup probably shouldn't be hiring extra PMs right now. But for a team running 4-5 parallel workstreams with AI-accelerated delivery, having one PM stretched across all of it sounds like a recipe for a very stressed PM and a lot of misaligned work. I've lived that particular version of chaos, and it's not fun for anyone. No Dedicated QA. But Testing Is More Important Than Ever. Anthropic doesn't have QA engineers embedded in their teams. Testing is a shared responsibility, which I know sounds like the thing every company says and few actually mean. What makes it credible here is the investment in evaluation infrastructure. Because engineers can generate code much faster with AI, the risk isn't slowness anymore. The risk is shipping bad things quickly. So the testing discipline has to scale proportionally. They follow the classic test pyramid: lots of fast unit tests, a smaller number of integration tests, and a minimal set of slow end-to-end tests. Nothing exotic there. Standard stuff. But the interesting layer is AI evals. Because Anthropic ships AI products, they have evaluation pipelines that measure model and product quality before anything reaches production. This isn't unit testing in the traditional sense. It's closer to building a second system whose entire job is to judge the first system. Engineers and PMs both contribute to it; it's not siloed. I think about this a lot when I see teams say "we write tests for all our AI features" and then point to a single function that asserts the LLM returns a non-empty string. That's not an eval. An eval is a pipeline that measures whether the output actually did the right thing across a representative set of inputs. It's hard to build and most teams skip it entirely. I get why, but it's a real gap. A simple but real version of what an eval harness looks like in Python: import anthropic from dataclasses import dataclass @dataclass class EvalCase: input_prompt: str expected_behavior: str # description, not exact string match tags: list[str] def run_eval(client: anthropic.Anthropic, case: EvalCase) -> dict: response = client.messages.create( model="claude-opus-4-5", max_tokens=1024, messages=[{"role": "user", "content": case.input_prompt}] ) output = response.content[0].text # Judge the output with a second model call judge_prompt = f""" Expected behavior: {case.expected_behavior} Actual output: {output} Did the output satisfy the expected behavior? Reply with PASS or FAIL and a one-sentence reason. """ judgment = client.messages.create( model="claude-opus-4-5", max_tokens=256, messages=[{"role": "user", "content": judge_prompt}] ) return { "input": case.input_prompt, "o...