I got tired of copy-pasting agent rules into every project. agent-skills is an open catalog of 65 engineering workflows for Cursor, Codex, and Copilot — install as files, improve in Git.
I merged a feature into a Spring Boot service last month — green CI, review approved, agent helped with most of the boilerplate — and the next day a teammate opened Cursor on a different repo and got completely different guidance for the same kind of REST endpoint. Different DTO rules. Different test advice. One project even had a .cursor/rules file that contradicted the AGENTS.md in the same tree. Nobody broke anything on purpose. We just never had a single place for "how our agents should work on Java backend stuff." Every repo had its own fragment: a chat export, a Notion page, a rule file someone wrote six months ago and forgot. That friction is what pushed me to build agent-skills — an open catalog of engineering skills you install as ordinary files. Cursor, Codex, and Copilot read them locally. You review them in Git. You fix them once and reuse them everywhere. Today it ships 65 active skills and 11 install packs across Java, .NET, PHP, React, Angular, architecture review, quality gates, and AI engineering. What We Were Actually Trying to Fix Ask a team what they want from AI coding agents and you'll get the same list. Consistent reviews. Sensible test strategy. Security that doesn't get skipped. Migrations that don't skip the scary steps. The implicit plan was: each developer curates their own rules, each repo accumulates its own AGENTS.md , and peer review catches the gaps. That's fine when the agent is a fancy autocomplete. It's not fine when the agent is writing half the vertical slice — controllers, tests, CI fixes — and the instructions are stale, contradictory, or missing entirely. The checklist said "reviewed and merged." Nobody was reviewing whether the agent playbook was any good. When Instructions Live in Chat Threads I've been using Cursor daily for a while, Copilot before that, and more recently Codex on some repos. The experience is the same across tools: the model is rarely the bottleneck. The bottleneck is whether it knows your conventions — package layout, quality gates, how you write acceptance tests, what "done" means on your team. Copy-pasting prompts from an old thread works until it doesn't. Someone updates Spring Boot and the pasted advice still mentions Boot 2 patterns. Someone adds Checkstyle and the agent keeps suggesting shortcuts that fail ./gradlew check . Someone disables a rule because it fought with another rule. Sound familiar? It's the same failure mode as "definition of done" drifting — except here the proxy is local markdown scattered across repos instead of a sprint checklist. I wrote about that gap in AI Broke Your Definition of Done . Agent instructions have the same problem: the activity (having rules) no longer guarantees the outcome (consistent, trustworthy behavior). The Stuff Copy-Paste Was Never Solving There's a whole category of agent guidance that doesn't survive copy-paste: Workflow order — inspect the build file before suggesting dependencies; run the smallest test command first. Stack-specific gates — Checkstyle, JaCoCo, ESLint, PHPStan, dotnet format , not generic "write tests." Planning vs coding — architecture review skills shouldn't fire during a typo fix; implementation skills shouldn't skip validation. Reviewable depth — thin routing rules that point to full references, not ten-page prompts in every repo. That last one matters. Agents need concise triggers and deep material when the task actually applies — same idea as writing acceptance scenarios before implementation, but for how the agent works . What a Shared Skill Library Actually Looks Like agent-skills treats each workflow as a skill : a folder with a workflow file, metadata, references, and an optional eval prompt. A skill is not a mega-prompt. Roughly: SKILL.md — what to inspect, in what order, what to deliver skill.yaml — domain, planning vs coding mode, packs, which agents support it references/ — checklists, commands, code patterns linked from the workflow eval/prompt.md — a realistic scenario to sanity-check behavior Examples I've been using in anger: java-spring-boot-service — REST vertical slices, DTOs, validation, transactions, tests java-quality-gates — fix CI when Checkstyle or JaCoCo fails, without disabling the gate system-architecture-review — boundaries, tradeoffs, migration risk technical-documentation-authoring — RFC, ADR, BRD structure that humans can actually review Skills declare modes : planning (design, review) or coding (implement, verify). Install planning packs for design sessions; coding packs for feature branches. Three zones in the repo The catalog separates concerns so contributors don't fight generated output: skills/ ← source of truth (you author here) registry/ ← packs, collections, backlog dist/ ← generated Cursor / Copilot / Codex output (never edit by hand) skillctl validates structure, builds vendor bundles, and installs packs. Because dist/ is committed , you can clone and install without Python — just the install script. Try It Without Becoming a Maintainer You don't need to understand the whole toolchain to get value. Clone once, copy a pack into your project: git clone git@github.com:josalero/agent-skills.git cd agent-skills ./scripts/install-from-clone.sh \ --dest /path/to/your-project \ --pack java-backend-pack Other packs that have been useful on real work: architecture-review-pack , quality-gates-pack , frontend-react-pack . For Codex or Copilot, add --target codex or --target copilot . Then ask for something concrete — implement an endpoint, fix a quality gate failure, run an architecture review — and notice whether the agent follows a workflow instead of generic stack overflow energy. Nothing phones home at runtime. These are files in your repo, same as your own rules. How to Contribute Without Owning the Whole Catalog I'm pretty explicit about this in the repo: bulk proposals are fine; active skills require intentional authoring. Easy wins if you want to help: Clarify a step in an existing SKILL.md Add a reference with a real command that worked on your stack Fix version drift (Boot 3.5 vs 4.0, ESLint flat config, whatever bit you) Adding a skill is more involved but mechanical: git clone git@github.com:josalero/agent-skills.git cd agent-skills pip install -e . cp -r templates/canonical-skill skills/my-new-skill # edit SKILL.md, skill.yaml, references/, eval/prompt.md # register in registry/packs/ and registry/collections/ make check make check runs validation, tests, build, and verifies dist/ is in sync. CI runs the same on Windows, macOS, and Linux. Active skills can't contain TODO in SKILL.md — same bar as "we trust this in production agent sessions." Full contributor docs: Authoring skills · Contributing The Deeper Point The real issue is familiar if you've been thinking about AI and delivery quality. We treated "having agent rules" as a proxy for "agents behave well on our stack." Copy-paste rules, drop an AGENTS.md , mark the story done. When the agent writes the code and reads the instructions from a stale fragment in each repo, the proxy breaks quietly. Everything still looks configured. Velocity looks great. The pain shows up later — in review noise, in CI fixes the agent should have anticipated, in three repos doing the same thing three different ways. agent-skills is my attempt to make the playbook shared, versioned, and reviewable — the same way we'd treat a library or a CI template, not a private prompt collection. I'm still expanding the catalog but I'm pretty sure "every repo rolls its own rules" stopped being enough a while ago, and most of us just haven't centralized the checklist yet. Repo: github.com/josalero/agent-skills If something's wrong or missing — that's a PR waiting to happen.