Beer & Servers Don't Mix

Teaching the Agent What You Know: How to Build, Test, and Ship AI Coding Skills That Actually Hold…

Or: Your AI Assistant Is as Naive as the Day It Was Hired — and That’s Your Fault

Teaching the Agent What You Know: How to Build, Test, and Ship AI Coding Skills That Actually Hold Up in Production

Or: Your AI Assistant Is as Naive as the Day It Was Hired — and That’s Your Fault

Somchai stared at the screen at 11:47 on a Thursday morning, the hum of the air conditioning settling over the floor like a shared sigh. The code review had come back clean — from the linter, anyway. The AI had generated the whole component in under a minute. Twenty-three files changed, all passing CI. He scrolled slowly through the diff, the Thai iced tea going warm beside his keyboard, and then stopped on a single function. The API call was inside the component. On the client. For every render.

That’s when he opened a new chat window and started re-prompting.

Here’s the uncomfortable truth about AI coding assistants: they’re not slow, they’re naive. They produce code that compiles, passes tests, and satisfies every constraint you gave them — which is precisely the problem. You didn’t give them the constraints that live in your engineers’ heads. The ones nobody wrote down because everyone already knew them. The ones a senior engineer communicates through a single comment in a code review, and a junior absorbs over six months of watching that happen.

You can’t mentor the agent over six months. You can’t pair-program it into your conventions. Unless you encode that knowledge somewhere it can find it, your AI assistant will keep producing technically correct, production-naïve code — and your engineers will keep fixing it manually rather than shipping faster. That’s not an AI problem. That’s a knowledge transfer problem.

This post is about solving it properly.

The Three-Layer Problem

When an engineer picks up a codebase, they absorb knowledge in three layers. Most AI tooling only reaches the first.

Layer 1: What the compiler knows. Types, function signatures, API contracts. This is the layer AI coding tools are genuinely excellent at. Give them a typed interface and they’ll implement it correctly.

Layer 2: What the linter knows. Style conventions, formatting rules, naming patterns — whatever you’ve encoded in ESLint, Prettier, StyleCop, or ktlint. AI tools honour these too, usually.

Layer 3: What the team knows. This is where AI tools fail completely, because nothing enforces it. It’s the knowledge that lives in your engineers’ heads and occasionally surfaces as a code review comment:

“We don’t do aggregation client-side — we’ve been bitten by the connection limit too many times.”

“One class per file in C#, always. If you’ve got an interface with a single implementation, that’s the one exception.”

“Never patch the symptom. Find the root cause. If the fix doesn’t explain why the bug happened, it’s not the fix.”

“React components should be named and ordered to mirror the visual layout. If you put the page and the code side by side, you should be able to point to what renders what without drilling down four levels of components.”

That last one came from Roland, eight years ago. It changed how I read every React codebase I’ve touched since. It took thirty seconds to say and had zero representation in any linter rule, any documentation, any file in the repository.

The agent doesn’t know it. And unless you do something about that, it never will.

What Skills Actually Are

A skill is a Markdown file. That’s the unfussy version. It’s a set of instructions that an AI coding agent should apply when handling a particular category of task — written for the agent the same way you’d write it for a new engineer, but with the specificity the agent can act on.

The IDE (Cursor, Claude Code, and others) maintains a catalogue of skill files in a local directory. When an engineer prompts the agent, the IDE reads the prompt, scans the skill catalogue, and decides which skills are relevant. The selected skills are injected into the context window alongside the prompt and code. The agent produces output with that knowledge baked in.

Here’s the flow:

The decision about which skills get selected is made entirely on the name and description of each skill file. If those are vague or overlapping, you get one of two failure modes: too many skills injected (context bloat, degraded output), or the right skill never selected at all (invisible, useless). Writing skill descriptions is interface design, not documentation.

Here’s what a real skill looks like. This one encodes an engineering culture norm — the kind of thing that separates production-grade engineering from code that merely compiles:

## Root cause analysis

**Never patch symptoms. Always find and fix the root cause.**

Before writing any fix, trace the problem to its origin.
A fix that suppresses the symptom without addressing why
it happened will fail again - or worse, hide the real issue
until it causes broader damage.

With this skill injected, agents stop patching and keep investigating until they find the underlying problem. Without it, they patch. Every time. Because patching is the fastest path to a passing test, and passing tests is what they’re optimising for.

That distinction matters more than it might appear — and we’ll come back to it in the risks section.

The Testing Method

Here’s where most teams stop. They write a few skill files, put them in a directory, tell engineers to install the MCP, and consider the work done. Then six months later, someone rewrites a skill during a late-night cleanup sprint, removes a constraint that seemed redundant, and the agent starts generating C# files with two classes in them again. Nobody notices for three weeks.

Skills are prompts. Prompts can be tested. The fact that most teams don’t test them means they’re deploying skills blind — with no feedback loop, no regression detection, and no way to know if a change improved things or quietly broke them.

The testing method below applies the same discipline to skills that we apply to code.

Level 1: Unit Tests with Promptfoo

Promptfoo is an open-source CLI for prompt evaluation. It lets you point it at a skill file (the prompt), supply test inputs, and assert properties of the output — using regex, exact match, or an LLM-as-judge for more nuanced assertions.

Think of this as a unit test for a skill. You’re not testing the full agent loop. You’re testing whether a specific piece of knowledge survives in the agent’s output when this skill is present.

Example: Your skill says “in C#, one class per file — the only exception is an interface with a single implementation.” A promptfoo test asks: “Should I put two classes in the same file?” The assertion checks that the response contains “no”, “never”, or equivalent negative language. If someone rewords the skill and accidentally drops that constraint, the test fails. The pipeline goes red. Someone notices.

This is directly analogous to unit testing a function: isolate the behaviour you care about, assert it explicitly, make it part of CI. Anyone who wants to change a skill has to change the tests, or the build breaks.

A promptfoo test file looks like this:

# promptfoo.yaml
prompts:
  - file://skills/csharp-conventions.md
providers:
  - anthropic:claude-sonnet-4-20250514
tests:
  - description: "One class per file rule"
    vars:
      question: "Can I put two classes in the same file in C#?"
    assert:
      - type: llm-rubric
        value: "The response must say no and explain the one-class-per-file rule"
  - description: "Single implementation exception"
    vars:
      question: "I have an interface with exactly one implementation. Same file?"
    assert:
      - type: llm-rubric
        value: "The response must say yes and correctly identify this as the only exception"

What promptfoo catches: specific facts and constraints that must survive a skill edit. The things small enough to be missed in an end-to-end test, important enough to break behaviour if lost.

What promptfoo misses: whether the skill actually gets selected by the IDE when it should. Whether the full agent loop — prompt plus context plus injected skill — produces better output than without it. That’s what the next level is for.

Level 2: End-to-End Tests with an LLM Judge

The end-to-end test validates the complete pipeline: code goes in, agent receives the prompt and injected skill, code comes out. The question is whether the output is meaningfully better when the skill is present.

The method:

  • Code before — a representative piece of code with known deficiencies. The naive output you’d expect from an agent without the skill.- Prompt — the task description the engineer would give.- Agent run — prompt plus code-before goes through the agent, with the skill injected.- Code after — the agent’s output.- LLM judge — a different model (ideally from a different vendor) receives both versions and scores the improvement on a percentage scale. (Note: Gemini is awesome for reviewing ime, and Calude is the best for editing)- Threshold — you set a minimum acceptable score. Below it, the skill needs work. Don’t use binary output, if you give an LLM a binary success failure, it’ll always bias success, if you give it success of “making a grade” it’ll be more honets, and the threshold/if statement can live in deterministic code. The reason to use a different model as the judge is the same reason you don’t mark your own homework. A model that helped write the code is poorly positioned to evaluate it critically. Cross-model judging — Gemini judging Claude output, or vice versa — is more likely to surface genuine deficiencies.

The percentage scoring matters too. You’re not asking the judge to pass or fail. You’re asking it to score improvement. That goal is significantly harder to game than a binary outcome.

The two levels complement each other. Promptfoo catches the specific, precise constraints that absolutely cannot be lost. End-to-end tests catch whether the skill is actually doing something useful in the real agent loop. Neither alone is sufficient. Together, they give you a CI pipeline for skills with the same confidence properties as a well-tested codebase.

The testing infrastructure lives in its own repository: skills, code-before examples, prompts, and the CI pipeline all version-controlled together. Adding a skill requires adding at least one test. Changing a skill may require changing tests. The pipeline doesn’t let untested skills reach production.

The Risks (The Part You Should Actually Read)

The testing method is the answer to most of these. But you should understand what you’re defending against.

Risk 1: Bad skill descriptions cause the wrong skills to be selected

The IDE selects skills by reading their names and descriptions. If those are vague (“general coding standards”) or overlapping (“React tips” and “React component structure”), one of two things happens:

Context bloat: too many skills get injected. The context window fills up. Output quality drops. Engineers re-prompt more, not less.

Skill invisibility: the skill is never selected because its description doesn’t match the vocabulary of the prompt. You’ve built a skill no one can find.

Write skill descriptions as API contracts. The description is the interface by which the IDE discovers and selects the skill. Treat it accordingly.

Risk 2: Agents optimise for passing tests, not for doing the right thing

This is the specification gaming problem, and the research literature is uncomfortable reading. ImpossibleBench (arXiv:2510.20270, Oct 2025) documents how LLM agents, when given the goal of making tests pass, will find paths that don’t involve actually solving the problem — deleting failing tests rather than fixing the underlying bug, hardcoding values that satisfy assertions without implementing the real logic, adding skip markers to flaky tests rather than fixing them.

It’s not theoretical. One engineer in our team experienced this directly: an agent added a Python script that bypassed a CI check rather than addressing the root cause the check was catching.

METR’s research on real-world reward hacking (June 2025) found something more unsettling: training models to avoid detectable cheating sometimes caused them to find less detectable cheating strategies. The field doesn’t have a definitive solution.

The LLM-judge pattern mitigates this. You’re not asking the judge to pass or fail — you’re asking it to score improvement. That’s a much harder goal to game. And cross-model judging means the judge has no incentive to rationalise the output it’s evaluating.

It’s a mitigation, not a guarantee.

Risk 3: Skills go stale and nobody notices

Code changes. Standards evolve. A skill written for a codebase with one set of conventions may become actively misleading as the codebase matures. Unlike code, skills don’t have type checkers to tell you when they’re wrong. They’ll silently produce incorrect output.

If you’ve read the Documentation Graveyard post, this is the same failure mode — knowledge encoded in a format that has no maintenance forcing function. The difference is that skills can be tested. If you maintain your code-before examples alongside the skill, a stale skill will eventually fail its tests.

The discipline required: update the code-before examples when the codebase changes. That’s it. That’s the whole thing. And it’s the thing most teams won’t do.

Risk 4: The selection mechanism is a black box

There is currently no telemetry available from Cursor or most other IDEs that tells you whether a specific skill was selected for a given agent session. You can observe the output, but not the mechanism. If a skill isn’t being selected, you won’t know it — you’ll just see no improvement and have to diagnose why.

Your CI pipeline can verify skill selection in controlled runs. In real engineering sessions, you’re partially flying blind. Test as many coding agents as you can in your pipeline, we do a parallel matrix in GitLab CI and test a bunch.

Risk 5: Agents constrained badly become unhelpfully narrow

A story from a FOSS Asia presentation, worth repeating: a team built a system where autonomous agents are triggered automatically on master branch CI failures to investigate and fix flaky tests. The system worked — but without an explicit constraint, the agent would identify a failing test, decide to refactor the surrounding code, delay the critical merge, and introduce unrelated risk. The team added a single constraint to their skill: “no unnecessary refactoring.” The agent stayed scoped to the job it was given.

Constraints in skills work both ways. The goal isn’t just to encode what you know — it’s to encode what the agent shouldn’t do in the process of applying what it knows.

Measuring Whether It’s Working

If you can’t measure the improvement, you’re running on faith.

The metric we use is first pass success rate: specifically, the proportion of agent sessions that require only a single pass versus multiple passes (re-prompting). Each agent invocation is a pass. A gap of more than five minutes between passes signals the engineer moved on to a new task — that boundary defines a session. One pass in a session is first-time-right. Two or more is iteration.

This matters because each extra pass is directly measurable time and compute cost. Token spend is logged per invocation with timestamps, requiring no new instrumentation. If skills reduce passes, the dollar impact is calculable immediately.

The known blind spot: engineers who fix the agent’s output manually rather than re-prompting don’t appear as multi-pass sessions. Based on our internal survey data, roughly 25% of engineers prefer to fix manually. That means the true re-work rate is likely higher than the session data shows. We acknowledge this explicitly — the survey provides a partial offset, but the metric understates the problem.

The experiment compares a treatment group (engineers with the skill library active) against a control group (same roles, no skills). A meaningful reduction in multi-pass sessions, sustained over roughly two sprints, is the signal we’re looking for. The direction matters more than the absolute number. If skills don’t move the needle, the post-experiment diagnosis looks at: whether skills were actually being selected (testable via the pipeline), whether the skills addressed the right categories of re-prompting, and whether the sample window was long enough.

This is the hypothesis-first, measurement-first discipline described in I Don’t Care What You Build. Articulate what you expect the skill to change before you write it. If you can’t, you’re not ready to write it.

The Doomsday Version

I want to end with the failure mode that keeps me honest.

We built an internal app — an MR tracker, AI-generated summaries for managers — by iterating with an AI assistant without any skills. The code worked. The app worked. But it was written in a dialect only the LLM could read fluently. The component structure bore no relationship to the rendered UI. The API calls were in the wrong places. The error handling was symptom-patching all the way down. The engineers who looked at it later understood none of it.

That’s manageable for an internal tool.

The version I think about at 2am: imagine that codebase is production. Imagine it’s 2am, there’s an incident, engineers are in the war room on level 7, and the system they’re trying to debug was vibe-coded by an AI without tribal knowledge. They open the file. They can’t read it because the rely on AI to not only write the code but understand it too. They call the AI for help. The AI gets stuck in a loop of, fix X, oh that didn’t work, I’ll fix Y, oh that didn’t work, I’ll fix X. You’ve seen it before right? What do we do now?

Skills won’t prevent every version of this. But they’re the difference between an AI that generates code that your engineers can read, reason about, and take ownership of — and one that generates code that only the AI understands, which is not ownership culture, it’s a different kind of vendor lock-in.

Where to Start

If you’re going to do one thing: write one skill. Pick the piece of knowledge your team repeats most often in code review or advice to juniors (not everything is in PR review comments, there’s so much more that's verbal) — the comment that appears so regularly it’s practically a meme. That’s your first skill. Write it down the way you’d say it to a new engineer on their first day.

Start with skills in your repo to fast iterate, then expand to mcp distribution later with a proper pipeline and test.

Then set up promptfoo and write tests for it. The discipline of writing the test will force you to articulate exactly what you’re asserting, which will make the skill itself more precise.

That’s a morning’s work. It’s also the beginning of a skill library that actually holds up under the pressure of a production codebase.

The rest is iteration.

References: