Spec-driven development means writing the spec before an AI agent writes any code, then keeping that spec as the thing everyone builds against. It’s a good practice. It also has a gap. The spec tells the agent what to do, and nothing checks that it did.
Instead of prompting an agent and correcting it as you go, you write down what to build first: the behaviour, the acceptance criteria, the constraints. The agent reads that and builds from it. When the code and the spec disagree, the spec wins.
Tools like Kiro, GitHub’s Spec Kit and Tessl have made the idea popular over the past year. Plenty of teams do it without any tool at all. They keep the spec in the repository and point the agent at it.
Birgitta Böckeler at Thoughtworks tried the main spec-driven tools and found the same thing in each. The spec is a document the agent is meant to read, and nothing compares the code against it. She “frequently saw the agent ultimately not follow all the instructions.” (martinfowler.com)
That matches what happened to us.
We build Traceway this way, and strictly. Nothing is built without an agreed spec. Each ticket has its own end-to-end test. A ticket’s original criteria are frozen when work starts, and any difference from what was agreed counts as a defect.
The agents had all of that and drifted anyway. They:
Of our last 120 merged pull requests, 53 carried staged proof or duplicate work. One carried 145 such files. Every one of them was forbidden in writing, in instructions the agent had read.
The agent hadn’t misread anything. It read the spec and did something else.
It helps to think about agent guardrails in three layers.
The spec, plus files like AGENTS.md or CLAUDE.md. They tell the agent what you want. They can’t stop it doing something else.
Checks that run as the agent works and can stop a step before it happens. Only some assistants allow these.
A check in your CI that reads the finished work and fails the pull request. It doesn’t care which assistant wrote the code, or whether a person did.
Spec-driven development lives in the first layer, which is why it informs and can’t enforce. If something genuinely has to hold, give it a check in the third layer. That’s the only one every assistant has to pass.
A spec answers what to build. It rarely says why, who agreed to it, or what it was supposed to achieve. Those usually live in a meeting, a thread or someone’s head.
A spec is also rarely one document. Behind a single ticket there can be a requirement, a design, an architecture call and a data or security rule, each agreed by someone different.
In Traceway each of those is a decision of its own type, with who approved it and why. Every decision links up to the strategy it serves and down to the work it triggers.
So the spec an agent builds from starts at an approved decision, whatever its type. The work can be checked against what was actually agreed, and traced back to the strategy that asked for it. Anything beyond it goes back to the person who approved it, as a proposal.
All three describe what should happen before the code exists. TDD writes it as tests. BDD writes it as plain-language scenarios that run as tests. Spec-driven development writes it as a document an AI agent builds from.
The difference is who’s implementing, and that’s also the weakness. A test fails when the code is wrong. A document doesn’t.
Put the check where the code has to pass anyway, which is the pipeline. Have it compare the work against the spec and the decision behind it, and fail the pull request when they disagree.
Start in warn mode, so you see what it would have caught before it blocks anyone. And make every failure say what to do instead. A check that can only tell engineers they’re wrong gets switched off.
That’s what Traceway for agents does. It’s in testing on our own build now and arrives in the product at the end of Q4 2026 or early Q1 2027. See how it works →
Yes. A written spec beats a prompt that’s gone the moment you send it. Just don’t assume the agent followed it.
The practice works with any of them. How closely each one sticks to the spec varies, which is the case for checking the result instead of trusting the instructions.
Those files are instructions. The agent weighs them against everything else in its context, and sometimes something else wins. Keep them. Just don’t treat them as a control.
People use both. “Design” usually means the spec includes screens and states as well as behaviour. The gap is the same either way.
Traceway checks your agents’ work against the decision behind it, in any assistant. Request a demo.
The first five customers join as Founding Customers, chosen by fit. Data is held in the United States, or in the European Union for a workspace set up there. Request a demo from anywhere.