Skip to content
Back to the blog

The spec is the source code, the code is the binary

When an agent generates the implementation from a specification, the spec becomes the artifact you version, review, and treat as source, and the generated code behaves like the binary a compiler emits. This essay argues that the architect stops reviewing the disassembly and starts reviewing the source, and that the review effort and the governance therefore have to move to the spec, with the guardrail layer as the place where its guarantees get compiled in.

When an agent writes the implementation from a specification, the question of what you are versioning changes its answer without anyone announcing it. You still open pull requests and comment on lines, but the diff you review is no longer of what you wrote: it is of what the agent produced when it read what you wrote. The spec is what you touched with your hands. The code is what came out of the machine. And that distinction, which sounds like a nuance, is exactly the same one we have been making for forty years between source code and the binary.

That is the thesis, plainly: the moment an agent generates the code, the spec becomes the source code and the code becomes the binary. What you version, review, and treat as the truth is the source; what gets compiled from it is a regenerable artifact you almost never look at. The problem is that the industry still reviews the binary as if it were the source, and that has a name: it is opening the disassembler to approve a change that is written on a line of C three folders up.

Why the compiler metaphor and not another?

Gregor Hohpe has a rule for telling whether a metaphor helps you think or only decorates, and he lays it out in The Mighty Metaphor: it is not enough for the image to look alike, the dynamics have to match. His full method, translating a problem into a domain where the reader already knows how to reason, is on his site. Let us apply his test, because the compiler is a dangerously comfortable metaphor and it is worth seeing whether it holds.

In a compiler, the source-to-binary relationship has rules that are not up for debate. You debug and version the source, not the binary. If you edit the binary by hand to fix something, you have fixed nothing: you have created a divergence the next build will erase, and now your source lies, because it no longer reproduces what runs in production. The binary is regenerable by definition; its whole value is that it comes out of the source as many times as needed. And a source you cannot compile back into that binary is not source: it is a long comment.

Swap "source" for "spec," "compiler" for "agent," and "binary" for "generated code," and the four rules hold one by one. You debug and version the spec. Patching by hand the code the agent generated is editing the binary: it works until the next regeneration, and meanwhile your spec no longer describes what runs. The generated code is regenerable, that is its entire point. And a spec you cannot regenerate the code from is not a spec, it is documentation with delusions of grandeur. The dynamics match, not just the image. That is why the model holds.

And because it holds, you can push it further than I am going to write here. If the generated code is the binary, then reviewing its diff line by line is diffing binaries, an exercise the classic world reserved for reverse engineering and nothing else. If the spec is the source, a spec without tests is source without a type checker: it compiles the same and fails just as late. I have argued neither of those and you are already seeing them. That is what separates a model from an ornament.

So is the spec design, or is it paperwork?

Here it helps to bring in someone who won this argument twenty-five years ago, before a single agent existed. Building on the argument Jack Reeves makes in What Is Software Design?, source code is not the construction of software: it is its design. What actually manufactures the product, Reeves holds, are the compiler and the linker, and the only documentation that fully satisfies an engineering design is the code listings themselves. Manufacturing, in software, is free and automatic; the engineering work lives entirely in the source.

That argument, which in its day served to dignify code against UML diagrams, fits like a glove in the agent era, only shifted up one level. If the compiler is what manufactures, and the agent now takes the compiler's place, then the spec takes the source's place: it is the design, the place where the engineering work lives. The generated code is the manufacturing, and manufacturing is once again free and automatic, which is exactly what turns it into inventory rather than an asset, as I argued in Code is inventory, not an asset.

What Reeves could not foresee, and where you have to add something of your own instead of translating him, is the architectural consequence. If the spec is the design, then the review effort and the governance have to move to the spec, because reviewing the binary was never the job. For decades we reviewed the source and not the binary, not out of habit, but because the source was the only place a design decision was legible. The agent does not change that logic: it makes it literal. The place where a decision is legible is now the spec.

Two models side by side. In the classic one, review lands on the source code; in the agent model, the habit keeps landing on the generated code instead of on the spec.

On top, review lands on the source, where the design lives. On the bottom, the habit has not moved: it keeps reviewing the binary.

How much does it cost to keep reviewing the binary?

Let us run the numbers, with the assumptions in plain sight so you can redo them with yours. This is not a real client's measurement; it is napkin arithmetic, the kind this blog does.

Suppose an agent generates 1,500 lines of code for each task you hand it, from a spec of 150 lines. The ratio is ten to one, and it is not far-fetched: describing the what almost always takes less than laying out the how. Now scale it to a team: five tasks a week are 7,500 lines of generated code against 750 lines of spec.

If you review the code, you review 7,500 lines a week that nobody wrote expecting a human to read them, because the agent optimizes for compiling, not for reading. If you review the spec, you review 750 lines that were written to be read and that hold the real decisions: a tenth of the volume, in the place where an error matters before it materializes. The point is not only "I save 90% of the review time," though you do; it is that the 750 lines of spec decide the 7,500 of code, and the 7,500 of code decide nothing, they only obey. Reviewing the effect instead of the cause is working ten times as hard to arrive late.

But I need to review the code because it is what runs in production, not the spec. The agent gets things wrong, hallucinates, pulls in a weird dependency. If I only look at the spec, that error slips past me.

It is the right objection, and it still conflates two things. That the error lives in the binary does not mean it gets fixed in the binary. When a compiler emits bad code, you do not patch the assembly: you fix the source, or you fix the compiler. If the agent hallucinates from a clear spec, the problem is the agent and it goes against the layer that governs it; if it hallucinates because the spec was ambiguous, the problem is the spec and it gets fixed there. Reviewing the generated code to catch the error is useful in exactly the same way reading the disassembly is when debugging: you do it when you suspect the compiler, not as the place where your work lives. Turning the exception into the method is what does not hold.

If your way of assuring quality is reading line by line what the agent generated, you are not reviewing code: you are reverse-engineering your own spec.

Two panels with the same illegible dump of bytes. The left one is labeled "reviewing the binary"; the right one, "reviewing the agent's code." They are identical.

Where do the spec's guarantees get compiled?

A spec that only lives in a document is a source nobody compiles: it expresses an intent and trusts the agent to honor it, which is about as practical as asking your developers to write bug-free code because you asked them nicely. For the spec to be real source, you need a layer that turns its boundaries into something the agent cannot cross even if it wants to, the same way a type system turns a programmer's intent into a compile error instead of a prayer.

That layer exists and has two faces. The spec-driven flow of Kiro treats the spec as the first artifact: you write it, you version it, and the code comes out of it, not the other way around. And a platform like Amazon Bedrock AgentCore is where the boundaries of that spec get compiled into operational guarantees: identity, permissions, and guardrails that define what the agent can touch before it acts. Between the two, the spec stops being a wish list and becomes executable source: what it says gets enforced, not hoped for. That is where the architect puts the effort, because that is where a decision takes effect. It is the same shift of control toward the boundary I described in Autonomy is a blank signature: you do not police every action, you define the boundary beforehand.

My suggestion

On Monday, take a single question to your next architecture review: what are we versioning as the source of truth, the spec or the code the agent generated from it? If the answer is "the code," you have a project that treats the binary as source, and it shows in the fact that nobody could regenerate it from scratch, because the spec, if it exists, no longer reproduces what runs. If the answer is "the spec," then the next question is where its guarantees get compiled, and who reviews those 750 lines with the care they used to spend on the 7,500.

And if you want a concrete test, run it this week: take a task the agent already solved, delete the generated code, and regenerate it from the spec alone. If it comes out equivalent, your spec is source and you can trust it. If it does not, you have just discovered that what you took for source was a comment, and that the truth lived in the binary you were about to delete. I concede it is a little vertiginous to delete working code to see whether it comes back. But if you really believe your spec is the source, regenerating the binary should be a piece of cake. ;)

References

  1. The Mighty Metaphor — Gregor Hohpe
  2. The Architect Elevator — Gregor Hohpe
  3. What Is Software Design? — Jack Reeves
  4. Kiro — spec-driven development
  5. Amazon Bedrock AgentCore — AWS