2026-10-05Gunner Technology
Build the Factory Backward From the Verdict
If agents can write more code than you can safely judge, another coding agent is not the next thing to build. The verdict is.
Start where shipping stops
Most teams begin automation at the exciting end. They give an agent a ticket, let it write the change, and celebrate when a pull request appears. That proves the agent can produce an artifact. It says almost nothing about whether the artifact deserves to reach a customer.
The useful starting question is harder: what evidence would let this change ship without a person rebuilding the whole argument from scratch? That question exposes the actual factory. You need a known input, an observable outcome, a comparison that cannot be quietly rewritten by the builder, and a release gate that acts on the result.
Build that verdict first. Then work backward toward the code. A factory designed this way does not treat verification as cleanup after generation. Verification defines the shape of the work the generator is allowed to attempt.
History can become a test rig
Running systems already contain useful questions. Past requests show the inputs the software actually received. Stored decisions, confirmed outcomes, and known failures show what happened next. When those records can be replayed safely, they give a candidate change conditions it did not get to choose.
Put the current version and the candidate version against the same bounded history. Compare their behavior. Read the changed cases instead of trusting one summary score. Run the current version twice when the system is nondeterministic, so ordinary noise does not masquerade as improvement. Keep the replay read-only and remove identifying data before it enters the test rig.
Replay is not truth. An old decision can be wrong, and a simulator can preserve the same bug as the system it judges. That is why the result needs an independent challenge and why intended changes need reviewed expectations. History gives you a demanding witness, not an infallible judge.
The builder does not write the answer key
An agent can write a feature, its tests, and a beautiful explanation of why both are correct. That combination feels complete because every piece agrees. Agreement is exactly the problem when all three pieces came from the same interpretation.
Expected behavior must come from somewhere the builder cannot revise to make the run pass. It can come from an approved specification, a stable production record, a hidden label, or a scenario maintained outside the implementation path. The mechanism matters more than the format. The answer key needs separate authority.
This does not mean a person must inspect every line. People choose the consequences worth protecting and approve changes to the expected behavior. Machines can perform the replay, comparison, routing, and enforcement every time. Human judgment stays at the boundary where a changed expectation becomes a changed promise.
Make proof part of the route
A one-off evaluation is useful during an experiment. A factory needs the same protection on an ordinary Tuesday when nobody remembers the clever script. Move the replay into the release route. Record the inputs, versions, allowed differences, and result. Fail closed when the evidence is missing or stale.
The gate should price the decision in consequences the business can understand. A change may reduce one kind of error while creating another. A new model may improve a score while costing enough to make the route uneconomical. A threshold may catch more failures while flooding the response path with noise. The winning candidate is the one whose whole result is acceptable, not the one with the prettiest aggregate.
When a candidate loses, kill it. Cheap code makes abandoned implementation affordable. Keeping a weak change because the team already built it is human project accounting leaking into a machine-scale system.
Proof is the durable asset
Models change. Agent tools change. The code they produce will increasingly be replaced instead of cherished. The durable part is the route that can decide whether a replacement preserves what matters and improves what should change.
Our prediction is that software teams will build less of their advantage around writing code and more around operating these verdicts. The repeatable work will move to agents. The people who remain will define acceptable outcomes, challenge the evidence, and decide when the factory is allowed to change its own rules. Roles centered on manually producing and checking every change will keep shrinking because the verdict can become machinery too.
Do not begin with the fastest builder you can buy. Begin with the decision you are unwilling to fake. Make that decision executable, independent, and hard to route around. Then connect specifications and agents to it from the other direction. The factory starts producing safely when every new piece of code is born with a way to lose.
In response to The Future of AI Is CI: Building a Software Factory From the Back by Emergent Minds | paddo.dev.