2026-10-05Gunner Technology

Quality Needs a Closed Loop

Agent-written code does not make quality disappear. It makes a weak quality system fail faster and a strong one learn faster.

More output reveals the quality system

When coding agents increase the amount of work entering review, every weakness downstream becomes visible. Requirements that leave room for three interpretations create three implementations. Slow test environments turn into queues. Optional review standards become arguments repeated on every pull request. Production surprises arrive faster because the factory is feeding them faster.
That does not prove agents produce worse software. It proves code generation was never the whole quality system. The system includes the request, the checks, the release decision, production observation, and the route back to the next change. Speed at one stage increases pressure on all the others.
Our position is simple: do not slow the agents until the old process feels comfortable again. Repair the machinery that cannot carry the new volume. Quality should scale with production, not depend on a shrinking amount of human attention spread across a growing stream of changes.

Each layer needs a different job

A specification review catches missing behavior before implementation. Unit tests protect local rules. End-to-end checks exercise the paths a user actually touches. A code review can spot a bad boundary or an unnecessary abstraction. Production monitoring catches the facts nobody predicted. These are not interchangeable badges called quality.
Trouble starts when a team asks one layer to compensate for another. High coverage cannot prove the requirement was right. Manual testing cannot protect every release at agent volume. A second model reading the same diff under the same assumptions is not independent evidence. A dashboard nobody routes into work is just a picture of decay.
Give every layer a named failure it is meant to catch and a signal the factory can evaluate. If two checks claim the same job, remove the weaker one or change its conditions. If a serious failure has no layer, add one before increasing throughput. The point is not to accumulate checks. The point is to close specific escape routes.

The builder cannot own the verdict

Agents make tests cheap to write. That is useful, but cheap tests can still preserve the agent's misunderstanding perfectly. If the same worker interprets the requirement, writes the implementation, invents the test cases, and declares the result correct, the whole chain can agree on the wrong thing.
Independent proof does not require a human to inspect every line. It requires conditions the builder cannot quietly redefine. Derive acceptance checks from the approved behavior before implementation. Run important paths in an environment the coding agent did not assemble for the demonstration. Enforce structural rules with deterministic tools. Route high-consequence changes through a separate authority.
Human judgment still matters, but it belongs at the boundary of consequence. People decide what the system must protect, which tradeoffs are acceptable, and when evidence is strong enough to ship. Spending that judgment on formatting, repetitive test setup, or comments a rule could reject is not caution. It is wasting the scarce part of the system.

Production must write back

Pre-release checks only cover what the factory already knows how to ask. Real use supplies the cases nobody named: an unexpected sequence, a dependency that degrades instead of failing, a technically valid action that creates the wrong business result. Monitoring is where those unknowns become visible.
Visibility is not the finish line. A production failure should create a durable change in the factory: a sharper requirement, a regression test, a tighter permission, a new route, or a changed release gate. If the response ends with one repaired incident, the factory paid for a lesson and then threw it away.
This is the difference between a stack of defenses and a closed loop. In a closed loop, evidence from the last release changes how the next request is specified, built, checked, and watched. The system gets harder to surprise. Without that write-back path, every layer begins each run with the same ignorance.

Quality has to become machinery

Teams often describe quality as a culture, which usually means good people are expected to remember good habits under pressure. That approach was fragile when changes arrived at human speed. It collapses when agents can produce more work than the same people can personally supervise.
Turn the habits into operations. Make incomplete requirements fail before build work begins. Make required checks impossible to skip. Record why an exception was allowed and when it expires. Measure which gates catch real defects and which only consume time. Let low-risk work move without ceremony, and make dangerous work stop at an authority boundary that cannot be routed around.
Agent-run delivery can produce better software, but not because the model cares more about quality. It can produce better software because the factory can apply more checks, more consistently, and carry every useful correction forward. The organizations that build that loop will increase output without surrendering control. The ones that rely on attentive reviewers will discover that attention does not scale.