2026-10-05Gunner Technology

Your Agents Need an Admission Gate

When agents create work faster than your system can accept it, more parallelism is not more capacity. It is a larger pile of unfinished decisions.

The backlog moved, but it did not disappear

Coding agents can open several workstreams while a person is still reviewing the first one. That feels like a capacity win because every stream looks active. Changes are being written. Tests are running. Pull requests keep arriving. Then the review queue grows, context switches multiply, and completed code waits for somebody to decide whether it is safe to become real.
The organization has not removed the bottleneck. It has moved the backlog from implementation into acceptance. Worse, the new backlog is harder to see. A list of unstarted tickets looks like unfinished work. A wall of plausible patches looks like progress, even when none of it has passed the decisions and evidence required to ship.
This is why generated output is a bad capacity measure. The useful unit is accepted change: work that met a decided requirement, survived independent checks, entered production safely, and produced evidence afterward. Everything else is inventory. More inventory can make a factory slower because every waiting change can go stale, conflict with another change, or consume another round of human attention.

Parallelism has a carrying cost

Starting work is cheap for an agent. Carrying that work is not. Each open stream holds assumptions about the codebase, the requirement, and the state of every neighboring change. The longer it waits, the more likely those assumptions drift. A fast agent can then spend its next turn repairing conflicts created by work the factory never had room to accept.
People pay the same carrying cost through attention. Ten patches do not create ten clean decisions. They create ten contexts to reload, ten sets of evidence to judge, and ten opportunities for one decision to invalidate another. If review depends on a few experienced engineers, agents can flood those engineers without delivering one extra customer outcome.
The answer is not to make the agents slower. It is to limit how much unfinished work the system admits. A factory should know how many changes each proof stage can carry, which work has priority, and what must finish before another stream begins. That limit is an operating control, not a failure of ambition.

Admit work against proof capacity

An admission gate asks a blunt question before generation starts: does the factory have a credible route to accept this change? The requirement must be decided enough to build. The evidence must be possible to produce. The relevant environment, permissions, and release path must be available. If the work will inevitably stop at a human checkpoint that is already full, starting it early does not create throughput.
This gate can be mechanical. Route low-consequence changes through deterministic checks and controlled release. Reserve human attention for product tradeoffs, unfamiliar architecture, irreversible operations, and risks the machinery cannot yet settle. When a repeatable issue reaches a reviewer, turn the correction into a rule or test so the next change does not buy the same decision again.
The gate should also be allowed to say not yet. That is important. Agent systems are often designed to look busy because idle machinery feels wasteful. But an agent waiting for capacity is cheaper than a reviewer sorting through stale work, and much cheaper than shipping a change whose evidence was rushed because the queue became politically inconvenient.

Measure the whole line

If you measure coding activity, your factory will optimize for coding activity. It will produce more patches, touch more files, and keep more agents occupied. None of those numbers tells you whether the line became better at turning a decision into a proven result.
Measure the time from admitted work to accepted change. Watch how long changes wait between stages, how often they return for missing decisions, how much reviewer attention they consume, and how often release or production sends them back. Those signals show where the operating system is weak. A growing review queue is not evidence that reviewers need to work harder. It is evidence that generation and acceptance are running at different speeds.
This changes the investment decision. You may not need another model or another agent. You may need clearer inputs, independent tests, smaller changes, safer permissions, or a release path that can judge routine work without a meeting. Improve the stage that limits accepted throughput, then raise the admission limit. Capacity should follow proof, not enthusiasm.

The factory sets the pace

Our position is that implementation will keep getting cheaper and more parallel. That will eliminate work built around translating settled instructions into code. It will also expose organizations that still depend on scarce people to inspect every machine-produced change by hand.
Human judgment still chooses the destination and the acceptable consequences. It should not serve as the universal queue between generation and release. Put repeatable standards into the machinery. Give proof its own stage and authority. Limit incoming work to what that system can honestly carry. Then expand the line by removing the next constraint instead of hiding it under more output.
A busy agent is not the goal. A busy reviewer is not proof of control. The goal is a factory that admits work deliberately and finishes it at the speed its evidence deserves. When generation outruns trust, the admission gate protects both throughput and judgment from a pile of work that was never ready to begin.
In response to Devs are coding faster. Coding reviews are eating the gains by CIO Dive - Latest News.