2026-09-28Gunner Technology

Rework Belongs in the Productivity Number

If an agent writes a change in ten minutes and the organization spends the afternoon repairing it, the ten-minute result is not productivity. It is an invoice sent to the next stage.

Gross output is not net progress

Most productivity measurements begin where the new tool is easiest to see. How quickly did it produce a patch? How many suggestions did someone accept? How much work appeared to leave the coding queue? Those numbers describe production, but they stop before the organization finds out what it bought.
The work can return as review comments, failed checks, reopened defects, support incidents, or a second implementation that replaces the first. It can also return quietly, through an experienced person rewriting the risky parts before anyone records a failure. Every one of those paths consumes capacity. Leaving them outside the measurement does not remove the cost. It only rewards the system for hiding it downstream.
A factory needs a net number. Start with accepted outcomes, then subtract the repair, rescue, and repetition required to keep them accepted. Code volume can rise while that number falls. That is not an inconvenient edge case. It is the first condition a serious productivity system should be built to detect.

Name the return path

Rework is too useful a signal to collapse into one bucket. A change rejected because the requirement was unclear points to a different failure than a change rejected for breaking an architectural boundary. A production rollback is different again. If the factory records all three as extra time, it knows the bill but not what to fix.
Attach the return to the stage that created it and the reason it came back. Missing intent belongs at the specification gate. A repeated security mistake belongs in an executable control. A change that passes local tests and fails under real conditions needs stronger independent proof. The label matters only if it sends the next attempt through a different route.
Do not let a successful rescue erase the original miss. When a person repairs the patch, preserve both attempts. When a stronger model takes over, record why the cheaper route failed. When a reviewer catches the same pattern for the fifth time, stop calling that review. It is an unbuilt factory rule wearing a person's calendar.

Compare routes, not impressions

A productivity claim needs a fair comparison. The question is not whether an agent can finish a selected task quickly under supervision. The question is whether one complete route delivers more durable outcomes than another under the same standard. That route begins with usable intent and ends after the result has survived the conditions that matter.
Keep the classes of work visible. A routine dependency update and an ambiguous product change should not share one average. Neither should a reversible presentation fix and a permission boundary. Compare similar work, include waiting and intervention, and hold the acceptance bar steady. Otherwise the measurement will credit the agent for receiving easier work or quietly lower the definition of done.
This does not require a universal productivity formula. It requires an honest local ledger: what entered, what was accepted, what returned, why it returned, what it cost, and whether the correction held. That evidence is enough to decide which route deserves more authority and which one is only creating faster inventory.

Make the metric change the machine

A dashboard that reports rework without changing the factory is another queue. Set operating consequences before the results arrive. If one failure class repeats, add a gate. If a route survives with little intervention, expand its authority. If repair cost crosses the boundary you set, narrow the work, buy stronger reasoning, or stop the route until its controls improve.
The same mechanism keeps the measurement from becoming a performance score for people. Ranking developers by accepted suggestions teaches them to accept suggestions. Ranking teams by output teaches them to move cost beyond the reporting boundary. The factory should measure routes and controls because those are the things the organization can actually redesign.
People still choose what counts as an acceptable outcome and which consequences are tolerable. They should not spend every week rediscovering the same rejection by hand. Once a judgment repeats, encode it. The useful metric is the one that moves that judgment out of a meeting and into machinery that applies it on every run.

Productivity is what survives

Agents will make software production look extraordinarily busy. More attempts will start, more code will appear, and more changes will reach the first checkpoint. Some organizations will call that productivity and hire people to absorb the repair. Others will build factories that reject weak work early, retain the reason, and improve the next route.
Our position is that the second group will need fewer people for coordination, routine review, and cleanup. That is not because mistakes disappear. It is because each repeated mistake becomes a control instead of a permanent job. Human judgment moves toward setting the destination, pricing the consequences, and changing the system when its evidence says the route is wrong.
Measure what survives, not what appears. Keep the return path visible. Charge every rescue to the route that needed it. Once rework sits inside the productivity number, the impressive demo and the useful factory stop looking like the same thing.