2026-10-05Gunner Technology

Model Choice Belongs After the Gates

The expensive model may break less. That does not make it your safest default. It means your factory has another routing signal.

A careful model is not a safety system

A model that causes fewer regressions is valuable. Buy that property when the work needs it. But do not confuse a statistical tendency with a control. The model can still miss part of a change, misunderstand a boundary, or spend more effort without producing more complete work. Carefulness lowers one kind of risk. It does not decide whether the result is safe to ship.
That distinction matters because teams keep making model selection carry responsibilities that belong to the delivery system. They choose the strongest option they can afford, give it a ticket, and hope capability will cover weak tests, vague scope, and thin review. When the change lands cleanly, the bet looks smart. When it fails, there is no mechanism to explain which protection was missing.
The factory has to assume every model is fallible in its own repeatable way. Tests catch known breakage. Independent review catches mismatches the builder rationalized away. Scope controls limit the blast radius. Release gates decide whether the collected evidence is enough. A more cautious model can improve the route, but it cannot replace the route.

Route by consequence, not reputation

The cheapest acceptable model is not the cheapest name on a price card. It is the least expensive route that can produce the required evidence at the required reliability. Sometimes that means a lower-cost builder behind strong tests and a separate review agent. Sometimes it means paying more for the builder because the code has sparse coverage, the run is unattended, or the cost of a regression is too high to absorb casually.
Those are operating conditions, not opinions about which model is smartest. A narrow, well-specified change with a mature test suite can tolerate different behavior from a wide change that crosses workspaces and leaves important outcomes unobserved. Sending both through the same model at the same effort level is not standardization. It is refusing to route.
Put the decision in machinery. Classify the change by affected surfaces, available proof, reversibility, and consequence. Then assign the builder, effort budget, reviewers, and release authority that match that class. The model earns a place on a route by its measured performance there. It does not earn permanent authority from a launch benchmark or a good week in somebody else's repository.

Thoroughness needs a map

Wide changes expose the limit of model shopping. If the ticket names twenty-six relevant files but the system gives the agent no map of the behavior crossing them, extra tokens can become a longer walk through the same incomplete understanding. The agent may work carefully inside the area it noticed and still leave the actual change unfinished.
The answer is not a heroic reviewer reconstructing the whole system at the end. Give the factory a change map before implementation begins. Name the contracts that move together, the consumers that must agree, the old behavior that must remain true, and the evidence that proves the transition happened everywhere. Break the work into pieces small enough to verify, then make completion depend on the map rather than the agent's confidence.
This is where human judgment belongs. People decide which consequences matter and which omissions would make the change unacceptable. Once those decisions are explicit, agents can perform the search, implementation, and proof repeatedly. Leaving the map in a senior engineer's head turns that person into a manual dependency and guarantees the next run starts ignorant again.

Measure the whole route

A model bill tells you what generation cost. It does not tell you what the change cost. Add the failed attempts, review work, test time, escaped defects, rollback effort, and delay before the result reached production. A cheaper builder that requires a strong review pass may still win. A premium builder that avoids repeated recovery may be cheaper on a fragile route. You cannot know by comparing token prices alone.
Run evaluations on your own work and keep the conditions visible. Record the task class, repository state, permitted tools, effort setting, gates, reviewer, and final outcome. Re-run important routes because model behavior changes. Promote a model when it improves the economics of a route, and demote it when the evidence moves. The name in the model picker is configuration, not strategy.
Our position is simple: choose the factory first and the model second. Build the proof, authority, and recovery path that the work deserves. Then use model differences to tune cost and risk inside those boundaries. Companies that do this will swap models without relearning how to ship. Companies that crown a winner will keep discovering that they outsourced their operating system to a moving target.
In response to Careful, Not Thorough: Claude Opus 5.5 vs GPT-6 Sol on Real Code by Emergent Minds | paddo.dev.