2026-10-03Gunner Technology
The Finish Line Is a Factory Interface
An agent that can work for hours does not need a longer prompt. It needs a finish line the factory can inspect.
Longer runs change the failure
Coding agents are getting better at carrying a whole task instead of handing back one clever fragment. You can give one a migration, an audit, or a feature and let it keep moving across files, tests, and review feedback. That is useful. It also changes what goes wrong.
A short assistant session usually fails in front of you. The answer is incomplete, so you ask another question. A long autonomous run can fail behind a polished summary. It may stop after the easy half, declare success when one check passes, or keep working after the useful job is finished. More endurance makes the boundary around the work more important, not less.
The practical response is to define done before the run starts. Name the artifact that must exist, the old path that must be gone, the checks that must pass, and the conditions that require a stop. That sounds like prompting advice. In a software factory, it is interface design.
Done must be observable
"Make the migration complete" is an intention. "Every endpoint uses the new client, the old client is absent, and the test suite passes" describes observable state. The second version gives the agent something to pursue and gives the factory something to verify after the agent says it arrived.
That last part matters. The builder's statement that the work is complete is evidence, but it is not the gate. Completion should be recomputed from the repository, the exported application, the test environment, or whatever system owns the truth. If the agent can satisfy done by writing the report that defines done, you built a ceremony around self-approval.
A good finish line also survives a change of model. It does not depend on one model recognizing the tone of your request or remembering a convention from earlier in the chat. The contract stays stable while the worker underneath it changes. That is how capability becomes replaceable without making the outcome negotiable.
Blocked is a routed state, not a mood
Long runs need an equally clear definition of blocked. An agent should not stop because it found something surprising, reached the end of a checklist, or has an update worth sharing. Those are normal states in the loop. It should stop when progress requires authority, information, or a destructive action it does not have.
This is where human judgment belongs. People decide which consequences require approval and which unknowns deserve escalation. They should not sit beside the run to answer every nonblocking question or repeatedly say "continue." Once the boundary is explicit, the factory can route routine uncertainty back into the work and reserve the human interruption for a real decision.
Status updates still matter, but they should travel beside the work instead of replacing it. A useful update says what changed, what remains, and what the system is doing next. A summary that ends with an obvious next step is not coordination. It is a manual handoff the factory failed to remove.
The work must outlive the session
The longer a run lasts, the less you can treat the conversation as the operating record. Context gets summarized. Processes restart. Another change lands first. A person redirects the task halfway through. If the plan, decisions, checks, and remaining work exist only in scrollback, recovery becomes an exercise in plausible reconstruction.
Put durable state where the next worker can read it. Keep the task list in an artifact. Record which gates passed. Preserve the reason for an approved exception. When a follow-up changes the destination, update the contract instead of relying on the latest message to overpower everything that came before it.
This is not paperwork for the agent. It is what lets the factory resume after interruption without inventing history. The transcript can explain the journey. The artifacts must say where the system actually is.
Autonomy is a control problem
A more capable model can think longer, use more tools, and recover from more local mistakes. None of that answers the operating questions: Who decides the destination? What proves arrival? Which uncertainty is safe to resolve automatically? Which action must stop at a human gate? What state survives if the run disappears?
Our position is that these questions will decide whether longer-running agents remove coordination work or merely hide it inside longer sessions. Teams that encode the answers will let agents carry whole outcomes. Teams that do not will hire people to watch the agents, restate the next step, and clean up after ambiguous finishes.
The model supplies endurance. The factory supplies the finish line, the stop conditions, the evidence, and the recovery path. Once those are explicit, you can hand over the whole task without handing over the authority to decide what success means.
In response to Getting the most out of Opus 5.5 in Claude and Claude Code by Claude.