2026-09-16
Trust Should Be Paid for Once
An agent that has to rediscover who it is, what it may do, and whether the system around it is telling the truth will spend your budget answering questions the factory should have settled already.
A blank session starts in doubt
Give a fresh agent a task and it has more to resolve than the task itself. Who asked? Which repository is real? What rules govern the change? Which tools are safe? What happens if the action is wrong? A capable model can investigate those questions, but investigation costs time and tokens. Worse, it can reach a different answer on the next run.
That makes a naked session an expensive operating unit. The worker spends part of every shift rebuilding the workplace before useful production begins. If the surrounding records conflict, the agent either stops, guesses, or verifies everything again. None of those outcomes is free, and a bigger model does not remove the uncertainty. It only gives you a more capable system for reasoning through it repeatedly.
Steve Yegge calls the durable alternative a seat: a long-lived role that holds expectations, history, authority, and boundaries while different models occupy it. The useful idea is not the office metaphor. It is that trust can become an input the factory assembles once instead of a conclusion every worker has to derive from scratch.
The role must outlive the worker
A production station should tell the arriving agent what outcome it owns, which context is current, what it may change, which evidence counts, and when it must stop. Those facts belong to the factory. They should survive a new session, a cheaper model, a failed attempt, and the person who originally knew how the route worked.
This is more than a reusable prompt. Prompts describe behavior; systems establish authority. A durable role needs scoped credentials, tool contracts, work state, acceptance gates, and a record of previous decisions. If the text says an agent cannot release while the credential still can, the credential is the truth. If the role says a test must pass while the agent can edit the test after failing it, the boundary is fiction.
Our position is that the model should be replaceable without renegotiating the job. The role owns continuity. The model supplies capability for the current attempt. When those are tangled together, changing models becomes an organizational reset. When they are separate, models can compete for work against the same authority and the same proof.
Trust is a record with consequences
Trust does not mean telling the agent that everything is fine. It means giving it records that remain true when checked. The task points to the actual revision. The permission matches the declared scope. The gate evaluates the artifact it claims to evaluate. The recovery log accurately records which external effects already happened. The agent can act because the machinery has removed reasons to re-litigate the environment.
One false record changes the economics. If the factory says a command is safe and it is not, every similar claim becomes suspect. The next agent has a rational reason to inspect the wrapper, the repository, and the execution history before moving. A shortcut that saved one uncomfortable admission creates a verification tax on every later run.
That is why corrections need to stay visible. Preserve the rejected attempt, fix the rule or tool that permitted it, and record what changed. Do not rewrite history into a clean story. A factory learns when later workers can rely on an honest account of failure and a mechanical change to the route.
Controls can also consume the factory
Every incident creates pressure for another rule. That instinct is understandable, but an unchecked pile of refusals can make ordinary work impossible. Controls overlap. Exceptions multiply. The agent spends more of its budget proving that it may begin, and eventually the factory is safer only because it has stopped producing.
A strong factory governs its controls as aggressively as its code. Every fence needs an owner, a consequence it protects, and evidence that it still blocks a real failure. Duplicate rules should collapse. Obsolete ones should disappear. A warning that never changes the route should either become a gate or leave. The goal is not the largest rulebook. It is the smallest set of mechanical refusals that keeps consequences inside acceptable bounds.
Human judgment belongs here. People decide which risks justify friction and which tradeoffs the business will accept. Agents can inventory controls, find contradictions, replay known failures, and show what a proposed deletion exposes. They should not quietly decide that a business consequence no longer matters because a rule is inconvenient.
Make confidence cheaper than doubt
Start with one repeated route. Write down the role, the permitted surfaces, the required evidence, and the stop conditions. Put enforcement outside the model. Then measure how much of each run is spent finding context, checking authority, recovering state, and repeating verification the system already performed. That overhead is factory waste, even when it appears as thoughtful agent behavior.
Move stable answers into durable machinery. Cache context that is genuinely reusable, but give it an owner and an expiration condition. Keep authority narrow enough that confidence cannot widen the blast radius. Route expensive capability to the decisions that need it instead of paying the strongest model to rediscover basic operating facts.
This will remove work people perform today. Coordinators who repeatedly explain the role, assemble the same context, chase the same approval, and reconstruct the same failure are doing jobs the factory should absorb. Human judgment moves up to defining roles, setting consequences, and deciding when a control has earned its cost. Trust should not live in a reassuring sentence or one careful employee. It should live in machinery that lets every qualified worker begin from the same honest ground.
In response to Seats and Sunsets by Steve Yegge — essays.