2026-08-26
Your Model Contract Is a Decaying Asset
The model capacity you locked in today can keep working perfectly and still become a bad deal. Capability moves, prices fall, and your contract does neither.
Two curves set the trap
Tim Davis calls the gap the token curve. On one side is the market price for completing a standard unit of inference work. On the other is the provider's cost to produce that work. Those curves don't move together. Competition can push the selling price down while financing, power, operations, and long compute commitments keep the production floor in place.
That distinction matters because demand can look enormous while the economics underneath it get worse. More tokens served does not guarantee more value per token. More infrastructure spending does not prove that the work produced by that infrastructure will hold its price. Volume, revenue, and capital commitments are three different measurements, and treating them as one hides the risk.
The uncomfortable part is duration. A provider can reserve years of capacity based on today's hardware and today's prices, then sell into a market where comparable capability gets cheaper every quarter. The machines have not failed. The contract has. Its cost is fixed while the value of its output keeps being marked down.
Portability turns stranded machines into competition
Different chips, generations, sites, and serving architectures already exist. They do not become useful substitutes just because they exist. A workload has to move between them, run at an acceptable fidelity, meet its service level, and produce evidence that the result is equivalent. Without that portability layer, spare capacity is only spare capacity on somebody else's island.
Once the layer works, the market changes. A buyer can compare qualified routes instead of buying a brand name or a single machine type. Older hardware can remain useful for the work it still performs economically. Specialized systems can take the phase they handle best. Capacity that was difficult to reach becomes supply, and new supply presses on the price of every locked commitment around it.
This is not a prediction that every chip becomes interchangeable. The opposite is more useful: every route earns a narrower authority. The system knows which workload, latency class, quality threshold, and operating conditions a route has actually passed. Portability comes from measured compatibility, not from pretending the differences disappeared.
The same curve reaches your software factory
Most software teams will not finance a data center. They will still inherit this market through model contracts, reserved throughput, platform commitments, and architecture built around one provider's assumptions. The danger is not only paying too much. It is allowing a purchasing decision to become a permanent routing decision inside the delivery system.
A software factory should buy capability, not loyalty. Planning, implementation, review, migration, and incident work do not need the same model or the same serving route. Each station should define the work it needs, the evidence required, the authority granted, and the current cost of an accepted result. Then the factory can move repeatable work when a cheaper qualified route appears.
The model is rarely the durable advantage. If swapping it breaks the line, the value was trapped in an integration. The durable asset is the harness that assembles context, limits authority, grades output independently, records failures, and routes the next attempt. That machinery lets falling inference prices become lower production costs instead of a quarterly rewrite project.
Price the accepted result
A cheap token route can still be an expensive work route. One model may use fewer dollars per call and trigger more retries. Another may cost more at generation and clear verification on the first pass. A third may be excellent at code transformation and unreliable at deciding whether the transformation belongs. Price alone cannot route the factory because the factory does not sell tokens. It produces accepted work.
Measure the full attempt: retrieval, planning, generation, tools, verification, rejected output, escalation, and recovery. Divide that cost by results that passed independent gates. Now a new model, hardware route, or provider can audition against the unit the business actually consumes. If it produces the same accepted result for less, move the work. If it only makes the meter look better, leave it outside the line.
This is also how portability stays honest. The builder cannot declare itself equivalent. Tests, policy checks, runtime evidence, and service-level measurements have to survive outside the route being evaluated. A boundary the vendor can redefine whenever the result misses is not a qualification standard. It is marketing copy with a benchmark attached.
Move judgment above the curve
Our position is that inference will keep getting cheaper for equivalent work, and factories will use that decline to automate more software delivery. Jobs built around repeatable production will disappear as the cost of assigning that work to agents falls. Preserving an expensive route to preserve the work around it will not stop that shift. It will only make the organization carrying it less competitive.
Human judgment still decides what deserves to be built, which consequences are acceptable, and what proof is strong enough to ship. Put those decisions above the routing layer. Everything below them should be able to move when capability and economics move. The factory should retain the specification, the controls, and the evidence while workers change underneath it.
Treat every long model commitment as a position that can lose value. Keep its authority narrow. Keep alternatives qualified. Keep the proof independent. The token curve is not just an infrastructure story. It is a warning that the worker you chose is temporary, even when the contract says otherwise.
In response to The Token Curve by Tim Davis | Blog.