The Economics of Open-Weight Inference
How open-weight demand can support the useful life of NVIDIA GPU families
GPUs are depreciated as if each NVIDIA generation retires the last. In The Economics of Open-Weight Inference, we test that assumption against model costs, hardware throughput, and Ornn’s rental data.
Open-weight models make inference portable. Long-running agents, batch evaluation, and reinforcement learning tolerate latency, so operators can send that work to whichever GPUs complete it most cheaply. Often, that’s older hardware.
Across eleven open-weight and eight closed models, the cheapest qualifying open-weight model completes a benchmark task at roughly one fifth of the cost of a comparable closed model. On gpt-oss-120b, the A100 produces output more cheaply than the H100 at spot and at three- and five-year term prices. Ornn’s five-year A100 term mark keeps 80% of its one-month price, against 44 to 60% for Hopper and Blackwell.
A new generation doesn’t make its predecessors obsolete. Older GPUs keep an earning role while they serve useful work competitively. The rental data come from the Ornn Compute Price Index and forward marks; the index methodology is published on Ornn Data.
@techreport{ornn2026openweight, author = {{Ornn Data}}, title = {The Economics of Open-Weight Inference}, institution = {Ornn}, year = {2026}, month = {sep}, url = {https://ornn.com/research/the-economics-of-open-weight-inference}}