
Ornn Data paper: open-weight demand keeps older GPUs earning
A new Ornn Data paper argues that a new NVIDIA generation does not by itself make its predecessors economically obsolete, and that open-weight inference gives older GPU families a multi-year earning life.
Ornn Data has published "The Economics of Open-Weight Inference," a paper arguing that a new NVIDIA generation does not, by itself, make the previous one economically obsolete. It examines how demand for open-weight models affects the useful life of older GPU families, from Ampere to Blackwell.
Hosted open-weight models and cost per task
Closed-model access runs through subscription allowances that a provider can reset, so a posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth the cost of a comparable closed model. The paper limits that claim: open models were cheaper at several sampled score thresholds, not every one, and the closed frontier still leads at the top of the range.
Self-hosting on rented hardware
Self-hosting has no posted per-token price. Ornn computes it as the GPU-hour rent on its spot index divided by the throughput an operator achieves, adjusted for utilization and reserved headroom. At full utilization the paper puts compute-only cost at $0.12 to $0.35 per million output tokens. The ranking reverses with the workload: on gpt-oss-120b, a sparse mixture-of-experts model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices, while the dense Llama-2-70B comparison favors newer hardware. Ornn notes that the A100 and H100 sparse inputs come from different third-party serving setups and that the dense A100 row is estimated.
What the rental market shows
Ornn publishes a settled daily rental index and term-price curves from one month to five years. The five-year A100 term price retains 80.2% of the one-month price, against 43.7% to 59.8% for Hopper and 53.8% for Blackwell, for a contract that would end when the Ampere family is 11.3 years old. A100 occupancy rose from 74% on 1 March 2026 to 90% on 1 September 2026 while listed capacity grew 13% and the spot index rose 20%, so the family's spot strength is not only a matter of fewer tracked GPUs being rentable.
What the paper does not claim
The paper states what its data cannot show. Forward marks are analyst-produced indicators, not executable quotes. The rental observations do not identify the contribution of open-weight demand, and the study does not establish that it caused A100 occupancy or rental prices; device retirements, dense inference, non-LLM work and financing effects remain alternative explanations. Ornn also discloses a commercial interest: it licenses its data and rents GPUs through Ornn Compute. The argument is narrower than the headline: older GPUs keep an economic role while suitable workloads stay competitive on them and operators remain free to deploy there.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.