Back
Tokens Too Cheap to Meter: How AI Inference Costs Keep Falling
SiTech AI Team2 წთ. საკითხავი

Tokens Too Cheap to Meter: How AI Inference Costs Keep Falling

An essay on jyn.dev argues that machine intelligence is getting cheaper by orders of magnitude every year, and that tokens will stop being the scarce resource long before quality does.

An essay on jyn.dev argues that the price of machine-learning intelligence is falling by several orders of magnitude a year, with no sign of slowing, and that quality and access — not token supply — will become the binding constraint.

The evidence

Epoch AI data cited in the post shows GPU power efficiency rising exponentially: on a logarithmic plot the trend slope is about 1.3, meaning efficiency doubles roughly every two years — a pace unseen since Moore's Law.

GPU power efficiency over time, from Epoch AI

Per-token prices for frontier models are not falling consistently, but cost per task is. On the pareto frontier tracked by Artificial Analysis, the intelligence axis in 2026 looks much like 2025, while the cost axis became two orders of magnitude cheaper over 2025.

Cost per task on the 2026 model pareto frontier

Inference engines improve 10–50% a year: vLLM 0.11.1 (December 2025) spends about 40% fewer joules per token than 0.5.4 (September 2024), and NVIDIA reports up to 50% efficiency gains from MLPerf 2.0 to 2.1. Architectures are shifting too — Mixture-of-Experts deactivates unneeded expert layers, letting a model be 7x smaller (6B to 0.8B parameters) at equal quality.

Where “too cheap to meter” comes from

TypeSafe AI's Jev emits no text; it picks among pre-chosen options and returns a probability — whether a shell command violates a system prompt, for instance. Its pricing page puts existing LLMs at $0.20–$10 per million input tokens, with output about five times that, against System One + Jev at $0.042 per million — $42 per billion — with output tokens free: roughly three cents to read five books.

Tooling already exploits it: jgrep returns a probability in about 200 milliseconds for roughly a thousandth of a cent, and the open-weight Laya runs locally.

What it changes

The essay also compares a model turn with local tools: GPT-5.6 Luna costs about 30 cents per million tokens, so a 10,000-token tool decision runs a third of a cent, while grep is roughly 4.5 orders of magnitude cheaper. The author does not expect a bubble burst: cheap volume does not make quality cheap, and OpenAI keeps building — five new Stargate sites.

He expects security to get harder, more renting of inference-optimised GPUs, and a shift in what makes software hard: product requirements, testing and interface design rather than algorithms. Users gain a fourth option beside use it, don't use it, or use a rival — telling an LLM to build it.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.