
Epoch AI: Cost of a Given Level of AI Performance Falls 47% a Quarter
An Epoch AI analysis finds that the cost of reaching a given level of AI performance dropped about 47% per quarter over three years — a 13-fold fall every year, faster than any other transformative technology.
The cost of reaching a given level of artificial intelligence performance fell by an average of about 47% per quarter over the past three years, according to a new Epoch AI report by Emberson and Roodman. That works out to a 13-fold decline every year — a faster rate than any other transformative technology in history. The economics blog Marginal Revolution highlighted the findings.
What the study measures
The authors analyzed five benchmarks covering mathematics, hard sciences and games of skill, and traced the "cost frontier": the cheapest model available at any moment that can still hit a given score. On average, that cost falls 47% per quarter, or 13 times a year. The pace depends on the task — game-based puzzles get cheaper more slowly, at 39–43% per quarter, while mathematics moves faster, at 50–52% per quarter.
The rate is not constant. When a new capability first appears it is, briefly, the state of the art, and that is when prices fall fastest: across the five benchmarks, cost drops 66% per quarter (75-fold a year) at that stage, then slows by half — to 32% per quarter, or 4.7-fold a year — two years later.
From o3 to GPT-5.6 Luna: 725 times cheaper
The report's clearest example is the GPQA Diamond test, a set of PhD-level questions in physics, chemistry and biology. Epoch AI estimates that in January 2025 OpenAI's o3 model reached a 75% score at an average cost of about 30 cents per question. Less than 18 months later OpenAI released GPT-5.6 Luna, which gets the same score for $0.0004 per question — roughly 725 times cheaper.
Why it matters
The authors note that models are not only getting smarter: a fixed level of intelligence increasingly requires far less inference spending to run. Marginal Revolution takes a competitive reading of the numbers — the threat from open-weight models is smaller than it appears, because frontier models do not merely outperform older ones, they also get rapidly cheaper to run at any given level of performance.
Epoch AI's authors caution that the findings come with caveats. Benchmarks are an imperfect proxy for useful work, and firms may train deliberately for known tests. The figures also assume a user who always picks the cheapest model that meets a target, which real users rarely do. Nor does a falling price mean falling spending: if a job demands as much cognition as a million benchmark runs, the bill adds up even at a penny per run.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.