Back
OpenRouter data shows AI agents widening lead in token use
SiTech AI Team3 min read

OpenRouter data shows AI agents widening lead in token use

OpenRouter data cited by a16z shows agents using 7.3 trillion tokens versus 1.4 trillion for humans in August, while cached prompts and memory demand complicate cost estimates.

Agents dominate OpenRouter token volume

AI agents generated 7.3 trillion tokens on OpenRouter in August, compared with 1.4 trillion for human users, according to an Andreessen Horowitz chart based on OpenRouter data. Futurum Group CEO Daniel Newman cited the figures six months after agent usage first surpassed human usage.

A chart from OpenRouter head of insights Peter Walker showed that, since the February crossover point, agent token usage increased 14 times, while human usage rose 2.8 times. The mixed category grew 4.7 times over the same period, according to the source's analysis. Because this category may include both agent and human behavior, the size of the agent lead may vary.

OpenRouter assigns each API key to an agentic, mixed or human category using a seven-signal weighted composite score. The classification considers factors including tool call rate, turn count and gap timing.

A strong trend, but not a complete market measure

The OpenRouter figures measure token volume rather than spending and cover only one platform. The trend was not completely consistent, with dips in April and July, while the presence of mixed traffic makes direct comparisons more difficult.

McKinsey's 2026 State of AI survey found that 40% of respondents from large organizations reported scaling AI agents, up from 27% a year earlier. Newman wrote that AI was currently used by AI five times more than by humans and predicted the multiple would reach 10 times and continue higher. That projection remains a forecast, while the OpenRouter data covers activity on a single service.

Cached prompts contribute to memory pressure

More than 85% of agent tokens came from cached prompts, a16z said, citing OpenRouter. The firm also reported that cached tokens accounted for nearly all relative growth in token usage. Cached prompts cost less than processing a prompt from scratch, but the stored information still has to remain in memory.

Models retain context in a KV cache, which is outgrowing GPU high-bandwidth memory capacity. a16z, an OpenRouter investor, linked this constraint to rising demand for HBM. Logs from a call center consultancy that tested DeepSeek on rented Nvidia H200s showed that 96% of all input during its own agents' September usage on Claude Code involved rereading old conversations. Combined with the OpenRouter figures, this suggests token counts can overstate bills, although the hardware cost of memory remains significant.

Micron expects RAM and storage shortages to worsen in 2027 and 2028, with customers paying more than they do this year. Memory manufacturers are prioritizing HBM for AI data centers, keeping the capacity needed for agent context in direct competition with other customers.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.