
Small models have arrived: why cheap inference changes the economics of AI products
Segment co-founder Calvin French-Owen argues that small, fast models are now good enough for real work. For investors, the main obstacle to more consumer AI companies is the cost of tokens.
Writing on August 26, Segment co-founder Calvin French-Owen argues that small, fast language models have quietly become good enough for real work. His main example is gpt-5.6-luna, which he has been using for weeks: it runs at around 100 tokens per second and moves comfortably across his codebase, email and knowledge base.
The cost barrier for consumer AI
French-Owen recalls a question he keeps hearing from investors: why are there not more consumer AI companies? His answer is token costs. The pre-AI playbook for consumer apps was to launch a cheap website, attract users, raise money to scale, and then build an ads marketplace — roughly the path taken by Google, Facebook and Snapchat. Adding AI changes that arithmetic, because every request carries real inference costs and the capital required rises sharply.
He illustrates the point with a personal benchmark: a prompt that asks an agent to research a person, work out what news they might like and build a personalized micro-site. With the previous generation of models in the Sonnet class, a single run cost about $1, which makes a $30-per-month subscription untenable. With luna, he says the results are decent and the average cost is about $0.10. He also notes that GLM 5.3 now sits on the Pareto frontier, giving buyers another option.
Cheap tokens and the shape of work
French-Owen extends the argument to business. On a recent hike, his Segment co-founder Peter described two kinds of work: the «IQ 180» kind, where someone finds a solution nobody had thought of, and «token spewer» work — staying responsive and pushing dozens of fronts forward at once. Peter has raised over $100 million for Charm Industrial and recently closed a Series A for Revoy, yet he estimates that about 95% of what he does falls into the second bucket. Most of the human tokens spent inside companies today look the same, and hiring skews toward that archetype.
His conclusion is not that frontier models matter less. He expects demand for top-end capability to keep compounding in fields that need genuine breakthroughs, such as engineering, hard science and model training. But he argues that demand for fast, cheap and good-enough models is about to take off, and that new harnesses, prompt-injection defenses, roles and permissions still need work before they are fully usable in business settings.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.