Why I'm still bearish on LLMs after Navier-Stokes
A widely discussed blog post argues that headline AI results do not amount to real autonomy: models still need heavy supervision, and the cost of the rigorous specifications that would fix it is a structural barrier for most firms.
The case against the hype
A post published on 15 September on the blog dank.systems argues that the excitement around large language models has outrun what the systems can actually deliver. The author writes that frontier labs are priced on the narrative that they will very soon produce "a fully automated drop-in replacement for most knowledge workers" — while current models still need "laborious oversight and guardrails on even the simplest tasks".
Headline demonstrations such as the Navier–Stokes work, FreeBSD remote code execution findings and the Hugging Face incident do not show that meaningful autonomy has been reached, he argues, pointing to software firms that keep employing engineers who would score below the models supervising them on today's benchmarks.
Narrow generalisation and the cost of specifications
The second claim concerns generalisation: models perform well only in a small neighbourhood of the tasks they were trained on, and even there "small perturbations within a covered class of task result in outright failure or reward hacking". Fixing that requires rigorous specification by domain experts, whose time is expensive — and specification is a separate skill of its own, so the overlap between people who know a domain and people who can specify it precisely is, he argues, very small.
The labour cost can exceed that of simply implementing an informal description. In hardware engineering, a typical CPU project has roughly three times as many specification and validation engineers as design engineers, and a 5:1 ratio is not unheard of, the post notes.
Mathematics as the best case, and human review as the fallback
Mathematical results are, in the author's words, the "absolute best case scenario" for work against a rigorous specification: the theorem statement is already a specification, audited for decades, and its formalisation in Lean is a translation of well-tested objects from mathlib. Even so, soundness bugs have previously let LLMs push bogus proofs through Lean's kernel.
The alternative, human review, does not scale to the volume of model output and is itself vulnerable to gaming; the author cites the xz backdoor and the "hypocrite commits" incident in the Linux kernel. If human review stays in the loop, the pace of production is capped by human attention.
Which firms can run autonomous AI
The post concludes that for most domains LLMs resemble "a cracked intern: quick and effective in the hands of an adult but not given run of the place". Only three classes of firms can accept fully autonomous use: those that absorb failure cheaply, such as rapid-prototyping teams; those with a narrow set of well-guarded tasks, such as controlled repetitive work or customer-service chat; and those that already pay for rigorous specification and validation, such as chip design and drug discovery.
The first two groups are price-sensitive and may not need frontier reasoning at all, the author argues, suggesting open models on cheap hardware. His conclusion is that the consequences will reach well beyond the frontier labs.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.