Back
Bigger models are not the answer: GPT-5.5 hallucinates on 86% of unknowns, GLM-5.2 on 28%
SiTech AI Team3 წთ. საკითხავი

Bigger models are not the answer: GPT-5.5 hallucinates on 86% of unknowns, GLM-5.2 on 28%

An analysis published on arrowtsx.dev argues the parameter race has plateaued: the MIT-licensed open model GLM-5.2 nearly matches GPT-5.5 on benchmarks while proving far more truthful.

A growing number of major AI labs are questioning the assumption that endlessly growing parameter counts and training data automatically produce better models. In an analysis published on arrowtsx.dev on June 18, the author argues that scaling has plateaued — and that sheer size now correlates with a different problem: models that will not admit uncertainty.

Benchmarks and the parameter race

The post opens with an unusual example. Claude Fable 5 was restricted by the US government three days after its release — described as the first US AI ban driven by national security, triggered by the risk posed by a single jailbreak. Meanwhile Z.ai's GLM-5.2, an open-weight model with 753 billion parameters and roughly 40 billion active, lands within 4 points of GPT-5.5 and 9 points of Fable 5 on the Artificial Analysis Intelligence Index. Opus 4.8 and GPT-5.5 are proprietary and conservatively estimated at 1–2 trillion parameters. When an MIT-licensed open-weight model comes that close to a closed model one and a half to two times larger, the author concludes, measurable intelligence has largely plateaued.

Hallucination rates

The second part of the analysis turns to truthfulness. On the AA-Omniscience benchmark, DeepSeek V4 Pro (1.6 trillion parameters, 49 billion active) records a 94% hallucination rate: asked questions it could not answer, it admitted not knowing only about 6% of the time. GLM-5.2 scored 28%, Opus 4.8 scored 36%, Fable 5 scored 48%, and GPT-5.5 scored 86%. Models trained on very large volumes of factual data, the author argues, learn to always produce an answer instead of saying "I don't know".

A test, and an unsolved trilemma

To illustrate, both models were given a complex Python question containing a clear architectural flaw, with high reasoning effort and temperature 1, served through OpenRouter. DeepSeek V4 Pro spent almost ten times as many reasoning tokens and still returned a confidently incorrect response. GLM-5.2 needed about 12 seconds and 800 reasoning tokens to recognise the impossibility of a single-threaded task handling multiplexed I/O without ever yielding. In a separate run, DeepSeek V4 Pro burned 3 minutes and 26 seconds in a reasoning loop before delivering a well-formatted but wrong solution.

The conclusion is that training and model selection should be designed around an unsolved trilemma: raw capability, uncertainty calibration, and computational efficiency. Choosing models by size or leaderboard position alone, the post warns, increasingly means choosing a system that will argue convincingly for an answer that is simply wrong.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.