Back
GPT-6 Astra, looped transformers and the debate over hidden reasoning
SiTech AI Team3 წთ. საკითხავი

GPT-6 Astra, looped transformers and the debate over hidden reasoning

OpenAI's new GPT-6 Astra tops benchmarks and computer-use demos, while reports that it relies on "looped transformers" have fuelled a debate over whether the model hides its chain of thought.

OpenAI released GPT-6 Astra in early September, and the machine-learning researcher Sebastian Raschka, who tested the model within days, calls it the strongest he has used.

In his review, Astra beats its GPT-5.6 predecessor almost everywhere — writing, math, coding — but the widest gap is in graphical demos and 3D rendering. On ARC-AGI-3, a benchmark mixing logic puzzles with generalisation, the model scores 99.9 percent, against 7.8 percent for GPT-5.6 Sol.

A frontier lead, with caveats

On the independent Artificial Analysis indices the picture is more measured: Astra sits at the frontier of the Coding Agent Index v1.4 and the Intelligence Index v4.2, but does not pull away by leaps and bounds. Because those evaluations mix harnesses — Stirrup for GDPval-AA and AA-Briefcase, Terminus 2 for Terminal-Bench v2.1 — a model tuned around a single primary harness may be underestimated.

Computer use as the new front

The most striking gains are in computer use, where the model operates software on a local machine through the ChatGPT and Codex apps. Demos circulating online include rendering New York City in Blender and painting with a mouse in a browser version of MS Paint. OpenAI reportedly bought tens of thousands of Mac Minis and Mac Studios to expose macOS to the model during reinforcement learning: screenshots go in, mouse and keyboard actions come out, and a harness executes them and returns fresh screenshots. Training itself runs elsewhere — Nvidia's chief executive said Astra was trained on about 100,000 Grace Blackwell GPUs. Astra remains a reasoning model, trained with reinforcement learning with verifiable rewards.

Looped transformers and hidden chains of thought

Two days before launch, The Information reported that Astra uses "recurrent depth," or a looped transformer. In such designs the same blocks are applied several times with shared weights: Nanbeige4.2-3B runs 22 blocks twice, ByteDance's Ouro repeats 48 blocks four times, and Mixture-of-Recursions routes each token through one to three passes. The trick raises effective depth while storing fewer weights, but compute and KV-cache costs stay close to a conventional model of the same depth; SMELT estimates 6.8 to 18 percent less training compute for the same validation loss.

Raschka doubts that looping hides reasoning. OpenAI has withheld most reasoning traces since o1, and Astra uses fewer tokens than GPT-5.6 Sol at equal accuracy — a pattern also seen between smaller models (Luna uses 80 percent more tokens than Sol for similar results) without alarms about interpretability. Astra's system card does note reduced monitorability relative to Sol, tied mostly to shorter traces. OpenAI's chief scientist, Jakub Pachocki, said the computation graph of frontier models is within a factor of two of GPT-4 and warned against a "race into unmonitorability kicked off by confused reporting."

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.