
8 Models, 43 Matches: Why Agent Leaderboards Measure the Wrong Thing
A leaderboard of 43 matches and eight models puts claude-sonnet-5 first on Elo, grok-4.6 first on win rate and claude-fable-5.1 first on average placement. The disagreement says more about the metric than about the models.
Forty-three matches, eight models and 173 Elo points between first place and last: the entire scoreboard of TinyAIArena, where language models pilot fighters through turn-based combat and every match is replayable round by round, as Yano.AI reported on 28 September 2026.
The top rating belongs to claude-sonnet-5 at 1063, the bottom to deepseek-v4-flash-0731 at 890, after 25 matches without a single win. Read past the rating column, though, and the table stops agreeing with itself.
Three Winners from One Table
Win rate puts grok-4.6 first at 34%, ahead of claude-fable-5.1 at 32%. Average placement puts claude-fable-5.1 first at 1.88 finishes per match, against 2.38 for the Elo leader. Damage dealt puts grok-4.6 first with 3,676, more than the 2,620 of claude-sonnet-5. Three defensible definitions of the best agent, three answers from one small table.
43 Matches Is a Sample, Not a Ranking
Every model starts at 1000 Elo, so early movement reflects match count more than capability, and qwen3.8-max-0902 sits at 995 after exactly one match. The top three are separated by 33 points across 43 matches, while one fight ended in six rounds and another ran twenty, against an average of 10.3.
Who Won Is Not the Same Question as Why
Automated failure attribution in multi-agent systems reaches 33.3% step-level accuracy under a dynamic configuration and 30.3% under a static one, against 66.7% and 65.9% at agent level. An arena reports who won, the agent-level equivalent; production debugging needs to know which action lost the fight, and that figure sits near one in three.
Restricting analysis to output fields alone drops agent-level accuracy from 62% to 51% and step-level accuracy from 28% to 16%. The gap is on the input side: not whether reasoning text appears in a transcript, but whether the decision context of every call is recorded.
What Teams Should Log First
A trace schema that closes the gap records the rendered prompt, the injected context, the tool call and its arguments, and the serving configuration beside every step. A model at temperature 0.0 is a different system from one at 0.1 or with top_k=50, so rank two agents under different templates, sampling settings and engines and the table ranks harnesses rather than models.
Generic metrics answer the wrong question: ROUGE measures whether wording overlapped a reference, BERTScore whether two sentences meant the same thing, and helpfulness ratings often never verify that the task completed. Arena.ai scores tool reliability and task completion, but it still cannot explain a loss.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.