
Kaggle AGI hackathon judging challenged over “gross inconsistencies”
A participant in Google DeepMind’s Measuring Progress Toward AGI hackathon on Kaggle has publicly challenged the awarding of the grand prizes, alleging thin use of the required benchmarking SDK, contradictory findings and undisclosed scores.
Winners of the Measuring Progress Toward AGI — Cognitive Abilities hackathon, a Google DeepMind featured competition hosted on Kaggle, were announced in mid-July. More than 1,000 teams had submitted benchmarks across five cognitive tracks. Within days, one participant published a detailed public objection to the way the prizes were awarded.
The complaint
Writing under the results announcement, the participant said they were presenting “evidence of gross inconsistencies regarding the evaluation process and selection of winners”. The post centres on MEDLEY-BENCH, a behavioural metacognition benchmark that took one of the four $25,000 grand prizes.
The critic argues that the entry used the Kaggle Benchmarks SDK — mandatory for the competition — so thinly that the model-comparison view exposes a single score with no insight into how the data was collected, and that the writeup contains “unsubstantiated claims”. Their main example: Finding #1 states that scale increases the “evaluation” measure while the “control” measure stays flat, even though both plotted lines rise together. The same team’s supplementary paper, the comment notes, reports that its base measures are highly correlated (ρ = 0.79–0.94), and a later insight contradicts the earlier core claim. The submission, the post adds, was padded with two AI-generated videos, an AI podcast, a website and a 20-page arXiv paper.
Requests for scoring transparency
Other participants raised narrower questions. One asked the organisers to release evaluation scores — ideally a full scored leaderboard, or at least the scores of the winning submissions — arguing that a hackathon built on measuring cognitive ability rather than asserting it should apply the same standard to its own judging. Another compared the Learning Track benchmarks by counting runs made with staff models, which he treated as a proxy for organiser attention, and asked which published criterion his own entry, ATLAS, had failed to meet.
Several commenters said the grand-prize entries are hard to verify: GAUGE reports roughly 200 items but shows a single run in the benchmark interface, and Metaproteus likewise displays one score. A separate objection concerned scope rather than scoring: the competition invited work on models that “reason, act and judge”, yet its operational structure was limited to five cognitive tracks in a largely prompt-and-response format, leaving no place for benchmarks grounded in physical interaction.
Why the argument matters
Four grand prizes of $25,000 each went to MEDLEY-BENCH, LearningBench, GAUGE and Metaproteus, while ten track prizes of $10,000 covered executive functions, learning, metacognition, social cognition and attention. Organisers thanked participants and said the quality of the submissions made judging “incredibly difficult”. The episode is a reminder that benchmark results, and the competitions that reward them, are only as useful as the evidence behind them.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.