← Back
SiTech Team⏱️ 6 წთ. საკითხავი

AI Chatbots Reading X-Rays Can Be Dangerously Confident Even When Wrong

AI Chatbots Reading X-Rays Can Be Dangerously Confident Even When Wrong

The RadLE 2.0 benchmark reveals AI models reading X-rays are dangerously overconfident when wrong — even Claude Fable 5 scored only 758/2000 vs human radiologists' 988.7.

What Is RadLE 2.0? Radiology's Last Exam for AI

In July 2026, the CRASH Lab at Ashoka University, India, released RadLE 2.0 — the second iteration of Radiology's Last Exam. This benchmark goes beyond conventional accuracy metrics. It tests whether AI systems can recognize when they should leave a diagnosis to a human professional. The fundamental question is not just whether a model gets the answer right, but whether it knows its own limits.

The evaluation ran 200 clinical cases across 16 different AI models and compared their performance against a panel of human radiologists. Out of a maximum of 2,000 points, human experts scored 988.7. The best-performing AI model managed 758.

This striking gap demonstrates that despite rapid progress, AI remains far from ready for independent radiological diagnosis. More concerning still: many models cannot tell when they are making a mistake.

The Scoring System — Why Honest Silence Beats Confident Guesswork

RadLE 2.0's scoring methodology marks a radical departure from traditional benchmarks. It rewards honesty and punishes overconfidence. Models must self-assess their answers on a confidence scale from 0 to 4. A correct answer with high confidence earns full points. A wrong answer delivered with high confidence deducts an equal number. Answering "I don't know" scores zero but carries no penalty — making it the safest option when uncertain.

This design addresses a critical flaw highlighted by a highly cited paper: as long as benchmarks only reward accuracy, AI models are effectively trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent. This insight is precisely what makes some models dangerous for real-world patient care.

No Single Winner — Claude Fable 5, Gemini 3 Pro, and Muse Spark 1.1

No single model swept every category. Anthropic's Claude Fable 5 performed best on reliable and safe answers, leading the primary composite metric. Google's Gemini 3 Pro achieved the highest raw accuracy among frontier models, nearly catching up to human performance on hit rate alone.

Meta's Muse Spark 1.1 distinguished itself by excelling at knowing when to hand a case off to a human radiologist. Meta had recently cut the model's hallucination rate nearly in half — it now more often refuses to answer rather than producing a wrong one. This is a rare and valuable trait in the current AI landscape.

Other frontier models trend in the opposite direction. Grok 4.5, for instance, hallucinates significantly more than its predecessor. The paradox is that it knows more, but it is also more convinced of its wrong answers — a dangerous combination in a medical context.

The Open-Weight Problem — Trying to Answer Everything

The study uncovered a particularly troubling pattern among open-weight models and those specifically fine-tuned for medical use. These models attempt to answer nearly every case, frequently with high confidence, and are wrong most of the time.

"Several models would have scored much better if they had stayed quiet more often instead of guessing," the researchers noted. This behavior is especially alarming given that patients are already uploading X-ray and MRI scans to chatbots and trusting their responses. The gap between open-weight models and human performance widens sharply, yet their overconfidence remains unchecked.

Real-World Risks — Patients Trusting Chatbots Today

More and more people are turning to AI chatbots for medical advice. A study published in npj Digital Medicine demonstrated that widely used chatbots frequently provide unreliable answers to medical questions. The research team criticizes executives and investors for publicly overstating AI capabilities. Claims that AI systems already diagnose better than 99 percent of doctors are mostly based on anecdotes or simulations, not rigorous clinical validation.

A 2025 Polish observational study revealed another dimension of risk — the "Google Maps effect." Doctors who regularly use AI during colonoscopies detect significantly fewer precancerous lesions when working without the tool. Detection rates dropped from 28.4 percent to 22.4 percent. Without the navigation aid, users are lost, and their skills atrophy from over-reliance on automation.

As recently as April 2026, a study of 21 state-of-the-art models showed they are not yet ready for unsupervised clinical use. The gap between laboratory performance and real-world reliability remains dangerously wide.

Radiology's AI Hype Cycle — From 2016 to 2026

Radiology has already been through one complete AI hype cycle. In 2016, AI pioneer Geoffrey Hinton famously declared that we should stop training radiologists because deep learning would soon take over the job. Richard Sutton and other prominent voices agreed. Nearly ten years later, radiologists remain overburdened, and Hinton has walked back his prediction. He reduced the profession to image analysis and overlooked the full complexity of the field.

OpenAI CEO Sam Altman spent years predicting that AI would replace human jobs at a terrifying pace, then recently reversed course, suggesting AI may have actually created more jobs. The common thread is clear: AI specialists understand their models but consistently overestimate how quickly entire professions can be replaced.

Conclusion — AI Must First Learn to Stay Quiet

The RadLE 2.0 authors are unequivocal: before AI makes independent decisions in medicine, it must know when it is better off not doing so. Recent studies on autonomous medical AI agents — MIRA for electronic health records and AMIE for simulated consultations — show that AI can keep pace with general practitioners in controlled settings. But the real world is far more complex than any simulation.

AI specialists may understand their models intimately, but they routinely overestimate the speed at which entire professions can be replaced. Much like the AI they build, even people do not always know when they would be better off staying quiet because they have stepped outside their own expertise.

RadLE 2.0 will be expanded on a rolling basis to include new models. A full scientific publication with cost analyses and an error taxonomy has been announced. The benchmark's first version, released in September 2025, painted an even starker picture: radiologists hit 83 percent accuracy while the best AI managed only about 30 percent. Within three months, Gemini 3 Pro surpassed the level of resident radiologists on raw accuracy — but the models still lack any sense of their own limits. And that, ultimately, is what makes them dangerous.

📖 Source