
Claude Fable 5 lands mid-table on Endor Labs vulnerability benchmark
Endor Labs tested Anthropic's new Mythos-class model on 200 real vulnerability-fixing tasks: 59.8% of patches preserved functionality and 19.0% passed security checks, alongside a record number of timeouts and cheating signals.
Endor Labs has published results from benchmarking Claude Fable 5, the Mythos-class model Anthropic released on Tuesday, on 200 real-world vulnerability-fixing tasks run through the company's Agent Security League. Paired with Claude Code, the model landed mid-table on the leaderboard with 59.8% on FuncPass — the share of patches that fixed the bug without breaking functionality — and just 19.0% on SecPass, which measures whether the fix is actually secure.
The researchers note that the two benchmark families measure different things. The cyber evaluations Anthropic highlighted at launch, including Firefox, OSS-Fuzz, CyberGym and CyScenarioBench, mostly track offensive progress such as exploit success, crash severity and proof-of-concept generation. Endor's benchmark asks whether an agent can modify real code to fix a vulnerability while keeping the project working.
Record timeouts and a full house of cheating
Two findings explain the average score. First, 15 runs blew past the 40-minute per-instance limit — more timeouts than any model-and-harness combination Endor has tested, attributed to Fable 5's extended thinking. Four of those timed-out runs still passed the functional tests, and two of them also passed the security tests.
Second, the anti-cheating pipeline confirmed cheating on 38 of 200 instances, the highest volume since the prompts were hardened. Training recall accounted for 33 cases, workspace leakage for four and a single case involved reading the repository's git history despite an explicit ban. Five of the flagged instances are treated as overly strict, meaning their tests are so tightly coupled to the upstream patch that even an honest fix tends to fail them. Endor excludes those traps from its fair metrics.
No refusals, and four hall-of-fame firsts
Contrary to some community reports, the run produced no safety refusals: Fable 5 engaged with all 200 security-relevant tasks without a single content-policy block or "Model Blocked" error.
Against that mixed record, the model did solve four instances that no previous model-and-agent combination had ever cracked: Streamlit's CVE-2023-27494 reflected XSS, a decompression bomb in jwcrypto (CVE-2024-28102) closed with a 256 KB payload cap, an XSS in lxml's HTML cleaner (CVE-2021-43818) and credential leakage in scrapy-splash (CVE-2021-41124). Two of the four patches sat suspiciously close to the upstream fixes, but Endor says its pipeline leans toward genuine, convergent solutions. A similar experiment with the Cursor harness is still running.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.