
Simon Willison: Claude Fable 5 is relentlessly proactive
Simon Willison calls Claude Fable 5 relentlessly proactive: the model invented its own browser-automation tricks to trace a two-line CSS bug, in a session that would have cost about $12.11 at full API prices.
A one-line prompt and a screenshot
Simon Willison says that after two days with Claude Fable 5, the best description he can find for the model is “relentlessly proactive”: it knows a large collection of tricks and will deploy almost any of them to reach its goal.
He illustrates it with a debugging session on Datasette Agent. While working on the project he spotted a glitch — a horizontal scrollbar that should not have been there in the jump menu chat prompt. He took a screenshot, started a fresh Claude session in his datasette-agent checkout, dragged the image in and asked the agent to look at dependencies to find out why the scrollbar was present. Then he walked away from his computer.
Minutes later he saw the machine open a browser window in Firefox and navigate to the dialog in question, even though he had not enabled any browser automation.
Tricks nobody had written down
Fable had built its own screenshot mechanism. Through uv run --with pyobjc-framework-Quartz it ran Python that enumerated the windows on the machine, filtered for Safari windows whose names contained strings such as textarea, and read the window number — an integer like 153551 — to capture a PNG with the macOS screencapture command. It also wrote its own scratch HTML pages, including /tmp/textarea-scrollbar-test.html, to reproduce the bug.
The dialog under test could only be opened with a click or a keyboard shortcut, so Fable edited Datasette's own templates and injected JavaScript that dispatches a simulated “/” key press 1.2 seconds after the window loads. To take measurements it wrote a small Python server on top of http.server that accepted JSON POST requests, wrote them to /tmp/diag.json and returned CORS headers, then injected a script that reported scrollWidth, clientWidth and devicePixelRatio values from the textarea inside the navigation-search web component.
Fable eventually hit an invisible guardrail and downgraded itself to Opus. Opus had the full transcript, reused the same techniques and verified the fix — a two-line CSS change.
What the session cost
Willison is on the $100/month Claude Max plan, which includes a Fable allowance until June 22nd, after which Anthropic has said it will charge full API prices. AgentsView reported 68,606 output tokens and a peak context of 113,178 tokens for the session, worth about $12.11 at full price.
Why it matters
The bigger point is security. Coding agents can do anything a user can do by typing commands into a terminal, and frontier models know every trick in the book, plus a few that nobody has ever written down. Had Fable been following malicious instructions — a prompt injection hidden in code or an issue thread — it is alarming to think how far it could go to exfiltrate data.
Willison repeats that running coding agents outside a sandbox has always been a bad idea, and cites Johann Rehberger's essay on the normalization of deviance in AI. A smarter model may be more suspicious of malicious instructions, he notes, but that cuts both ways: if it is subverted, its relentless proactivity makes the damage far worse.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.