
TryAI pits Claude Fable 5 against GPT-5.6 Sol in a $100 music video arena
An open-source harness gave two frontier models the same song, a hard budget and ffmpeg, then left them to produce a full music video on their own. All four runs finished unaided.
TryAI has published a build-off in which two frontier models, Claude Fable 5 and GPT-5.6 Sol, each received the same song, a hard dollar budget and a set of tools, and were left to produce a complete music video without human intervention. The aim was to observe how tool use differs between models on a long-horizon, open-ended task.
How the arena works
Each model ran an autonomous tool-calling loop with six tools: plan, web_search, get_budget, generate_image, generate_video and run_command. Only the two generation tools spend money, and the model picks its own FAL or Replicate model and parameters. The shell supplies ffmpeg and ffprobe for analysing audio, cutting clips and muxing the final cut; when the budget hits zero, paid generation is refused but editing continues. Every message, tool call, charge and error was logged, and the harness is open source on GitHub.
All four runs — two models at $25 and at $100 — received identical inputs: Bruno Mars and Mark Ronson's “Uptown Funk”, a short text description and a time-stamped lyric transcript.
What the runs cost
All four finished without hitting a step or time limit and produced a valid full-length video with the original song. Wall-clock time ran from 38m56s for Claude Fable 5 at $100 to 49m39s for GPT-5.6 Sol at the same budget. Metered generation spend was $24.30 and $48.60 for the Claude runs and $23.18 and $36.57 for Sol; adding model tokens at $10/$50 per million input/output tokens for Fable 5 and $5/$30 for Sol, the totals reach $41.29, $27.45, $39.82 and $73.65. Runs generated between 46 and 80 distinct clips.
Given a free choice, the models diverged. Three of the four runs went pure text-to-video; only GPT-5.6 Sol at $25 built an image-to-video pipeline, generating stills with FLUX schnell before animating them. At $100, Sol mixed Wan 2.5, Veo 3.1 Lite and Hailuo 2.3 Standard in a single run, while Fable 5 used Seedance 1.0 Pro and delivered the only 1920x1080 output. Neither model touched Replicate, although both keys were available.
What still goes wrong
TryAI is blunt about the results: none of the videos were great. Character and story consistency struggled in all four runs, with recurring characters drifting between shots and no run holding a coherent storyline. The models read lyrics literally — the line “make a dragon wanna retire” produces an actual dragon on screen — and tempo matching stayed weak, with cuts landing on the beat but the motion inside each clip rarely matching the song.
Editing was mostly a single pass: the models concatenated and muxed their clips but rarely re-cut or added effects, and none seriously reviewed its own footage. TryAI adds that $100 was probably too much budget, since neither model wanted to approach the cap.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.