
DeepSeek-V4-Flash leaves preview with stronger agent benchmarks and Responses API support
DeepSeek has moved the DeepSeek-V4-Flash API into public beta with the same architecture as the preview and only new post-training, plus native support for the OpenAI Responses API format and Codex.
V4-Flash moves out of preview
DeepSeek has released the official DeepSeek-V4-Flash API in public beta. The calling convention is unchanged: developers keep the same endpoint and set the model name to deepseek-v4-flash. According to the company's changelog, the released checkpoint — DeepSeek-V4-Flash-0731 — keeps the same architecture and size as V4-Flash-Preview and was only re-post-trained, so the gains come from training rather than from a new design.
DeepSeek opened the V4 family earlier this year with two tiers, V4-Pro and V4-Flash, exposed through both the OpenAI ChatCompletions interface and the Anthropic interface. This latest update touches the Flash API only: the changelog says the V4-Pro API and the app and web models are unchanged, and that the official release of V4-Pro will follow soon.
Agent and coding benchmarks
DeepSeek reports results across a set of agentic and coding evaluations: Terminal Bench 2.1 at 82.7, NL2Repo 54.2, CyberGym 76.7, DeepSWE 54.4, Toolathlon Verified 70.3, Agents' Last Exam 25.2 and Automation Bench (Public) 25.1. On its internal test sets it reports DSBench-FullStack 68.7 and DSBench-Hard 59.6.
The company notes that Code Agent tasks in public benchmark sets were run with the DeepSeek Harness minimal mode — a framework it says will be released soon — at the maximum effort level, with top_p 0.95 and temperature 1.0. DSBench-FullStack is an internal full-stack development set and DSBench-Hard an internal hard-problem coding set, so both are self-reported numbers.
Responses API and Codex support
The API now natively supports the OpenAI Responses API format and is specifically adapted for Codex, with a one-click configuration script in the documentation. DeepSeek says the results far exceed V4-Pro-Preview on agentic benchmarks — a claim that, like the table above, comes from the vendor's own harness.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.