DeepSeek has pushed a post-trained V4-Flash update into public beta with native support for the Responses API, making the model available as a custom provider across Codex CLI, the ChatGPT desktop app and the Codex extension for VS Code.
Launch Tracker: Read the verified DeepSeek V4-Flash-0731 launch record for the source checklist, pricing snapshot and availability details.
The release, identified on DeepSeek’s pricing page as DeepSeek-V4-Flash-0731, keeps the architecture and parameter count of April’s V4-Flash preview. The change is in post-training. DeepSeek says that work produced much stronger agent and coding results, led by an 82.7 score on Terminal-Bench 2.1.
Cline quickly drew attention to the jump: the April preview scored 56.9 on the same benchmark, making the new company-reported result a 25.8-point increase. The more practical news for developers is that the updated model can now speak the protocol Codex uses.
What DeepSeek released
DeepSeek announced the official V4-Flash API as a public beta on July 31, 2026. The company’s changelog says the API call remains the same: developers select deepseek-v4-flash to reach the latest version.
- Release: DeepSeek-V4-Flash-0731 public API beta
- What changed: post-training, with the architecture and model size unchanged
- Availability: DeepSeek API and MIT-licensed weights on Hugging Face
- Codex support: native Responses API compatibility
- Context: one million tokens, with a listed maximum output of 384,000 tokens
- Current pricing: $0.14 per million uncached input tokens and $0.28 per million output tokens
DeepSeek says the V4-Pro API and its app and web models are unchanged. The 0731 checkpoint itself, however, is no longer limited to DeepSeek’s hosted API. Kimmonismus highlighted the weight release, and DeepSeek’s official DeepSeek-V4-Flash-0731 repository confirms a public, ungated MIT-licensed release with 48 Safetensors weight shards. The model card calls 0731 the official release that supersedes the preview. It retains a 284-billion-parameter mixture-of-experts architecture with 13 billion parameters active per token.
The benchmark jump is broad, but still a vendor claim
DeepSeek’s release table compares V4-Flash-0731 with both the earlier V4-Flash preview and V4-Pro preview. The updated Flash model leads both previews across every agent benchmark shown.
Selected company-reported results (V4-Flash-0731 vs V4-Flash Preview):
- Terminal-Bench 2.1: 82.7 vs 56.9, up 25.8 points
- NL2Repo: 54.2 vs 39.4, up 14.8 points
- Cybergym: 76.7 vs 38.7, up 38.0 points
- DeepSWE: 54.4 vs 7.3, up 47.1 points
- Toolathlon-Verified: 70.3 vs 49.7, up 20.6 points
- Agents’ Last Exam: 25.2 vs 15.8, up 9.4 points
Those gains are unusually large for a release that leaves the base architecture untouched. They also need context. DeepSeek ran the public code-agent evaluations with its forthcoming “DeepSeek Harness” in minimal mode, using maximum reasoning effort, top_p=0.95 and temperature=1.0. The harness has not yet been released, so outside researchers cannot reproduce the complete setup today.
Two other scores in the announcement, DSBench-FullStack at 68.7 and DSBench-Hard at 59.6, come from DeepSeek’s internal test sets. They are useful as company-reported indicators, not independent evidence. Kingy.ai has not rerun these benchmarks or tested V4-Flash-0731 in a production repository.
Why the Codex support matters
Codex clients use OpenAI’s Responses API rather than the older Chat Completions format. V4-Flash-0731 now supports that interface natively, so developers can add DeepSeek as a model provider instead of routing requests through a compatibility layer.
DeepSeek’s Codex integration guide provides an automated setup script and a manual configuration path. Both create a custom model catalogue and add a DeepSeek provider to Codex’s shared ~/.codex/config.toml file. Because Codex CLI, the desktop app and the VS Code extension read that same file, one configuration makes the model available across all three clients.
Only V4-Flash works with Codex at launch. DeepSeek says V4-Pro support is expected in early August 2026. In the ChatGPT desktop app, locally configured models appear under the generic “Custom” label, so users should verify the selected provider and model in their configuration rather than relying on the picker label alone.
The price makes the result hard to ignore
DeepSeek currently lists V4-Flash at $0.0028 per million cached input tokens, $0.14 per million uncached input tokens and $0.28 per million output tokens. The company also lists a one-million-token context window and support for thinking and non-thinking modes.
That combination makes V4-Flash-0731 an obvious candidate for cost-sensitive agent workloads: long repository context, native tool-oriented API support and a low token price. It does not prove the model will outperform a more expensive competitor on a particular codebase. Terminal agents are sensitive to scaffolding, prompts, tool permissions, latency limits and the exact task mix. A team choosing a coding model should rerun its own repository-level evals before moving production work.
DeepSeek also says it plans to introduce peak-hour pricing at twice the regular rate, although it has not announced the effective date. The current pricing page should be treated as a live reference rather than a permanent tariff.
What to watch next
The first question is reproducibility. Releasing the DeepSeek Harness would let independent evaluators test how much of the gain comes from the model and how much comes from the agent scaffold. The second is practical self-hosting. The 0731 weights are available under MIT, but Kingy.ai has not tested a local deployment or verified parity with DeepSeek’s hosted API. A 284-billion-parameter mixture-of-experts model remains a substantial hardware and serving project.
The third is V4-Pro. DeepSeek says Pro’s Codex support is due in early August and that a broader Pro release will follow. For now, V4-Flash-0731 is the only DeepSeek model officially wired into Codex, and the only new checkpoint covered by this announcement.
For developers, the sensible next step is a controlled trial: configure V4-Flash through DeepSeek’s official guide, run a representative set of repository tasks, and compare completion quality, tool reliability, latency and total token cost against the model already in use.
Kingy Launch Brief
Put the week’s verified AI launches in your inbox.
One source-checked edition every Friday, with a clear try, watch or skip verdict. After subscribing, check your inbox and confirm your address.
Free · Fridays · Double opt-in · Unsubscribe anytime
