Here is the honest state of the DeepSeek V4 Flash benchmark picture on launch day. As of 2026-07-31, almost every number you will see for DeepSeek-V4-Flash-0731 is vendor-reported: DeepSeek's own table lists Terminal Bench 2.1 at 82.7, NL2Repo at 54.2, Cybergym at 76.7, and DeepSWE at 54.4, produced with DeepSeek's not-yet-released Harness in minimal mode at maximum reasoning effort (changelog, 0731 model card). The one same-day independent signal is Artificial Analysis, which scored the 0731 release 50 on its Intelligence Index, up from 40 for the previous Flash (Artificial Analysis) - with launch-day caveats of its own. This page, written by independent DeepSeek fans not affiliated with DeepSeek, walks through the scores, the settings behind them, and what still needs verification.
The Short Benchmark Verdict
What the current scores support
The current evidence supports three specific statements. First, DeepSeek reports large agent-benchmark gains for the 0731 release over the April Preview in its own evaluation: DeepSWE moved from single digits to 54.4 and Cybergym roughly doubled in the vendor table (changelog). Second, one independent evaluator confirms a real jump: Artificial Analysis reran its own nine-benchmark harness on 2026-07-31 and measured Intelligence Index 50 versus 40 for the prior Flash, with the agentic GDPval-AA v2 Elo rising from 1189 to 1559 (source). Third, DeepSeek's table shows 0731 outperforming the V4 Pro Preview on the listed agent benchmarks - a scoped claim that holds only within DeepSeek's own harness and settings. Together these justify "worth testing on your own agent tasks," not more.
What they do not prove
The current scores do not prove that DeepSeek V4 Flash outranks any named competitor overall, or that the vendor numbers will reproduce in your harness. The headline Terminal Bench 2.1 result of 82.7 is DeepSeek-reported, and as of 2026-07-31 DeepSeek does not appear at all among the 17 entries on the benchmark's own official leaderboard, where the top listed pairs score around 83.8 (tbench.ai leaderboard). Two of the nine vendor rows use internal DeepSeek datasets that nobody outside the company can rerun, and the harness used for the public rows had not been released at verification time, so even a motivated third party could not reproduce the exact setup on launch day. Absence of contradiction is not confirmation; it is simply too early.
DeepSeek-Reported 0731 Results
Public agent benchmark scores
DeepSeek's July 31 release materials report the following scores for DeepSeek-V4-Flash-0731 on public agent benchmarks (changelog, model card). All rows are vendor-reported.
| Benchmark | DeepSeek-V4-Flash-0731 (vendor-reported) |
|---|---|
| Terminal Bench 2.1 | 82.7 |
| NL2Repo | 54.2 |
| Cybergym | 76.7 |
| DeepSWE | 54.4 |
| Toolathlon-Verified | 70.3 |
| Agents' Last Exam | 25.2 |
| AutomationBench Public | 25.1 |
The table is dominated by terminal, repository, and tool-use agent tasks, which matches DeepSeek's framing of 0731 as an agent-focused post-training update; the model itself kept the Preview's 284B total / 13B activated architecture (see the 0731 release page for what changed). DeepSeek's footnotes state that the public Code Agent tasks were run with the DeepSeek Harness in minimal mode at maximum reasoning effort. Treat every cell as "DeepSeek says," pending independent reruns.
Internal DSBench results
Two additional rows in the same table - DSBench-FullStack at 68.7 and DSBench-Hard at 59.6 - use DeepSeek's own internal test sets: an internal full-stack development set and an internal coding-agent hard-problem set, per DeepSeek's footnote (changelog). Internal datasets are not automatically dishonest; vendors often build them because public benchmarks saturate or leak into training data. But they are unverifiable by construction: no outside party can inspect the tasks, check the grading, or rerun the evaluation. The practical rule is to read the DSBench rows as directional evidence of where DeepSeek focused its post-training, and to exclude them from any cross-vendor comparison, because no competing model was graded by a neutral party on those tasks either.
The Test Settings Behind the Scores
DeepSeek Harness minimal mode
DeepSeek's footnote says the public Code Agent benchmark tasks were executed through "DeepSeek Harness minimal mode" - an agent scaffold that, at verification on 2026-07-31, DeepSeek had not yet published (changelog). This matters more than it might sound. Agent benchmarks like Terminal Bench do not test a bare model; they test a model plus the loop that feeds it terminal output, retries failures, and decides when to stop. An unreleased harness means the single biggest uncontrolled variable in the headline numbers is currently invisible, and the vendor scores cannot yet be separated into model quality and scaffold quality. Once DeepSeek ships the harness, the setup becomes reproducible and the numbers can be checked - watch for that release before treating 82.7 as settled.
Maximum reasoning effort and sampling settings
The same footnote pins the generation settings: reasoning effort at max, temperature 1.0, top_p 0.95. Two consequences follow. First, the scores represent the model's ceiling configuration, not its default or cheapest one - max effort produces more reasoning tokens, which costs more money and time per task (the API guide explains how the low, high, and max effort levels are set). If you run DeepSeek V4 Flash at low effort to save cost, you should not expect the published numbers. Second, temperature 1.0 sampling means individual runs vary; the footnote does not state how many repeats were averaged per task. Any private rerun should therefore fix the same settings, report the repeat count, and expect run-to-run variance rather than a single deterministic score.
Why Harness Choice Matters
Model quality versus agent scaffolding
An agent benchmark score is a joint measurement of model and scaffold, and the scaffold's share is large. The scaffold decides what context the model sees, how tool errors are surfaced, how many retries are allowed, and when the task is declared done. The same weights can pass a task in one harness and fail it in another purely because one loop recovered from a truncated file write and the other did not. That is why this site keeps vendor scores, independent scores, and community observations in separate buckets, and why "DeepSeek V4 Flash scored X" should always trigger the follow-up question: in whose loop, at what effort, with how many retries?
What community harness comparisons found
Community tests - small-sample and not controlled experiments, but instructive - repeatedly show harness effects comparable in size to model effects. One r/LocalLLaMA user ran the same self-hosted DeepSeek V4 Flash on the same coding task through Claude Code, OpenCode, and Pi and reported similar diff quality but large differences in wall time, tool calls, and tokens (community report, 2026-07-26). Another user watched the model fail a bug in one harness and fix it in two attempts after switching harness (community report, 2026-04-27). A blogger reported a roughly 20-point grade swing for V4 Pro from a harness swap alone (community observation, 2026-05-04 - about Pro, not Flash 0731, cited only to illustrate the dynamic). If you use the model through Claude Code, our Claude Code setup guide covers the harness-specific behavior.
Independent Evidence Available at Launch
The Artificial Analysis signal
Artificial Analysis is, as far as our launch-day search found, the only third party that published its own quantitative evaluation of the 0731 release on 2026-07-31. Using its Intelligence Index v4.1 - nine benchmarks including Terminal-Bench 2.1, GPQA Diamond, SciCode, and its own agentic GDPval-AA v2 - it scored DeepSeek V4 Flash 0731 at 50, ten points above the prior Flash version, and reported the GDPval-AA v2 Elo at 1559, up from 1189 (article). Two of its side findings matter for buyers: the model is unusually verbose, consuming about 210M output tokens across the eval suite versus a median of roughly 62M for comparable models, and its measured hallucination-rate improvement did not come with accuracy gains. Verbosity feeds directly into real cost - see the pricing breakdown for why output tokens dominate agent bills.
Why launch-day metadata needs verification
The Artificial Analysis signal carries explicit launch-day caveats. At retrieval on 2026-07-31, its providers sub-page stated that no provider benchmark runs existed yet for this model entry, listed output speed and time-to-first-token as no data available, and tracked only DeepSeek's first-party API (providers page). Earlier the same day the page still labeled the model proprietary even though the MIT-licensed 0731 weights were already public on Hugging Face - a concrete example of metadata lagging a same-day release. None of this invalidates the Intelligence Index run itself, which used a disclosed methodology. It does mean secondary fields on a launch-day page - license labels, speed numbers, provider lists - should be rechecked before being quoted.
Preview Scores Versus 0731 Scores
Different model revisions
DeepSeek-V4-Flash-Preview (April 24) and DeepSeek-V4-Flash-0731 (July 31) share an architecture but are different model revisions: DeepSeek states 0731 was re-post-trained with the same structure and size (changelog). Their benchmark tables are also largely different suites: the Preview card reported, for example, GPQA Diamond 88.1 and LiveCodeBench 91.6 at max effort (Preview card), while the 0731 table is agent-benchmark-centric. Do not stitch these tables into one "before and after" chart, and do not attribute Preview scores to 0731 or vice versa. Where DeepSeek's own July 31 comparison table lists both revisions on the same benchmark - Preview at 61.8 and 0731 at 82.7 on Terminal Bench 2.1 - the comparison is at least internally consistent, but it remains vendor-run.
Different benchmark and harness versions
Even the same benchmark name can hide a version change. Terminal Bench 2.1 keeps the same 89-task collection as 2.0 but repaired 28 tasks that had broken evaluation - drifted dependencies, too-tight resource budgets, instruction/test mismatches (tbench.ai changelog). Secondary reporting on the 2.1 release notes that one unchanged model/harness pair improved by 12.1 percentage points purely from those repairs. So the Preview card's Terminal Bench 2.0 score of 56.9 at max effort (Preview card) cannot be subtracted from 0731's 82.7 on version 2.1 to compute an improvement: the delta mixes a real model change, a benchmark revision effect of unknown size, and possibly different harness versions. This page never merges 2.0 and 2.1 numbers in one comparison, and you should distrust any chart that does.
How to Reproduce a Fair Benchmark
The minimum test card
A benchmark result is only as useful as its test card. To make a DeepSeek V4 Flash run comparable - to the vendor table, to another model, or to your own run next month - record at minimum: the exact model revision (DeepSeek-V4-Flash-0731, not just the stable deepseek-v4-flash API ID, which now silently serves the new revision), API or local weights plus quantization, harness name and version, reasoning effort, temperature and top_p, the context length actually configured (see the 1M context explainer), the task set and its version, and the run date. Every one of those fields has already caused real disagreements in community results this year. A score without this card is an anecdote.
Raw outputs, failed runs, and repeat counts
Publish what failed, not only what passed. With temperature 1.0 sampling - DeepSeek's own documented benchmark setting - individual runs vary, so a single pass or fail per task is mostly noise; report how many repeats you ran and the pass rate across them. Keep raw transcripts, including tool calls and error output, so others can check whether a "failure" was the model, the scaffold, or a broken task environment - the exact ambiguity that forced Terminal Bench to repair 28 of its 89 tasks in version 2.1. And do not quietly drop crashed runs: a run that hit a rate limit or a harness bug is deployment data, even if you exclude it from the headline pass rate with a stated reason.
Cost, latency, and retries
Record total tokens (input, cached input, reasoning, output), wall time, and retry count per task, because these change the verdict as much as pass rate does. Artificial Analysis measured the 0731 model generating about 210M output tokens across its eval suite versus a roughly 62M median for comparable models - a 3x-plus verbosity gap that per-token prices hide. At DeepSeek's official rates as of 2026-07-31 ($0.14/M cache-miss input, $0.28/M output - see current prices, source: official pricing page), a verbose model can still be cheap in absolute terms, but a fair comparison against a terser competitor must multiply tokens by price, not compare price sheets. Max effort raises both quality and cost, so benchmark at the effort you will actually run.
Benchmark Score Versus Real Task Value
Pass rate and accepted diffs
For coding agents, the number that predicts real value is not a leaderboard score but the rate at which the model produces changes you actually accept. A diff can pass a benchmark's tests and still be unacceptable in a real repository - wrong style, needless rewrites, deleted comments, or a patch that fixes the symptom instead of the cause. Community reports on DeepSeek V4 Flash consistently describe it as a strong executor that benefits from explicit plans and acceptance criteria, with users adding rules like "do not mark complete until tests pass" (community observations, small samples, various harnesses). When you evaluate it, count accepted diffs on your own repositories alongside any public score.
Cost per successful task
The metric that ties everything together is dollars per accepted result: total spend across all attempts, including failed runs and retries, divided by tasks that actually succeeded. A model with a lower per-token price but more retries or triple the output verbosity can cost more per success than a nominally pricier model - one launch-day commenter measured the model needing several times more tokens than a competitor for the same work (community comment, unverified), while others report full days of agent coding for under a dollar. Both can be true on different workloads. That is also the right frame for the Flash-versus-Pro question: the Flash vs V4 Pro comparison matters less as a score duel than as a cost-per-success question on your tasks.
DeepSeek V4 Flash Benchmark FAQ
What is the Terminal Bench 2.1 score?
DeepSeek reports 82.7 on Terminal Bench 2.1 for DeepSeek-V4-Flash-0731, measured with its unreleased DeepSeek Harness in minimal mode at max reasoning effort (changelog). This is a vendor-reported number: as of 2026-07-31 the official tbench.ai leaderboard does not list DeepSeek among its entries, so treat 82.7 as DeepSeek's claim pending a leaderboard listing or independent rerun.
Are the 0731 benchmarks independently verified?
Mostly not yet. As of 2026-07-31, the only independent quantitative evaluation our search found is Artificial Analysis's Intelligence Index run (score 50, methodology disclosed). No independent rerun of DeepSeek's specific agent-benchmark table exists - the harness DeepSeek used was not yet public, and two DSBench rows use internal datasets. That search covered general web search plus the tbench.ai leaderboard on launch day; arena-style leaderboards were not individually checked.
Does DeepSeek V4 Flash beat V4 Pro?
Only in a narrow, dated sense: DeepSeek's July 31 table shows 0731 outperforming the V4 Pro Preview on the listed agent benchmarks, in DeepSeek's own harness. That is not an all-task verdict, it says nothing about a future official Pro release, and Preview-era Pro results on knowledge benchmarks were stronger than Flash's. The DeepSeek V4 Flash vs V4 Pro page covers the comparison properly.
Can Terminal Bench 2.0 and 2.1 be compared directly?
No. Version 2.1 repaired 28 of the 89 tasks that had broken or inconsistent evaluation, and secondary reporting notes an unchanged model/harness pair gained 12.1 percentage points from the repairs alone (tbench.ai). Any Preview-to-0731 delta that crosses the 2.0/2.1 boundary mixes model improvement with benchmark revision effects of unknown size.
Why does reasoning effort matter?
Because the published scores were produced at max effort, and effort level changes results dramatically. On the Preview model card, GPQA Diamond moved from 71.2 without thinking to 88.1 at max effort (Preview-revision data, source). Max effort also generates far more reasoning tokens, so it costs more per task. Benchmark and budget at the effort you will really use; the API guide shows how to set it.
Which benchmark is most relevant for coding agents?
Of the public rows, Terminal Bench 2.1 and DeepSWE are closest to real agent coding work - multi-step terminal tasks and repository-level software fixes. But harness sensitivity means no public score replaces a small pilot: five to ten bounded tasks from your own backlog, run in your own harness at your production effort level, with cost and retries logged, will tell you more than any launch-day table.