Independent Benchmarks Confirm DeepSeek V4.1 Flash as the Top Open-Weight Model: Vals.ai and Artificial Analysis Results
When a model lab publishes benchmark numbers for its own model, the natural response is skepticism. Self-reported scores are displayable evidence, not independent verification. DeepSeek’s September 10 launch of V4.1 Flash came with an impressive benchmark table — Codeforces 3471, Terminal-Bench 2.1 at 90.6, DeepSWE v1.1 at 74.2 — but these were DeepSeek’s own numbers, measured with DeepSeek’s own evaluation harness.
Within days, two independent evaluation organizations published their own results. Vals.ai, a benchmark platform that runs models through a proprietary suite of agentic and reasoning tests, ranked V4.1 Flash as the new number one open-weight model on its Vals Index. Artificial Analysis, another independent evaluation service, assigned it an Intelligence Index of 40 and reported the highest AutomationBench-AA score of any model tested.
The independent results confirm what DeepSeek claimed, with important nuances that the self-reported numbers do not capture.
Vals.ai: The Open-Weight Champion
Vals.ai evaluated DeepSeek V4.1 Flash across its full benchmark suite on September 10, 2026. The headline result: 57.86% on the Vals Index, making it the top open-weight model and ranking 15th overall among all 56 evaluated models (including proprietary systems from OpenAI, Anthropic, Google, and others).
The cost-effectiveness story is more striking than the accuracy number alone suggests. V4.1 Flash costs $0.30 per test on the Vals Index. The next-best open-weight model, Kimi K3, scores 57.81% — just 0.05 points behind — but costs $6.47 per test, roughly 21.6x more expensive. V4.1 Flash is the cheapest model in the Vals Index top 15.
Where V4.1 Flash Dominates
The Vals.ai results reveal specific areas of exceptional strength:
| Benchmark | V4.1 Flash Score | Rank | Notable Comparison |
|---|---|---|---|
| Vals Index | 57.86% | 15/56 | #1 open-weight, $0.30/test vs Kimi K3’s $6.47 |
| Code Migration | 45.62% | 9/58 | #1 open-weight; GLM 5.3 costs $24.91 for 44.22% |
| SkillsBench | 69.80% (with skills) | 1/34 | #1 overall, not just open-weight |
| Vibe Code Bench | 84.74% | 7/96 | #2 open-weight, 40x cheaper than Kimi K3 |
| Terminal-Bench 2.1 | 74.53% | — | #2 open-weight across three full trials |
The SkillsBench result deserves special attention. V4.1 Flash is not just the top open-weight model on this benchmark — it is the top model overall, beating every proprietary system tested. SkillsBench measures the model’s ability to use external tools and APIs to complete complex, multi-step tasks. The fact that an open-weight model leads this category is a signal that DeepSeek’s post-training pipeline, which includes large-scale agent task synthesis and reinforcement learning, is producing genuinely superior agentic behavior.
The Cost-Per-Task Revolution
The Vals.ai results expose a pattern that extends beyond raw accuracy. V4.1 Flash is not just cheaper per token than its competitors — it is cheaper per completed task, because it finishes tasks in far fewer tokens.
Compared to its predecessor, DeepSeek V4 Flash 0731, V4.1 Flash gains 4.3 points on the Vals Index. But the efficiency story is more dramatic: despite having list prices two to four times higher than the previous Flash generation, V4.1 Flash costs less per test on most agentic benchmarks because it completes the same tasks with substantially fewer tokens. It also roughly halves latency on Vibe Code Bench, EMB (embedding benchmark), Legal Research, and Harvey’s Legal Agent Benchmark.
This is the Engram and CED architecture paying off in production economics. The asymmetric prefill/decode split (8B/16B active parameters) and the 890-bytes-per-token KV cache footprint are not just engineering achievements — they translate directly into faster, cheaper task completion.
Artificial Analysis: Intelligence Index and Beyond
Artificial Analysis published V4.1 Flash results on September 10, assigning it an Intelligence Index of 40. For context, the median open-weight model scores 18 on this index. Top proprietary models cluster in the 50-55 range after the v4.3 methodology tightening. V4.1 Flash sits at 40 — above the open-weight median by more than 2x, within striking distance of the proprietary frontier.
AutomationBench-AA: The Highest Score
The most notable Artificial Analysis result is on AutomationBench-AA, where V4.1 Flash scored 68.9% — the highest of any model tested. This benchmark tests agents on 657 business workflows across simulated applications, measuring whether the agent can follow business rules while navigating complex application interfaces. This is not a coding benchmark or a reasoning benchmark; it is a workflow automation benchmark that directly mirrors enterprise use cases.
Scoring highest on this benchmark is a signal that V4.1 Flash is not just a strong model — it is a strong agent. The difference matters. A model generates text; an agent takes actions, manages state, handles errors, and completes multi-step workflows. The post-training pipeline that DeepSeek used for V4.1 Flash — including large-scale agent task synthesis and reinforcement learning in synthesized environments — appears to have produced agent capabilities that transfer to real-world business automation.
Speed: 214 Tokens Per Second
Artificial Analysis also measured V4.1 Flash’s generation speed at 214.4 tokens per second — roughly three times the market average. This speed advantage is a direct consequence of the CED architecture’s asymmetric parameter activation. During decode (the token generation phase), the model activates 16 billion parameters — a fraction of the 552 billion total. The reduced active parameter count means less computation per token, which translates to higher throughput.
The speed advantage has compounding effects on cost. At 214 tokens per second, V4.1 Flash completes a 10,000-token reasoning trace in approximately 47 seconds. A model running at the market average of ~70 tokens per second would take 143 seconds for the same output. For agentic workloads that generate long reasoning traces before delivering final answers, the speed differential multiplies into dramatically lower total task cost and latency.
The Codeforces 3471 Story
The Codeforces rating of 3471 deserves separate examination. Codeforces is a competitive programming platform where participants solve algorithmic problems under time constraints. A rating of 3471 places V4.1 Flash in the top tier of human competitive programmers — the “International Grandmaster” level, which fewer than 200 human programmers have achieved.
Independent verification of this score is available through BenchLM.ai’s Codeforces leaderboard, which tracks published scores from model providers. As of September 18, 2026, V4.1 Flash leads the leaderboard at 3471, followed by V4 Pro 0813 at 3206 and dots3-note Preview at 3056. The gap between first and second is 265 rating points — a substantial margin in competitive programming.
However, BenchLM.ai notes important caveats. All six entries on the Codeforces leaderboard are provider self-reports; zero come from the benchmark owner or an independent run. The scores should be read as displayable evidence rather than verified results. This is not a criticism of DeepSeek specifically — every provider on the leaderboard self-reports — but it means the Codeforces number, while impressive, has not been independently replicated.
Where V4.1 Flash Does Not Lead
The independent benchmarks also clarify where V4.1 Flash falls short of the proprietary frontier, and where its predecessor V4 Pro still holds advantages.
Deep Reasoning and Knowledge Tasks
On GPQA Diamond (graduate-level science reasoning), V4.1 Flash scores 90.9 — strong, but below V4 Pro’s 92.4. On the text subset of Humanity’s Last Exam, Flash scores 39.1 versus Pro’s 42.7. These are knowledge-intensive reasoning tasks that reward deep, careful analysis over fast, efficient generation.
The gap is consistent with the architectural design. V4.1 Flash is optimized for agentic workloads — input-heavy, tool-using, multi-step tasks. V4 Pro, with its symmetric architecture and larger active parameter count, retains an edge on pure reasoning tasks that do not benefit from the asymmetric compute split.
Specialized Domain Benchmarks
On certain specialized benchmarks, V4.1 Flash ranks lower. MedCode (medical coding) ranks it 47th of 93 models on Vals.ai. LegalBench ranks it 52nd of 145. These are domains where domain-specific fine-tuning matters more than general reasoning capability, and where the open-weight model has not received the specialized post-training that proprietary alternatives may have.
The Open-Weight Implication
The Vals.ai and Artificial Analysis results collectively establish something that was previously uncertain: an open-weight model can lead specific benchmark categories — not just coding or reasoning, but agent automation and workflow execution — against proprietary alternatives from labs with substantially larger compute budgets.
This matters for three reasons:
-
Reproducibility: The model weights are publicly available on Hugging Face under an MIT license. Any researcher can download them, evaluate them, and build on them. The benchmark results are verifiable by anyone with sufficient compute.
-
Cost accessibility: At $0.15 per million input tokens (off-peak, cache miss) and $0.60 per million output tokens, V4.1 Flash is accessible to individual developers, small startups, and academic researchers. The cheapest proprietary model in the Vals Index top 15 costs substantially more per test.
-
Self-hosting: DeepSeek has explicitly invited organizations planning large-scale deployments with 2,000+ GPUs to contact them for deployment support. The model can be self-hosted, eliminating per-token API costs entirely for organizations with sufficient infrastructure. NVIDIA has already published deployment recipes for V4-Flash on B200 and H200 GPU configurations, and AMD has published a playbook for running DeepSeek V4 Flash on Ryzen AI Halo platforms using the ds4 inference engine.
What the Independent Results Mean for the Industry
The Vals.ai and Artificial Analysis results confirm a shift in the AI landscape that has been building since DeepSeek R1’s open-source release in early 2025. The gap between open-weight and proprietary models is no longer a chasm. On agentic and coding benchmarks, it has closed to a few percentage points. On cost, open-weight models have opened a gap in the opposite direction — 20x cheaper per test than the nearest open-weight competitor, and substantially cheaper than any proprietary model in the top tier.
The implication is that model differentiation is shifting. Accuracy alone is no longer the primary axis of competition. The questions that matter now are: How cheaply can the model complete a task? How fast can it deliver the answer? How reliably can it follow complex multi-step instructions? And — increasingly — can I run it on my own hardware?
On all four questions, V4.1 Flash’s independent results suggest the answer is shifting in favor of open-weight models. Whether that lead holds when the next generation of proprietary models — GPT-6 Astra, Claude Fable, Gemini 3 — arrives is the question the benchmark platforms will answer next.