In August 2026, the ARC Prize Foundation released independently verified benchmark results for DeepSeek V4 Flash 0731, and the numbers sent shockwaves through the AI community. DeepThink’s reasoning engine, powering the V4 Flash model, achieved 89.0% accuracy on ARC-AGI-1 at just $0.02 per task and 61.4% on ARC-AGI-2 at $0.04 per task — a result that redefines what’s possible in cost-effective AI reasoning.
Why ARC-AGI Matters
The ARC-AGI benchmark series is not your typical AI evaluation. Designed by François Chollet, creator of the original ARC challenge, these tests measure fluid intelligence — the ability to solve entirely novel problems without relying on memorized training data. Each task presents a grid-based visual reasoning puzzle with a handful of input-output examples, requiring the model to infer transformation rules from scratch.
- ARC-AGI-1: Tests basic pattern induction across roughly 800 tasks
- ARC-AGI-2: A significantly harder benchmark with multi-step reasoning and symbolic interpretation, launched in March 2025
- ARC-AGI-3: An interactive benchmark for agentic intelligence, where models must explore environments autonomously
Human participants solve ARC-AGI-1 tasks in about 30 seconds on average, and ARC-AGI-2 tasks in about 300 seconds. For years, frontier models struggled to exceed 20% on ARC-AGI-2 — until reasoning models like DeepThink came along.
The Breakthrough Numbers
ARC Prize evaluated DeepSeek V4 Flash 0731 across three reasoning-effort tiers: Low, High, and Max. The results at Max effort tell a compelling story:
| Reasoning Effort | ARC-AGI-1 Score | Cost per Task | ARC-AGI-2 Score | Cost per Task |
|---|---|---|---|---|
| Low | 84.0% | ~$0.01 | 46.0% | ~$0.02 |
| High | 87.0% | ~$0.015 | 56.0% | ~$0.03 |
| Max | 89.0% | $0.02 | 61.4% | $0.04 |
These are not vendor-published numbers — they are independently verified by the ARC Prize Foundation, lending them enormous credibility.
Key Takeaways
- On ARC-AGI-1, V4 Flash Max is among the top-tier results while costing fractions of what competitors charge.
- On ARC-AGI-2, the harder benchmark, V4 Flash achieves 61.4% — a score that compares favorably to Kimi K3 and approaches GPT-5.6 Luna, at a fraction of the cost.
- The gap between ARC-AGI-1 and ARC-AGI-2 scores (27.6 percentage points) confirms that ARC-AGI-2 remains the more challenging benchmark for the industry.
How DeepThink Powers This Performance
The secret to V4 Flash’s remarkable cost-performance ratio lies in its Mixture-of-Experts (MoE) architecture and DeepThink’s reasoning optimization:
- 284 billion total parameters, but only ~13 billion active per token — active parameters drive inference cost
- Compressed Sparse Attention (CSA) and Hierarchical Compressed Attention (HCA) mechanisms balance performance with computational efficiency
- DeepThink’s chain-of-thought reasoning enables transparent, multi-step problem-solving without excessive token overhead
- Context caching reduces costs for repeated prompts in the 1-million-token context window
The architecture achieves what was previously considered impossible: near-frontier reasoning capability at commodity pricing. Cached input costs drop to as low as 0.02 yuan per million tokens.
Comparison: How V4 Flash Stacks Up
On the ARC Prize cost-vs-score chart, the comparison with leading models reveals a paradigm shift:
- Against Kimi K3: V4 Flash at Max effort beats Kimi K3 on the cost-vs-score tradeoff, despite costing significantly less per task.
- Against GPT-5.6 Luna: V4 Flash achieves roughly comparable scores at a fraction of Luna’s cost.
- Against DeepSeek V3: The MoE architecture reduces compute per token to approximately 10% of V3’s requirements, with KV Cache size reduced to just 7%.
On the Artificial Analysis Intelligence Index v4.1, V4 Flash Max scored 50 while Claude Opus 4.8 Max scored 56. However, running the entire benchmark suite cost $72.02 for V4 Flash versus $3,752.55 for Opus — a 52x cost difference for a 6-point score gap.
What This Means for Developers
The ARC Prize results have immediate practical implications for AI developers and enterprise teams:
1. Democratizing Reasoning Quality
Previously, state-of-the-art reasoning was reserved for organizations with deep budgets. V4 Flash’s pricing makes frontier-level reasoning accessible to startups, individual developers, and cost-sensitive enterprises.
2. Agent Workloads Become Economically Viable
Agentic tasks — where AI systems autonomously read files, execute code, and complete complex workflows — are token-intensive. V4 Flash’s low cost per task makes long-horizon agent tasks practical for the first time. A single agent workflow that would have cost hundreds of dollars with flagship models can now run for pennies.
3. The Efficiency Frontier Moves Faster Than Intelligence
The ARC-AGI results demonstrate that the efficiency frontier is advancing more rapidly than raw intelligence. DeepThink-powered models are closing the gap on reasoning benchmarks faster than competitors can reduce their prices, creating a self-reinforcing cost advantage.
4. Self-Hosted Alternative
For organizations with data privacy requirements, V4 Flash’s open weights enable local deployment on consumer hardware. A 284B-parameter MoE model with DeepThink reasoning can run on a laptop — no dedicated GPU required.
The Bigger Picture
The ARC Prize benchmarks are more than just numbers on a leaderboard. They represent a fundamental shift in the AI industry:
- From capability-driven to cost-driven competition: When a model can score 89% on a top-tier reasoning benchmark while costing $0.02 per task, the market dynamics change entirely.
- From benchmark chasing to practical deployment: Developers are no longer choosing models based solely on maximum capability. They’re optimizing for the most reasoning per dollar.
- From brand loyalty to performance-per-dollar: As one OpenAI executive discovered when promoting price cuts on social media, the community response was clear: “We want DeepSeek’s prices, not yours.”
Looking Ahead
The ARC Prize results for V4 Flash are a milestone, not a ceiling. As DeepThink continues to optimize its reasoning engine and as hardware improves, we can expect:
- Even lower cost-per-task for high-reasoning-effort workloads
- Better performance on ARC-AGI-2 and emerging ARC-AGI-3 benchmarks
- Wider adoption of DeepThink-powered pipelines across scientific research, enterprise automation, and developer tooling
- A continued shift in industry pricing baselines toward the cost-performance frontier DeepThink has established
Conclusion
DeepThink V4 Flash’s ARC Prize results — 89% on ARC-AGI-1 at $0.02, 61.4% on ARC-AGI-2 at $0.04 — are more than just benchmark numbers. They signal a new era where high-quality AI reasoning becomes a commodity, accessible to everyone from independent developers to Fortune 500 enterprises.
The efficiency frontier has moved. The question is no longer “who has the smartest model?” but “who delivers the most reasoning per dollar?” And for now, DeepThink is leading that race.