DeepSeek V4 Flash 0731 Tops ARC-AGI-2 at 61.4% for Pennies per Task
DeepSeek V4 Flash 0731 hits 89% on ARC-AGI-1 and 61.4% on ARC-AGI-2 at a fraction of a cent per task, shaking up the ARC Prize leaderboard.
BY FOUNDERBUILT AI NEWS
DeepSeek newest model, V4 Flash 0731, is making waves on the ARC Prize leaderboard. At maximum reasoning effort, the model scores 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private. What makes the result notable is the cost: roughly two cents per task on ARC-AGI-1 and four cents per task on ARC-AGI-2. DeepSeek continues its pattern of publishing strong open-weight results at a fraction of the price of frontier rivals.
The model ships with three reasoning variants, Max, High, and Low effort, letting developers trade accuracy for latency and cost. The Low effort variant still clears 84% on ARC-AGI-1, while High lands at 87%. ARC-AGI-3 results are not yet published. The model page also links to an accompanying paper and model weights, keeping with DeepSeek open approach. The release was flagged on Hacker News within hours, with over 300 points and 180 comments as the community compared it to previous ARC-AGI entries.
For founders, the takeaway is straightforward. Capable reasoning is getting cheaper fast, and ARC-AGI scores are becoming a standard spec sheet item. If a two-cent-per-task model can hold 89% on a benchmark designed to resist memorisation, that changes what is practical to build with small budgets. Expect other labs to answer with their own optimised variants in the coming weeks. Check the ARC Prize results page for full per-task breakdowns.