DeepSeek has just reminded the AI industry that a high-performance model does not necessarily have to cost a fortune. Without any press release or spectacular announcement, the Chinese lab qu
DeepSeek has just reminded the AI industry that a high-performance model does not necessarily have to cost a fortune. Without any press release or spectacular announcement, the Chinese lab quietly deployed the final version of its flagship model. Behind this update lies a shift. The performance gap is narrowing, while the price gap is widening. For Web3, autonomous agent developers, and companies facing the growing cost of AI, this new equation could reshuffle the deck. DeepSeek is no longer just trying to compete but is attacking where its competitors are most vulnerable, the price.
In Brief
- DeepSeek deploys the final commercial version DeepSeek-V4-Pro-0813 without official announcement, replacing the preliminary April version.
- Claude Fable 5 leads the Chinese model by only 5.3 % on average over nine benchmarks, a superiority that drops to 2.8% excluding a specific test.
- Anthropic’s API shows a cost 4,500% higher than DeepSeek’s, with an average bill of $30 vs. $0.65 for an equivalent volume.
- Executing an agent task costs more than $31 on Claude Fable 5, versus about $0.04 on the DeepSeek architecture.
This update carried out by DeepSeek, which reignites the global race for artificial intelligence, resulted in a discreet adjustment to the publisher’s pricing grid, where the mention deepseek-v4-pro now identifies the final commercial version DeepSeek-V4-Pro-0813. Until now, the published independent tests only relied on a preliminary version deployed in April. The Chinese company had reminded on July 31 during the launch of its V4-Flash variant that its Pro API remained unchanged and that the final model would follow shortly.
The technical sheet on Hugging Face still displays the V4 series as a preliminary version, meaning no external lab has yet independently evaluated this 0813 version. Of the ten agent evaluation benchmarks published by the company, the competing model Claude Fable 5 from Anthropic holds only a 5.3% average lead on nine compared tests.
The detailed analysis of the evaluations reveals decisive nuances on the real level of performance. The overall 5.3% gap is mainly explained by the test “Humanity’s Last Exam” without tools, where DeepSeek scores 42.7 against 53.3 for Claude Fable 5. Isolating this test showing a 10.6% gap, the American rival’s average superiority collapses to 2.8% on the rest of the tests. It should be noted that DeepSeek measured these results on its own technical infrastructure, via the minimal mode of its “DeepSeek Harness” framework configured at maximum effort level with high creativity. Two of the ten tests, named “DSBench-FullStack” and “DSBench-Hard”, are based on internal test sets without verifiable public ranking.
Here is the summary of performance metrics published by the lab :
- Claude Fable 5 maintains an average lead of 5.3% over nine compared benchmarks ;
- The gap drops to 2.8% when excluding the “Humanity’s Last Exam” without tools test (42.7 vs 53.3) ;
- DeepSeek ranks first directly on two of the ten presented agent benchmarks ;
- The tests rely on the minimal internal “DeepSeek Harness” framework, including two tests without public ranking.
A Pricing Chasm Between the Chinese Alternative DeepSeek and Proprietary Models
On the pricing front, the confrontation turns into an industrial show of strength when examining API access costs. DeepSeek maintains its pricing at $0.435 per million input tokens and $0.87 per million output tokens, with a cost of $0.003625 for cache input. Meanwhile, Anthropic charges Claude Fable 5 at $10 per million input tokens and $50 per million output tokens. Weighting these figures on standard usage ratios, the overall bill comes to about $0.65 at the Chinese maker versus $30 at the American giant, a massive 4,600% markup. This financial gap further widens at the task execution level, the American model reflecting longer and generating more text.
Measurements made by Artificial Analysis thus price the average task at $3.15 for Fable 5 versus 3 cents for V4-Flash, a factor of 105. Clément Delangue, CEO of Hugging Face, estimates the total bill at “more than $31 per task” at Anthropic versus “about $0.04” for its competitor.
Within Anthropic’s own catalog, the presence of Claude Opus 5 complicates the situation, the latter outperforming Fable 5 on most benchmarks at half its price. This price asymmetry thus begins to undermine the coherence of American proprietary product ranges.
Join the ‘Read to Earn’ programThis link uses an affiliate program.Towards a Redistribution of Technological Cards Driven by Open Source
This dynamic fits into a movement where Chinese open-weight labs, like Kimi whose K3 model outperformed Fable 5 and GPT-5.6 Sol at launch, drastically reduce inference costs.
The availability of DeepSeek on Hugging Face under the MIT free license gives developers the possibility to immediately verify this performance on their own servers. This transparency strongly contrasts with the closed ecosystems of Silicon Valley giants, paving the way for permanent community auditing. Ultimately, this democratization of low-cost computing power could transform the global application ecosystem.
It would allow Web3 infrastructures, DeFi protocols, and autonomous agent developers to integrate advanced reasoning capabilities without enduring the rent imposed by American proprietary giants. However, dependence on internal benchmarks invites a nuanced analysis. The future of the market will depend on the ability of open source players to maintain this level of efficiency while offering equivalent security and reliability guarantees in the long term.