Summary
- DeepSeek has transitioned its deepseek-v4-pro to the 0813 build, marking the general release of a model that has been in preview since April.
- In tests across nine benchmarks, Claude Fable outperforms DeepSeek by an average of 5.3%. DeepSeek won two of these comparisons.
- The pricing for Fable 5 is $10 per million input tokens and $50 per million output tokens, while V4 Pro is priced at $0.435 and $0.87, respectively.
On Wednesday, DeepSeek released the finalized version of its flagship model without any formal announcement. The update was noted on the API pricing page, which now indicates the model version as DeepSeek-V4-Pro-0813.
The usage costs for this model remain at approximately $0.435 for input tokens and $0.87 for output tokens per million. While the pricing structure has not changed, the underlying model weights have been updated.
Since its initial rollout in April, DeepSeek V4 Pro has been offered at a price significantly lower than GPT-5 Pro, at 98% less. All previous evaluations conducted by independent labs were based on a preview version. DeepSeek acknowledged this on July 31, stating that the Pro API was “unchanged” and that a formal release would follow shortly. The model card on Hugging Face continues to describe the V4 series as a “preview version.”
As such, the benchmarks circulated widely were based on a version of DeepSeek that the company did not consider final. No independent evaluations of the 0813 version have been reported yet.
Analysis of DeepSeek’s Benchmarks
The company provided a comparison across ten benchmarks. In eight of these, where Fable 5 or another model excelled, the differences were minimal.
When averaging Fable 5’s performance across the benchmarks, it leads by 5.3%. Excluding the Humanity’s Last Exam without tools—where DeepSeek scored 42.7 compared to 53.3 for Fable 5, a 10.6% gap that skews the overall results—the average lead narrows to 2.8%.
The public pricing for both models reveals a stark contrast. Fable 5 charges $10 for every million input tokens and $50 for output, while V4 Pro charges only $0.435 and $0.87, with cached input at just $0.003625. This results in a blended cost comparison of approximately $30 for Fable 5 versus $0.65 for V4 Pro—an astounding difference of about 46 times, or 4,600%. Such a disparity is crucial for businesses that utilize AI tools extensively.
When it comes to cost per completed task, the gap widens further, as Fable 5 processes longer and generates more extensive outputs. According to Artificial Analysis, the cost per benchmark task is $3.15 for Fable 5 compared to just 3 cents for V4-Flash, making it about 105 times cheaper. Hugging Face CEO Clément Delangue noted this difference as exceeding $31 per task for Fable 5 against roughly $0.04 for V4 Pro. There is no existing per-task cost for the 0813 version yet.
The competitive landscape is further complicated by Anthropic’s offerings, as Claude Opus 5 outperforms Fable 5 on most benchmarks while being priced at half the cost.
DeepSeek conducted its own evaluations using infrastructure yet to be disclosed. The company mentioned in July that the DeepSeek Harness minimal mode, which operates at maximum capacity with heightened creativity, would be released soon. Two of the ten benchmarks, DSBench-FullStack and DSBench-Hard, are internal tests without public leaderboards for verification.
Despite these complexities, the trend remains consistent. Chinese open-weight labs are achieving results comparable to those of American counterparts but at a fraction of the cost. For instance, Kimi K3 surpassed both Fable 5 and GPT-5.6 Sol upon its release, while DeepSeek and Xiaomi have managed to reduce frontier costs by 99%, contrasting with rising expenses in U.S. labs. DeepSeek’s model weights are MIT-licensed and available on Hugging Face, allowing for independent verification with ease.
