DeepSeek-V4-Flash leaves preview and challenges the Opus 4.8 performance crown

More than expected from a compact model
DeepSeek has moved DeepSeek-V4-Flash out of preview and added it to its API. The new model not only outpaces the previous version, V4-Pro-Preview, but also nips at the heels of Opus 4.8, the heavyweight, expensive flagship model. The jump in performance is significant, especially when the price tag is considered.
Data that speaks about agents
On the DeepSWE benchmark, which tests autonomous software-engineering capabilities, V4-Flash achieves a 54.4 percent success rate, leaving GLM-5.2 behind and coming close to Opus 4.8. In TerminalBench, which simulates command-line work, it reaches 82.7 percent, again above GLM and level with Opus. The results show that a compact, inexpensive model can run complex automation tasks at nearly the same level as the market’s pricey offerings.
Post-train makes the difference
The most striking detail is that the architecture and size of the model remain identical to the previous version. All improvement comes from post-train work—targeted fine-tuning after the initial training phase. No parameters were added and the model was not enlarged; it simply learned to use its existing capacity more effectively.
A price that drives the market crazy
Pricing continues DeepSeek’s aggressive strategy. Input tokens cost $0.14 and output tokens $0.28, making the model tens of times cheaper than Opus 4.8. Users who employ caching—preserving context across calls to save costs—pay only $0.0028 per token. In the current market, V4-Flash becomes a serious competitor to models such as GPT-5.6 Luna and Terra, which deliver similar performance at much higher prices.
What's next
Despite the Flash version’s success, DeepSeek has not forgotten the full model. The company promised to release the complete DeepSeek-V4-Pro “as soon as possible” (ASAP), meaning competition in the agents-and-code segment is set to heat up further in the coming months.