A DSpark paper signed by Liang Wenfeng and released in June 2026 said DeepSeek-V4’s online service saw an 85% increase in generation speed under real-world traffic.
According to PANews, the paper said the improvement was not solely due to hardware upgrades, but was achieved by using confidence-based scheduling to reduce compute waste from ineffective verification.
The report described large-model competition as shifting from a focus on parameter scale toward system-level engineering centered on inference efficiency and computing cost.