DeepSeek says its V4.1 Flash model outperforms its prior flagship while cutting inference costs and lifting speed, using a 552-billion-parameter Mixture-of-Experts design and a new Causal-Encoder-Decoder architecture. Strip away the benchmark boasts against rivals like Kimi K3, and the real signal is strategic: China's labs are competing not on raw capability alone but on the unit economics of inference. When only a fraction of parameters activate per query, the cost curve bends downward fast, and that is where the competitive war is now being fought.
Globally, this pressures the pricing assumptions of frontier vendors. If a Chinese open-weight lineage keeps closing the quality gap while undercutting token costs, the premium that OpenAI, Anthropic, and Google command on API pricing becomes harder to defend for routine coding and agentic workloads. The risk for Western incumbents is commoditization from below: enterprises increasingly route high-volume, low-differentiation tasks to whatever is cheapest and good enough, reserving premium models for the hard cases. Coding and cyber benchmarks matter here because those are exactly the workloads that scale into millions of daily calls.
For Japan, the calculus is more constrained than in most markets. Data-residency rules, procurement caution, and unresolved questions around Chinese-origin models mean many Japanese enterprises and government-adjacent buyers will hesitate to deploy DeepSeek directly, regardless of price. But the indirect effect is real: cheaper open-weight models reset the reference price for every AI proposal an SIer puts in front of a client. NTT Data, Fujitsu, NEC, and the broader SIer tier will find it harder to justify premium managed-AI margins when a client's CTO can point to a model delivering comparable coding output at a fraction of the cost.
The pragmatic path for Japanese SIers is to treat the model layer as increasingly interchangeable and move value up the stack: domain integration, security review, Japanese-language tuning, and the systems-integration glue that connects an LLM to legacy ERP and mainframe estates. RPA vendors face a sharper version of the same squeeze, as capable low-cost coding models absorb the scripted-automation tasks that once justified standalone tooling. Local development teams benefit most concretely: falling inference costs make it viable to embed AI assistance across more of the SDLC, but leadership should build an abstraction layer that lets them swap models as this price war continues, rather than hard-wiring to any single vendor.