DeepSeek disclosed DSec, an elastic compute platform that spins up roughly 3 million ephemeral sandboxes per day, peaking above 380,000 concurrent instances at creation rates over 5,000 per second, to train agentic reinforcement-learning systems.

The headline number distracts from the strategic point. As frontier labs shift from static next-token prediction toward agents that act, observe, and iterate, the bottleneck migrates from raw GPU hours to the environments where agents practice. You cannot teach an agent to file a bug, execute a trade, or navigate a codebase by feeding it text; it needs a safe, disposable world to fail in millions of times. DSec is essentially a factory for those worlds. This reframes the AI arms race: the durable moat is increasingly the orchestration layer that provisions, isolates, and tears down execution environments at planetary scale, not the weights themselves. Weights leak and get distilled. A pipeline that reliably manufactures 3M clean sandboxes a day is far harder to copy.

This also connects directly to the week's security narrative around autonomous agents breaching real systems. The same sandbox infrastructure that trains a helpful coding agent trains one capable of probing live targets. The gap between a training environment and a production attack surface is thinner than most enterprises assume, and DSec-class tooling lowers the cost of building agents that act in the world, for better or worse.

For Japanese enterprises and SIers, the implication is uncomfortable but clarifying. The domestic conversation has fixated on which foundation model to adopt, framing AI strategy as procurement. DSec suggests the higher-value work is environmental: building the simulated ledgers, mock ERP instances, and synthetic operational sandboxes where agents can be trained and validated against a company's actual workflows. SIers like NTT Data, NRI, and Fujitsu already possess something the model labs lack, deep, messy knowledge of how Japanese banks, manufacturers, and government systems actually run. Packaging that domain reality into reproducible agentic training and evaluation environments is a defensible business the hyperscalers cannot easily replicate from abroad.

The RPA incumbents face a sharper reckoning. Rule-based automation assumed a stable, scripted world; agentic RL assumes a world you can simulate and let the agent learn. Japanese firms that treated RPA as the ceiling of automation should treat sandbox-based agent training as the next floor. The winners will be teams that stop asking which model is smartest and start building the environments where their agents earn competence, plus the governance to ensure those same capabilities are not turned against them.