The center of gravity in AI infrastructure is quietly shifting. For three years the operative metric was accelerator count — more GPUs, larger clusters, bigger training runs. Agentic AI breaks that logic. When an autonomous agent chains dozens of reasoning steps, calls tools, and re-queries context to complete a single task, the binding constraint stops being how much compute you own and becomes how many useful tokens you extract per watt and per dollar. H3C framing the debate around tokens at Apsara is a signal that the hardware vendor conversation is migrating from procurement volume toward throughput economics.
Globally, this reframes the winners. If token efficiency — not headline FLOPS — determines unit cost, then the advantage flows to whoever optimizes the full stack: inference-optimized silicon, KV-cache management, speculative decoding, batching, and model routing that sends easy queries to small models and hard ones to frontier models. It also changes the buyer's math. An agentic workflow that quietly burns tokens on redundant reasoning can turn a promising pilot into a P&L liability. The next phase of enterprise AI spending will be governed less by capability envy and more by cost-per-completed-task discipline.
For Japanese enterprises, this is the more welcome framing of the AI race. Japan has been structurally cautious on capex-heavy GPU buildout — constrained by power, datacenter siting, and a corporate culture that demands demonstrable ROI before scaling. A token-efficiency paradigm rewards exactly that caution. Firms that never won the raw-compute arms race can compete on how tightly they engineer agent workflows, prune unnecessary reasoning loops, and match model size to task value.
The implication for SIers is sharper still. The domestic integration business has long monetized headcount and system delivery; the emerging opportunity is token governance — architecting agent pipelines where every reasoning step is justified, caching is aggressive, and model routing is deliberate. An SIer that can guarantee a client's agentic system delivers outcomes at a predictable per-token cost holds a defensible position that generic cloud resale cannot match. Expect FinOps to expand into a distinct 'TokenOps' practice.
RPA vendors face the clearest strategic fork. Legacy rule-based automation is being reframed as agentic, but naive LLM-driven agents are far costlier per action than deterministic scripts. The durable design for Japan's process-automation market is hybrid: cheap deterministic execution for the predictable 80%, expensive reasoning reserved for genuine exceptions. Teams that treat tokens as a metered utility, not free compute, will define the economics of the next automation cycle.