The headline warning from IDC and IEIT Systems is that the global compute supply-demand gap will keep widening through 2030, with high-bandwidth memory as the primary chokepoint. The more interesting story sits underneath that number: the nature of demand is changing, not just its size.

Training a frontier model is a lumpy, project-based workload. Running fleets of AI agents is the opposite. Agents chain reasoning steps, call tools, retry, and hold context across long-running tasks, which means each unit of real work consumes far more inference cycles than a single chatbot query ever did. That converts AI compute from a capital-expenditure spike into a continuous operating load that scales with headcount and workflow adoption. HBM is where this bites hardest, because agentic inference is memory-bandwidth-bound, not just FLOP-bound. You cannot solve it simply by fabricating more logic dies when the constraint is stacked DRAM and advanced packaging capacity that only a handful of suppliers control.

The global implication is a bifurcation of pricing power. Hyperscalers with locked-in HBM allocation and custom silicon will ration capacity internally and prioritize their own agent products, while everyone downstream faces volatile inference costs. Expect inference pricing to stop falling as smoothly as it has, and expect "agent economics" to become a board-level line item rather than an engineering footnote. The Italian nuclear vote and the broader energy scramble are the same story viewed through the power grid: memory and megawatts are now the two hard limits on how far agents can scale.

For the Japanese market, this constraint cuts both ways. On the supply side, Japan sits unusually close to the bottleneck: materials, packaging chemistry, and the equipment that enables HBM stacking run through Japanese firms, and Rapidus-era ambitions gain strategic weight precisely because advanced memory and packaging are the scarce inputs. That is a genuine, durable opportunity rather than a hype cycle.

On the demand side, the picture is harder for Japanese enterprises and SIers. Domestic AI infrastructure remains thin relative to US hyperscalers, and a widening global compute gap means Japanese buyers compete for allocation from a weaker position, often paying in a soft yen. SIers built on labor-arbitrage and multi-year system-integration contracts should treat agent inference cost as a first-class design constraint now. The teams that win will architect for efficiency by default: aggressive caching, smaller task-specific models, on-prem or edge inference for steady workloads, and reserved capacity contracts negotiated early. RPA vendors face the sharpest repricing. Deterministic rule-based bots are cheap to run; replacing them with agentic reasoning is powerful but far more compute-hungry, and in a constrained-HBM world the naive "agent-ify everything" pitch will collide with unit economics. The pragmatic path for local dev teams is hybrid: keep deterministic automation where rules suffice, reserve agents for genuinely ambiguous work, and instrument token and memory cost per task from day one.