The question of whether GPUs keep owning AI compute is really a question about vertical integration. Google's TPU line, Amazon's Trainium and Inferentia, and OpenAI's move into custom silicon share one motive: the largest buyers of accelerators no longer want to pay a merchant markup on their single biggest cost line. When you are training and serving models at hyperscale, designing your own chip stops being a science project and becomes a margin decision.
The strategic shift is from raw performance to workload-specific economics. GPUs win on flexibility and a mature software moat, which is why they still dominate training and any workload that changes shape month to month. But inference is where the volume and the recurring cost sit, and inference is predictable enough to hard-wire into an ASIC. Expect a bifurcation: general-purpose GPUs for the frontier and experimentation, purpose-built silicon for the steady-state serving that pays the bills. The real battleground is not transistors but compilers, kernels, and the toolchains that decide whether a custom chip is usable outside its owner's walls.
For the market, this pressures pricing power at the top of the stack and rewards whoever controls the software layer between model and metal. It also raises a supply-chain point: custom silicon still routes through the same advanced foundries and HBM suppliers, so diversifying chip designs does not diversify the physical chokepoints.
For Japan, the implication is uncomfortable and clarifying at once. Japanese enterprises and SIers overwhelmingly consume AI compute through foreign hyperscalers, which means the cost curve, the roadmap, and the availability of accelerators are all set offshore. As hyperscalers steer customers toward their in-house chips, Japanese teams face a lock-in they did not choose: workloads tuned for one vendor's silicon are expensive to move. SIers building AI platforms for regulated clients should treat accelerator portability as an architecture requirement now, abstracting the serving layer rather than betting on any single instance type.
There is also a domestic opening. Rapidus's advanced-node ambitions, Arm's position inside SoftBank, and existing local accelerator efforts give Japan more of a hand in this game than it had a cycle ago. The pragmatic near-term play for Japanese dev teams and RPA-heavy operations is cost discipline: profile inference workloads, right-size models, and negotiate for the cheaper in-house silicon tiers instead of defaulting to premium GPUs. The winners in Japan will be the integrators who turn compute volatility into a managed, portable service.