d-Matrix unveiled Raptor, an inference accelerator that bonds a TSMC 4nm compute die face-to-face onto a custom DRAM die at a 36-micron pitch, aiming for roughly 100 TB/s of bandwidth per card. The bet underneath it matters more than the spec sheet.

Generative inference is increasingly memory-bound, not compute-bound. Feeding large models means moving weights and KV-cache data fast enough to keep expensive logic busy, and that is exactly where HBM-based GPUs hit a wall of cost, power, and supply scarcity. By moving memory into the same 3D package rather than beside the die, d-Matrix is attacking bandwidth-per-dollar and bandwidth-per-watt, the two numbers that actually govern the economics of serving tokens at scale. If it holds up in production, the strategic implication is a widening split between training silicon, where Nvidia's moat is deep, and inference silicon, where the field is genuinely contestable. Hyperscalers and inference-heavy startups have every incentive to fund a second source that loosens HBM dependency.

The risk is equally real. Custom DRAM plus advanced hybrid bonding is hard to yield, hard to test, and hard to scale, and a startup competing on packaging complexity is exposed on cost curves the moment volumes matter. Software maturity, not silicon, usually decides these races.

For Japan, the read is less about buying Raptor and more about who supplies the shift. Advanced packaging and 3D stacking pull directly on Japanese strengths: bonding and dicing equipment from firms like Disco and the broader back-end toolchain, plus materials for interposers, thin-wafer handling, and thermal management. Every architecture that trades monolithic GPUs for stacked compute-on-memory raises demand for exactly the process steps Japan's equipment and chemicals makers dominate.

Japanese enterprises and SIers should track this as an inference-cost story. As memory-centric accelerators mature, the price of running inference on-premise or in domestic clouds could fall enough to reshape build-versus-buy math for regulated sectors that resist sending data offshore. SIers designing AI platforms should avoid hard-coding GPU-only assumptions and instead architect for heterogeneous inference back-ends, because the winning silicon for serving models may not be the winning silicon for training them.