d-Matrix's Raptor stacks DRAM directly under the compute die, dropping the PHY interface to claim SRAM-class bandwidth at roughly a tenth of HBM's power draw. The pitch is aimed squarely at inference, where KV-cache growth is quietly becoming the dominant cost center.

The strategic point is not the spec sheet but the timing. HBM has become the single most contested resource in AI hardware, and its scarcity is now spilling into consumer pricing, with Amazon raising device prices citing memory costs. When a memory shortage reaches the checkout page for a Kindle, it signals that the constraint is structural, not cyclical. Any credible path that sidesteps HBM's power and supply bottleneck will attract serious attention from hyperscalers desperate to decouple their roadmaps from a three-supplier oligopoly.

The hard part is that 3D DRAM lives or dies on manufacturing. SK hynix, Samsung, and Micron are all pursuing logic-DRAM stacking, and a startup proposing to fuse memory beneath compute needs a foundry and packaging partner willing to co-develop a non-standard flow. That is a multi-year commitment, and incumbents have every incentive to protect HBM margins. Raptor's real test is whether it becomes a licensable architecture that a memory giant adopts, rather than a standalone product.

For Japan, this is more relevant than it first appears. The country's semiconductor revival is concentrated in materials, equipment, and advanced packaging rather than leading-edge logic. Firms in photoresists, bonding, and test are the enablers any 3D DRAM approach depends on, and a shift toward hybrid bonding and die stacking plays directly to that strength. Rapidus and the domestic packaging push should watch which memory architecture wins, because it determines where the next capex flows.

For Japanese enterprises and SIers, the near-term signal is cost. Rising memory prices push inference economics in the wrong direction just as firms move from pilots to production. That strengthens the case for smaller task-specific models, aggressive quantization, and on-prem or edge inference where power budgets are tighter. SIers who can architect around memory efficiency, rather than assuming abundant HBM in the cloud, will hold a real advantage as component costs stay elevated through the buildout cycle.