A survey of 170 enterprises finds two-thirds now run AI in production, yet fewer than half can rigorously track what that compute costs. Performance and GPU availability have quietly outranked total cost of ownership in the buying decision.
The more interesting story is not the demotion itself but why it feels safe. Teams under production pressure treat cost as a problem to solve later, once the workload ships. The trouble is that AI infrastructure does not behave like traditional cloud: idle GPU capacity burns money continuously, and the survey's finding that most self-operated GPUs run at half utilization or less means a large share of enterprise AI spend is producing nothing. When you cannot see unit economics, you cannot renegotiate, right-size, or route workloads intelligently. The appetite for AI-specialized clouds as the next investment target, despite almost no one using them today, suggests procurement is being driven by narrative and peer pressure rather than measured return. This is how a capability boom turns into a margin problem two budget cycles from now.
Globally, the winners will be the platforms and tools that make inference economics legible. Expect a fast-maturing market for AI FinOps, GPU scheduling, and cost-attribution tooling, and a hard reckoning at firms that scaled first and instrumented never. Reliability as the top success metric is defensible, but reliability without cost visibility is just expensive uptime.
For Japan, this is both a warning and an opening. Japanese enterprises are traditionally disciplined on capex, but AI adoption anxiety can override that instinct, and many lack the internal MLOps depth to measure token-level or GPU-level cost on their own. That gap is precisely where SIers should move. Rather than reselling GPU capacity or model access as a pass-through, the durable margin lies in becoming the accountability layer: workload placement, utilization optimization, and transparent cost governance built into managed AI services. An SIer that can prove a client's cost per inference and cut idle capacity delivers something the hyperscalers do not package neatly.
RPA and automation vendors face a parallel shift. As Japanese firms bolt generative AI onto existing RPA estates, the per-run cost of an LLM-backed task is far less predictable than a deterministic bot, and finance teams will demand visibility they currently don't have. Local development teams should treat cost instrumentation as a first-class requirement now, not a cleanup task, because retrofitting observability onto a live AI stack under production pressure is exactly the trap this data describes.