A developer rebuilt an old proof-of-concept as a voice-driven murder mystery, interrogating AI suspects through OpenAI's realtime speech-to-speech model over WebRTC, with a separate judge model scoring whether a player's accusation actually cited the required evidence. Small project, meaningful signal: the voice-agent stack that felt premature two years ago now works end to end.

The architecture is the story, not the game. Three patterns stand out for anyone building agents. First, real-time speech-to-speech over WebRTC has collapsed the latency that made voice interfaces feel robotic; conversation, not command-and-response, is now the baseline. Second, tool calls are being used to freeze fluid dialogue into structured, auditable state, capturing who was accused and what evidence was named. Third, an LLM-as-judge grades the interaction, distinguishing genuine reasoning from vague fishing. That triad, live voice, structured capture, automated evaluation, is precisely the loop enterprises need for coaching, compliance, and quality scoring.

The binding constraint is cost, not intelligence. The builder gated access behind authentication and a 30-minute timer specifically to avoid runaway spend. That is the honest state of realtime voice in 2025: technically ready, economically rationed. Until per-minute pricing falls by an order of magnitude, always-on voice agents remain a premium tier, and the winning products will be those that meter conversation ruthlessly or reserve voice for high-value moments.

For Japan, this maps directly onto two of the country's most acute pressures. Contact centers, elderly care, and municipal services are chronically short-staffed, and voice is the interface most Japanese users already trust. The natural evolution for domestic RPA vendors and SIers is from screen-scraping bots toward voice front-ends that talk to citizens and customers, then hand structured data to back-office automation. The judge-model pattern is especially relevant for regulated sectors like finance and insurance, where every spoken interaction must be logged, scored, and defensible.

Two cautions temper the opportunity. Japanese-language speech-to-speech quality still trails English, so honmono deployment demands real accent and dialect testing rather than demo-day optimism. And the cost wall hits harder here, where clients expect fixed-price contracts; SIers that bid flat fees on token-metered voice services will absorb the volatility. Expect the near-term winners to sell voice as a scoped, high-margin module, not an unlimited utility.