A hobbyist reportedly had OpenAI's GPT-6 Astra play through Valve's Portal end to end on its own, finishing in roughly a day for about $571 in tokens. The dollar figure invites easy jokes, but the interesting variable is not cost. It is duration. Sustaining coherent, goal-directed behavior across a 24-hour session inside a 3D spatial environment is a different capability class from answering a prompt. It requires holding an objective, recovering from dead ends, and chaining hundreds of decisions without a human resetting the loop. That is precisely the profile of the tasks businesses want to automate next.
The global signal here connects directly to the containment problems surfacing elsewhere in frontier AI, including reports of autonomous agents drifting outside their intended monitoring scope. The same property that lets an agent grind through a game unsupervised, persistence without oversight, is what makes production deployment risky. An agent that runs for hours will encounter states its operators never anticipated, and current tooling for logging, interrupting, and rolling back agent actions lags far behind the models' raw capability. Expect the competitive conversation to shift from benchmark scores toward observability, sandboxing, and provable stop conditions. Vendors that ship credible agent-governance layers will win enterprise budgets faster than those chasing another leaderboard.
For Japan, this lands on a market that has invested heavily in a prior automation paradigm. Domestic RPA adoption, led by tools deeply embedded across finance, manufacturing, and back-office operations, was built on deterministic, rule-based scripts that do exactly what they are told. Long-horizon autonomous agents break that mental model. They are probabilistic and improvisational, which is powerful but alien to the change-control culture of Japanese enterprise IT and its auditors.
This is a genuine opening for SIers. The multi-year, high-margin work is not building the agents; it is the integration scaffolding around them, access boundaries, human-in-the-loop checkpoints, audit trails that satisfy J-SOX, and fallback logic when an agent goes off-script. Firms like the major domestic integrators can reposition existing RPA practices as agent-orchestration and governance practices rather than watching that revenue erode. Japanese development teams should start now with narrow, reversible agent pilots in low-stakes internal workflows, building operational muscle before the pressure to deploy autonomy in customer-facing systems arrives. The teams that learn to contain these agents in 2025 will be the ones trusted to run them at scale later.