Alibaba's Qwen released Qwen-UI-Agent, a foundation model that operates real phones, desktops, and browsers by reading screens and mimicking clicks, typing, and swipes, pausing for confirmation before sensitive steps like payments.
The strategic shift here is the move from API-mediated automation to direct visual screen control. For a decade, automating software required either an integration point or a brittle scripted workflow tied to fixed coordinates. A model that perceives any interface the way a human does collapses that dependency. It puts Alibaba on the same battleground as Anthropic's computer-use tooling and other agentic efforts, and reinforces a pattern worth watching: Chinese labs are shipping capable, openly framed models fast, which compresses pricing and narrows the differentiation window for every vendor selling automation. Benchmark scores matter less than the trajectory — screen-native agents are becoming a category, not a demo.
The caveat is reliability. Visual agents remain probabilistic, and the gap between a strong benchmark and a workflow you trust with production payments is wide. The confirmation gate before sensitive actions is an admission that autonomy still needs guardrails, and that governance layer is where enterprise value and risk will concentrate.
For Japan, this lands directly on the RPA economy. Domestic adoption of tools like WinActor and UiPath runs deep, and a large share of SIer revenue comes from building and maintaining screen-scraping bots that break whenever a UI changes. A model that adapts to interface shifts on its own erodes the maintenance-contract logic those businesses rely on. SIers that treat this as a threat will defend legacy billing; those that treat it as leverage will reposition around agent orchestration, exception handling, audit trails, and the human-in-the-loop controls Japanese enterprises demand.
There is also a genuine opportunity. Japan's biggest automation blocker is the mass of legacy systems with no APIs — mainframe green screens, aging web portals, packaged software no vendor will reopen. Screen-native agents can, in principle, drive exactly these systems. The near-term winners will be teams that pair the technology with rigorous scoping, security review, and confirmation workflows, rather than those chasing full autonomy. Executives should fund controlled pilots on low-risk internal processes now, and pointedly avoid pointing an unsupervised agent at anything that moves money.