Cactus released Needle2, a 14MB agentic LLM that runs a full session in 28MB of RAM at 45M parameters and 2-bit compression, hitting 500 tokens/sec on a Raspberry Pi 5 and 300-700 on sub-$200 phones.
The timing is the story. As Wall Street lines up hundreds of billions for centralized AI infrastructure, Needle advances the opposite thesis: most of the world's 21 billion connected devices will never touch a datacenter, and the 1.5 billion Macs and PCs that dominate "edge AI" discourse are a rounding error against budget phones, microcontrollers, wearables, and small robots. The economic argument is power, not accuracy. On an always-on assistant, every MFLOP is milliwatt-hours, and a model that spends 70 versus 87-164 changes what battery-constrained hardware can host at all. The strategic bet underneath is narrower and sharper than it looks: if consumer intelligence is framed as typed functions, the only hard task is mapping a messy sentence onto the right call, which needs no world knowledge and no open-ended prose. That is why 45M parameters can trade wins with larger small models. It also means this is not competing with frontier reasoning; it is competing with the cloud round-trip.
For Japan, the fit is unusually clean. This is a robotics, factory-automation, and consumer-electronics economy where data-residency rules, plant-floor latency, and offline reliability make cloud dependency a liability rather than a convenience. On-device tool-calling lets manufacturers and device makers embed intelligence without egress costs or compliance exposure in finance and healthcare deployments.
For SIers, the interesting lever is customization economics. A model that fine-tunes on a Mac or PC in minutes to hours, from a handful of samples, turns bespoke vertical intelligence into a repeatable engagement rather than a research project. Combined with a confidence-score threshold that acts locally and escalates to a bigger model only when unsure, it maps neatly onto Japanese enterprise risk tolerance: cheap, contained, auditable, with a defined fallback.
RPA vendors should read this as pressure. Much of Japanese RPA automates the same problem Needle targets, translating unstructured input into structured actions, but does so with brittle rules and, increasingly, cloud LLM calls. An on-device model that returns structured output against a passed-in schema is a credible substitute for the perception layer, potentially collapsing cost and privacy concerns at once. The open question is durability: 2-bit, 45M-parameter models leave little margin for edge cases, so the near-term win is bounded, schema-driven tasks, not general autonomy.