Advocacy groups have asked the US Federal Trade Commission to examine how AI companies source physical books for model training, raising alarms about destructive scanning practices. The framing is emotive, but the strategic substance is dull and important: the industry is running out of cheap, high-quality text.

The global signal here is a shift from an era of scraping toward an era of acquisition. Web data is picked over, increasingly polluted with synthetic output, and legally contested. Books remain one of the last reservoirs of dense, edited, long-form prose, which is exactly what improves reasoning and reduces hallucination. That scarcity is why frontier labs are quietly moving to buy, license, or physically digitize catalogs. An FTC inquiry would not stop this, but it would push data provenance from a back-office concern into a governance and disclosure requirement, much as privacy audits became standard after early data scandals.

Expect three downstream effects. First, a widening cost gap between labs that can afford clean licensed corpora and everyone else, reinforcing incumbency. Second, a compliance market for data lineage, consent tracking, and rights clearance. Third, mounting pressure on the "train first, apologize later" posture that defined the last three years.

For Japan, the implications cut in a favorable and a cautionary direction. Japan's copyright law is unusually permissive on text and data mining for machine learning, which has been sold as a competitive edge for domestic model builders. But permissive input rules do not shield Japanese firms whose models are deployed globally; a US regulatory standard on data sourcing effectively becomes an export requirement. Enterprises building on domestic LLMs should assume their data supply chains will be audited by overseas customers, not just local regulators.

For SIers and enterprise IT teams, this reframes a procurement question. When integrating third-party models into client systems, the unexamined risk is not model quality but training-data provenance liability flowing downstream to the integrator and the end customer. Japanese SIers should start treating data lineage documentation as a contractual deliverable, and RPA and automation vendors handling scanned documents should separate clearly licensed enterprise content from ambiguously sourced corpora. The firms that build provenance discipline now will win the enterprise deals that compliance-sensitive buyers are about to start demanding.