For years, the future of artificial intelligence was represented through immaterial images: cloud computing, neural networks, weightless streams of data. Then the pallets appeared. Millions of books were purchased, transported to warehouses, stripped of their spines, scanned page by page and finally sent for recycling. The scene retains something both industrial and archaic, because the most advanced machines on the planet still depend on objects made of paper, glue and ink. This is not merely a technological curiosity or a sudden sentimental return to matter. The printed book offers a reserve of human language relatively protected from the synthetic contamination spreading across the web. Long regarded as the medium of an age in decline, paper is now being treated as a raw material.
The Anthropic case made this reversal visible. The company purchased millions of volumes and produced digital copies of them through a destructive process; at the same time, it had assembled a library of more than seven million works downloaded from pirate archives. In June 2025, federal judge William Alsup distinguished between the two operations. Training models and replacing an individually purchased book with an internal digital copy were deemed transformative uses, whereas the creation of a permanent library through pirated copies was not recognised as fair use.
That distinction was followed, on 20 July 2026, by the approval of a $1.5 billion settlement concerning 482,460 works, with compensation of approximately $3,000 per title. The affair therefore cannot be reduced to a straightforward exoneration of the artificial intelligence industry. It has instead delineated a still-unstable legal frontier, within which the technological function of a work does not erase the question of how it was acquired. The route by which content entered the system consequently remains legally relevant, even when its subsequent use is regarded as transformative.
Behind the copyright dispute lies a problem that the law alone cannot resolve. Artificial intelligence is encountering the material limits of genuinely human knowledge, which is neither infinite nor homogeneous, nor automatically renewable. The web is increasingly populated by texts, images and comments produced or reworked by generative systems; the distinction between an original document, an imitation, a synthesis and a duplication becomes less certain, even as the overall body of content continues outwardly to expand.
When new models learn indiscriminately from the outputs of previous models, less frequent information may disappear, diversity contracts and the representation of reality loses some of its peripheral regions. The phenomenon has been called model collapse, an effective expression, though perhaps too technical to convey all its implications. Linguistic quality deteriorates, and it becomes more difficult to encounter the rare case, the deviation, the nonconforming thought: everything that occupies a marginal position within statistical distributions yet still contributes to the construction of knowledge. The result may be a representation that is narrower and less responsive to exceptions, while continuing to display a convincing linguistic form.
For businesses, the problem takes on a less abstract form. In recent years, managerial attention has focused on computing power, models and applications, while data has often been treated as an abundant and interchangeable resource. A verified, documented and legally usable archive may instead acquire greater value than the algorithm that interrogates it. Provenance affects product quality, legal risk, reputation and the ability to defend a competitive advantage. A company that feeds its systems with uncontrolled content may obtain rapid answers while simultaneously building decisions on opaque foundations. With well-classified proprietary data, even smaller models may produce more reliable results. The difference now has concrete industrial consequences. Shutterstock, for example, has established multi-year agreements to license images, video, music and metadata for training purposes, combining the availability of content with certainty over rights and compensation mechanisms for contributors. Reddit has presented its repository of more than one billion posts and sixteen billion comments as an advantage built on continuously updated human conversations, developing a dedicated licensing business. Quantity, structure, freshness and recognisable origin now enter the same economic calculation. Human data, long collected as a free by-product of digital life, is entering a market in which it can be selected, authorised and traded.
The issue extends beyond the boundaries of the individual company. Technological sovereignty depends on the availability of semiconductors, energy and data centres, but also on the linguistic and cultural reserves from which intelligent systems are built. A country that fails to organise and digitise its documentary heritage may find its collective experience filtered through models trained elsewhere, according to linguistic hierarchies and interpretative categories defined by others. The Institutional Books project has indicated a possible alternative by converting nearly one million public-domain volumes from Harvard’s collections into a documented corpus of approximately 242 billion tokens spanning more than 250 languages. From this perspective, libraries, archives and universities assume an infrastructural role, although the relationship between digitisation and the private appropriation of public knowledge remains unresolved.
At the corporate level, the issue returns in different terms. Data quality can no longer be entrusted exclusively to technical departments, because it shapes tools used in personnel selection, credit assessment, research, marketing and industrial planning. Boards of directors and senior management will need to understand the origin of their corpora, distinguish human data from synthetic data, establish preservation criteria and document authorisations. The person responsible for data therefore ends up safeguarding something more than an information resource and must ensure a degree of cognitive continuity within the organisation. Untraceable datasets and contaminated archives may remain invisible for a long time, like certain off-balance-sheet liabilities, until they produce an error, a dispute or a loss of trust.
The physical destruction of books retains a symbolic force that the legal distinction does not exhaust. A volume that is purchased and converted into a file does not necessarily disappear as content, but it ceases to exist as an object that can be accessed, transferred and preserved by a community. A library preserves works so that they may continue to be read by people who cannot be predicted; a dataset selects information so that it can be processed by a machine for a specific purpose. Both organise knowledge according to different responsibilities and time horizons. When private capital removes books from the market in order to destroy them, what appears efficient within a production process may become culturally irreversible, especially in the case of rare editions, minor translations or texts belonging to underrepresented languages.
The rush for books published before the explosion of generative AI nevertheless suggests something the industry may not have intended to declare. To produce the future, machines continue to need a past written without them, with its hesitations, contradictions and deviations. The decisive asset of the artificial economy may lie in the content that artificial intelligence is able to generate, but also in the human capacity to produce, recognise and preserve materials with an identifiable origin. Information is now abundant, whereas provenance remains a more limited resource. In an environment populated by copies without genealogy, a growing share of economic and cultural value will depend on the ability to recognise who produced something, under what conditions and under whose responsibility.

