ISBNdb, a service housing metadata for 113 million publications, has introduced a controversial offering aimed at AI firms: bulk purchases of books, up to a million copies, for digitization and eventual disposal. Following media coverage of this scheme, the company quickly removed the service description from its website, asserting that no deals had been finalized. According to representatives, the initiative was merely a "market demand check."
In another market segment, platforms ActID and New Claw are willing to pay individuals between $15 and $700 for the rights to use their likenesses as characters in AI-generated series and advertisements.
ForkLog investigates the reasons behind the need for labs to acquire "organic" data, the phenomenon of "model collapse," and what other human elements machines might require.
Model Collapse
The term "model collapse" refers to the degradation of AI systems as they learn from data generated by other neural networks rather than from human sources. With each new training cycle, the outcomes become increasingly average and predictable, gradually detaching from the original reality that large language models (LLMs) were meant to reflect.
The concept was first formalized by an international research team led by Ilya Shumailov, which published a groundbreaking paper titled "The Curse of Recursion: Training on Generated Data Causes Models to Forget" in May 2023.
Shumailov and his colleagues mathematically demonstrated that regular incorporation of synthetic data leads models to first struggle with accurately reproducing information about rare events and phenomena, as well as subtle patterns, ultimately resulting in degradation: diversity of responses declines, and errors accumulate.
The scale of the issue is underscored by statistics. In April 2025, researchers from the web analytics company Ahrefs examined 900,000 newly indexed pages and articles on Google. They found that over 74% of these materials showed signs of auto-generation.
According to an estimate by the independent research institute Epoch AI, the entire pool of text suitable for training may be exhausted between 2026 and 2032.
This scenario implies a rapidly diminishing reservoir of sources free from machine-generated text for future generations of models. Consequently, books published before 2022 are considered particularly valuable, as they are among the few sources that are likely devoid of synthetic noise or intentional traps.
Some authors have resorted to a countermeasure known as "data poisoning," embedding symbols in the text that the model either fails to read or interprets as hidden commands.
Research from a collaborative team including the UK’s AI Safety Institute, Anthropic, the Alan Turing Institute, the University of Oxford, and ETH Zurich has preliminarily demonstrated that LLMs can be subtly trained on malicious patterns using a critically small amount of "poisoned" data—just 250 documents. Conventional alignment methods are inadequate to recognize and eliminate these vulnerabilities.
451 Degrees at Anthropic
The future of books is not clear for all who sell them.
In August 2024, three prominent American authors—Andrea Bartz, Charles Greibner, and Kirk Wallace Johnson—filed a class-action lawsuit against Anthropic, alleging copyright infringement in the training of the Claude family of models.
According to court documents, the company invested tens of millions into an "unofficial" project named Panama, led by former Google Books manager Tom Tervey. Anthropic allegedly purchased print book copies, cut their spines with hydraulic knives, and scanned them en masse for subsequent disposal, all quietly conducted without public announcements, akin to a standard logistical operation.
In June 2025, Judge William Alsup ruled that the digitization of legally acquired print books and the subsequent use of digital copies for training language models falls under the doctrine of fair use.
Now, it is sufficient to purchase a physical book, retain the receipt, and verify the origin of the training data through documentation.
According to 404 Media, this practice has expanded beyond individual startups; data suppliers like ISBNdb have turned the acquisition and destruction of print books into a turnkey service.
For over two decades, ISBNdb has operated as the largest international metadata database for printed books, housing over 113 million titles. Traditionally, libraries, bookstores, and publishers have utilized this service.
On July 21, 2026, journalist Emanuel Maiberg uncovered a special promotional section on ISBNdb’s website: "Searching for Printed Books to Meet Your Dataset Needs for Your LLMs."
On this page, the company offered AI labs wholesale brokerage services, including:
- facilitating the purchase of printed books released before 2022 in bulk, ranging from 1,000 to 1 million copies per order;
- assurances that the books are "clean" from AI noise and data poisoning algorithms;
- legal protection through strict non-disclosure agreement terms and the legality of acquiring books from the secondary market.
The platform's management explicitly stated:
“The perception problem is real. ‘AI company destroyed two million books’ — that’s not the headline that will gain sympathy from the public.”
As part of their investigation, 404 Media surveyed independent booksellers to ascertain whether this process had begun. They confirmed that since April 2026, they had witnessed abnormal purchases. One used book dealer noted that his sales surged from twenty copies to several hundred in a week.
Purchasers showed little interest in genre or value—the primary criterion was the presence of an ISBN code. They acquired even grossly overpriced or rare foreign editions, completely disregarding the cost.
However, beyond the financial gain, the excitement of the frenzy was limited: the same seller noted that unique editions were being lost irretrievably in this scheme—the last remaining copy might end up under a scanner rather than on someone’s bookshelf.
After the investigation was published, the story spread rapidly online.
On July 28, ISBNdb removed the promotional page from its site. Two days later, the management released a statement changing their rhetoric:
Source: ISBNdb.“Facts: ISBNdb has never purchased, scanned, or sold books—for AI training or anything else. We do not train AI models. This page was merely a test of market demand; such a service was never launched.”
In a follow-up article by 404 Media, journalist Samantha Cole responded to the events, recalling how a company that offered AI giants protection from reputational scandals was itself caught attempting to monetize mass book scanning, prompting a hasty removal of traces of the initiative following the exposure.
The disparity between the cost of raw materials and the finished product in this deal is stark. A book that is destroyed right after scanning can be valued at $2 to $5 on the secondary market, while a model trained on millions of such copies is valued in the billions.
Need More
In the past three years, the Internet Court of Guangzhou has handled approximately 700 cases related to the unauthorized use of individuals' likenesses through AI. In March 2026, a Beijing court issued an unprecedented ruling: the use of someone's likeness in AI-generated deepfakes without permission is illegal, even if the image has been modified.
This judicial trend reflects the rapid growth of the AI content market. Journalists Kinlin Lo and Viola Zhou detailed how an industry of AI mini-dramas has emerged in China. Their data indicates that over 95% of the 128,000 such films released in the first three months of 2026 were produced using artificial intelligence.
Short mobile series in China boast a multibillion-dollar audience. The demand for recognizable faces for such projects has given rise to a new type of deal: likeness licensing.
Platforms like ActID and New Claw function as typical marketplaces: companies pay individuals between $15 and $700 for licensing their images. Users upload photos taken in specialized studios, after which producers select faces based on gender, age, and categories such as "neighbor girl," "brutal," or "supermodel," filtered by project genre.
Source: Rest of World.Long Liu, head of operations at New Claw, explained the platform's logic simply to Rest of World:
“Selling rights to use their photos allows them to earn while continuing their offline careers.”
New Claw entered this market after its previous production direction suffered significant losses. Liu stated that since the end of the previous year, traditional drama production has decreased, advertising budgets have shrunk, and even luxury brands have begun transitioning to AI content.
The other platform, ActID, founded in Shenzhen in March 2026, has registered around 800 users, with approximately 300 agreeing to license images for AI production. Prices range from $15 to $74 per episode, with a 10% commission for the platform.
ActID’s marketing director, Camilla Yan, noted in a media comment that AI series do not require exceptional acting skills. She explained the selection logic:
“Producers need attractive faces for lead roles, but they also require distinct characters, such as elderly people.”
The platforms aim to protect individuals from the kind of issues that arose with the AI drama platform Hongguo, owned by ByteDance. In April 2026, two influencers accused the company of stealing and altering their faces. A series with 40 million views had to be taken down.
Since the beginning of 2026, the Chinese tech giant has removed over 85,000 videos that improperly used others' faces and voices.
Beijing lawyer Ile Deng commented to Rest of World that the market for likeness licensing is a positive step towards establishing clear commercial and legal rules. However, Deng believes that platforms cannot entirely eliminate risks: once biometric data is in circulation, it may lose long-term control by its owners. In particular, there is no guarantee against unauthorized likeness replacement, illegal data collection, or the reuse of uploaded images for AI training.
“Many agencies still pay individuals mere dozens of dollars for data collection. However, if you closely examine contracts, the licensing terms are often so vague that it’s impossible to understand who and for what purpose will ultimately use that likeness,” the lawyer explained.
In practice, this control has proven fragile.
Individuals who sold licenses for their likenesses later found that their digital clones were used in misleading advertising or in materials that could be classified as political propaganda—contexts that the original consent hardly anticipated.
Two notable cases involved British citizens who sold rights to their digital avatars to the startup Synthesia:
- In 2021, Dan Dewhurst signed an official contract with the AI platform, expecting that his digital copy would be used for educational purposes. In 2024, the actor learned from a Guardian investigation that his "clone" had become the main "host" on a fake news YouTube channel called House of News. This project was backed by government-affiliated structures in Venezuela that were engaged in spreading disinformation. The avatar, speaking in English, was promoting a "tourism boom and economic prosperity" while masking the real economic crisis;
- A model and creative director from London, Mark Torres, provided data to create a digital twin, undergoing a full motion-capture session. However, his likeness was later used without permission in propaganda videos supporting Ibrahim Traoré—a dictator in Burkina Faso who seized power in a military coup.
https://youtu.be/TqNXqbTUpQ8
Both cases have set precedents for the British actors' union Equity, which is now pushing for legislative bans on the transfer of uncontrolled rights to digital likenesses without clear limits on usage context.
Don’t Forget the Receipt
Books and faces are not the only resources needed by AI companies.
A similar situation is developing in the voice data market. In June 2026, members of the SAG-AFTRA union ratified an agreement expanding protections against the use of synthetic voice copies without consent—following the same licensing principle as with likenesses, but within a collective agreement.
Robots require examples of natural human movements in real-world scenarios. Current physics simulations still do not accurately replicate how a hand grips a mug or how a foot finds footing on uneven ground, creating a situation for synthetic data for robots that mirrors the AI model collapse: lacking quality real examples, the system struggles with rare and complex scenarios. According to MIT, over $6 billion was invested in humanoid development in 2025, with companies like Scale AI and Encord hiring operators to record everyday movements on camera, while Tesla pays up to $48 an hour for people to repeat routine actions in motion capture suits to train its Optimus machines.
Setting aside the differences in medium—texts, faces, voices, movements—the underlying scheme remains the same across the board. Previously, such data was taken without consent: archives were scanned, photos scraped, and voices trained on publicly available sources. Now, payment is involved. However, the structure of the deal has not changed: the asset still permanently leaves the owner's control, but now with a receipt in hand.
Books and human likenesses may seem like distinct categories, but for the AI industry, they converge into one type of resource—data for training models. Companies are eager to purchase both, yet the market is structured in a way that offers no protections; once rights are transferred, individuals cannot fully control how their voice, face, or created content will be used in the future.
