The recent decision by the agent of Nigerian author Jerry Falade to withdraw his novel from consideration has raised concerns about the implications of artificial intelligence (AI) in content creation. The manuscript had attracted interest from 14 publishers, with the highest bid exceeding $2 million, but suspicions regarding the use of AI led to the termination of the deal on July 30, 2026.
On August 2, 2026, the European Union implemented a law aimed at addressing such disputes through the EU AI Act, specifically Article 50, which mandates the labeling of synthetic content, marking a significant step in global regulation.
ForkLog explored whether those who utilize AI for content creation should be worried, while others may welcome the changes, particularly those fatigued by unexpected AI-generated outputs on platforms like YouTube.
Global Regulation of AI Content
Since autumn 2025, China has been operating under its own set of AI content labeling regulations. Major platforms such as WeChat, Douyin, Weibo, and RedNote have quickly implemented automated tagging systems, complaint tools, and user alerts, positioning China ahead of other jurisdictions in this area.
On June 9, 2026, New York State introduced the Synthetic Performer Disclosure Law (S.8420-A), which regulates the use of AI-generated actors and digital replicas in advertising and media. Violations can result in fines ranging from $1,000 for the first offense to $5,000 for subsequent infractions.
Moreover, the key transparency provisions of Article 50 of the EU AI Act came into effect on August 2, 2026, establishing several critical requirements:
- Machine-readable labeling and digital watermarks: Providers of generative AI systems must label any created or altered audio, video, photo, and text content in a machine-readable format, allowing for its identification as synthetic (including invisible watermarks and signed metadata);
- Disclosure of interaction: Users must be explicitly informed when they are interacting with an AI system, such as chatbots, unless it is evident from the context. Additional mandatory notification requirements are also introduced for emotion recognition or biometric categorization systems;
- Labeling of deepfakes: Any generated materials representing real people or significant public events must have a clearly visible label.
On the same day, California's AI Transparency Act (SB 942) also took effect, with its launch postponed from January 1 to August 2 specifically to give developers time to adapt.
Regulators across various countries have moved in parallel to respond to the same pressure: the increasing volume of generated content has outpaced any voluntary industry restrictions.
Some market participants had already prepared for the deadline. TikTok integrated "smart" metadata as early as January 2025, tagging over 1.3 billion video materials since then. In the latter half of that year, the platform removed 51,618 videos lacking the necessary synthetic media labeling, a 340% increase compared to the same period in 2024.
The law did not introduce new censorship rules for AI content but rather formalized obligations that some major platforms had already been voluntarily adhering to in response to a global audience.
Three Levels of Protection
All laws regarding AI content share a common technical requirement: information about its origin must be retained even after copying, compression, or publication on another platform. The industry employs three approaches to achieve this:
- C2PA: This standard allows for the attachment of tamper-proof information regarding the origin of an image, video, or other file, detailing who created it, what tools were used for processing, and what changes were made. Content Credentials, based on this standard, serve as a digital "label" visible to users. The C2PA was developed by a coalition formed in February 2021 including Adobe, Arm, BBC, Intel, Microsoft, and Truepic, and by December 2025, it had expanded to over 200 organizations, including Google, OpenAI, Meta, Sony, TikTok, and Amazon.
- Invisible watermarks: These are embedded at the pixel level of images or sound waves. In August 2023, Google DeepMind introduced SynthID, a tool that adds labels to content generated by its Imagen model. In 2024, Meta FAIR's research team collaborated with the French institute INRIA to develop AudioSeal, capable of marking speech with precision up to 1/16,000 of a second, which remains intact even in short fragments cut from longer recordings.
- Statistical text marking: This approach embeds watermarks directly into the text generation process by altering token selection probabilities, subtly shifting the model toward certain options over others. The finished version, known as SynthID-Text, was unveiled by Google DeepMind on October 23, 2024, integrated into the chatbot Gemini.
While C2PA is the most mathematically robust solution, it has vulnerabilities; files can be edited using software that does not support this standard, potentially removing the manifest container.
Midjourney notably does not embed C2PA markers, opting instead for a simple text tag, trainedAlgorithmicMedia, in standard metadata that can be easily removed.
The first two methods effectively label images, audio, and video, but they do not guarantee that the label will survive file processing. The statistical marking method remains less reliable, as it can identify signs of machine generation but does not provide conclusive proof that the text was AI-generated.
The Neural Network Debut
Just days before the new transparency requirements of the EU AI Act took effect, a scandal highlighted the challenges of determining text origin. The controversy centered around Jerry Falade, whose debut novel attracted bids from 14 publishers, with Minotaur, an imprint of Macmillan, winning with an offer exceeding $2 million.
Falade's agent, Mark Gerald, founder of the Brooklyn-based boutique literary management and production agency Europa Content, had been negotiating for months and told The Bookseller:
“The book was stunningly good — everyone who read it fell in love with it.”
Following the auction results, film and TV studios aggressively pursued the rights for adaptation.
Doubts about authorship first emerged from an editor during the bidding process. Europa Content decided against running the manuscript through AI detectors, as Gerald noted that such tools are notorious for false positives, and Falade's explanations were deemed convincing.
On July 29, a final meeting with the author led to what Gerald described as "some details of his story changing." The following day, Europa Content terminated the contract. Notably, the letter to publishers did not assert a proven fact; it merely stated that they could no longer "confirm the text's origin" and that the situation "evoked a sense of guilt."
Falade completely denied any allegations of "AI plagiarism," asserting that he had used neural networks solely for fact-checking and text editing. In a comment to The Guardian, he accused the publishing industry of discrimination:
“Three Black authors received major publishing deals this year, and all of them had their agreements subsequently terminated or suspended due to AI-related suspicions.”
One such case involved Mia Ballard's horror novel "Shy Girl," which Hachette canceled for release in the U.S. after rumors of AI involvement in its writing, although Ballard herself denied the accusations. Another example is the viral novel "Daggermouth" by Halla Mikaela Wolf, which was deemed 60% AI-generated by Pangram Labs.
Source: Jerry Falade's Instagram page.In all three cases, there was no technical confirmation of AI usage. The decisions made by publishers and public reactions were based on probabilistic assessments from automated detectors, stylistic suspicions, and reputational risks.
Content labeling was intended to protect authenticity — to delineate where human effort ends and machine work begins. In the case of text, the system operates in reverse: where formal certainty is fundamentally absent, the vacuum is filled not by a presumption of innocence, but rather by the publisher's reputational panic, which prefers to retreat rather than risk being on the wrong side of a scandal.
The Open Source Blind Spot
Similar to the U.S. attempts to ban Chinese large language models (LLMs) with open weights following the Kimi K3 incident, the challenges surrounding the labeling of generated content are more complex than they appear at first glance. Open Source does not directly oppose regulation but significantly complicates it.
A significant portion of generative models physically lies beyond the reach of regulators. General-purpose neural networks like Qwen, Llama, GLM, and DeepSeek, as well as specialized models for image generation such as Stable Diffusion and FLUX.1 can be downloaded, run locally, and modified arbitrarily, including disabling built-in labeling.
Even where protection functions properly, it fails under routine manipulations: screenshots, format changes, and re-uploading to another platform. It is even more challenging with text.
On August 11, 2026, Anthropic announced the integration of watermarks in Claude. The text does not contain hidden markers, extra spaces, or special metadata. The algorithm works through controlled probability shifts when selecting synonyms. For instance, when the model chooses between the words "overcast" and "gray," a special pseudo-random key nudges it toward one specific option.
However, the marker becomes unreadable with deep paraphrasing, translation into another language, significant shortening, or transfer to another, especially open, model. Furthermore, reliable statistical recognition requires a sufficient volume of text; short phrases of a couple of words remain undetectable.
This leads to another non-obvious trend: users accustomed to seeing labels on synthetic content begin to interpret their absence as proof of authenticity, even though the label may have vanished during transmission — whether accidentally or not.
Realities
Systems currently operating in the market must align their labeling practices with regulatory requirements by December 2, 2026 — four months later than the original date, specifically to allow developers time to adapt existing products.
There are exceptions: artistic, satirical, and clearly creative works are not subject to this requirement. Memes, parodies, and arthouse short films are formally outside the law's scope. The regulator recognizes that insisting on a watermark for an obvious caricature is pointless.
YouTube illustrates how this plays out in practice. If a creator does not disclose their use of AI, but the platform detects pronounced signs of photorealistic synthesis, a label is automatically added. This can be disputed through YouTube Studio, except in two cases where the label becomes permanent: if the video was created using built-in tools like Veo or if the file already contains C2PA metadata.
Platforms are motivated to navigate the legal nuances for three primary reasons:
- Regulatory: To avoid fines and other risks in jurisdictions with mandatory norms;
- Reputational: To avoid being at the center of a scandal involving deepfakes or electoral misinformation, as has occurred with several major services;
- Commercial: Advertisers demand to know precisely alongside what content their budgets are placed, and synthetic videos without identifying markers do not meet these criteria.
A similar dynamic is currently observable in music streaming.
In April 2026, French service Deezer reported that AI tracks now account for 44% of all new daily uploads. Approximately 75,000 such compositions are added daily to the platform — over 2 million per month. However, a survey of 9,000 participants across eight countries revealed that 97% of users could not determine which compositions were created by humans and which by neural networks.
On August 11, 2026, Spotify announced the launch of the AI Persona badge: starting mid-September, profiles of artists whose personas appear convincingly generated will be marked in searches, on artist pages, and in playlists, while their tracks will default to disappearing from recommendations.
This initiative was prompted by the case of Velvet Sundown — a fictional band that garnered hundreds of thousands of listens from 1970s rock enthusiasts until its own social media confirmed that it was a fully synthetic project all along.
On September 3, another side of the same business will feel pressure: the music generator Suno is introducing download limits on tracks: no more than seven for free and up to 60 per month. The formal reason is a global settlement with Warner Music Group, requiring the company to limit mass file exports to combat slop.
Ultimately, Falade proved nothing — neither his innocence nor the claim that the allegations stemmed from bias. The system that regulators have built over several years across three continents has mandated the industry to confirm the origins of content: while this task is more or less resolved for images, sound, and video, it remains unaddressed for text.
For those using artificial intelligence in their daily work — from copywriters to programmers and editors — the main threat lies not in fines from regulators but rather in a presumption of guilt and corporate paranoia. If an employer, client, or internal platform algorithm suspects a person of "synthetic" work, they will have to prove authenticity independently, as there is currently no reliable technical alibi for text.
Labeling will complicate the lives of legitimate businesses and force giants like YouTube and Spotify to filter output stringently, but open AI models and simple metadata removal leave a wide avenue for spammers.