Summary
- Recently disclosed documents stem from the New York Times' copyright lawsuit against OpenAI and Microsoft.
- A memo from Microsoft in 2023 cautioned that the public might view AI models "hoovering up" their work as an unprecedented form of theft.
- CEO Satya Nadella stated that any paywalled content should be licensed by those wishing to utilize it.
Internal discussions among Microsoft staff raised the question of whether OpenAI's utilization of news articles constitutes "the largest theft of labor in human history," potentially leading to a detrimental cycle that could undermine the AI models they are developing. This information was revealed in court documents made public on Thursday, as reported by the New York Times.
These documents are part of a lawsuit filed by the Times against both companies in late 2023, which has since seen the involvement of eleven additional publishers. OpenAI has consistently contested the allegations and has been compelled to maintain 20 million ChatGPT conversation logs as part of the legal proceedings. Judge Sidney Stein of the Southern District of New York is currently considering motions for summary judgment, with documents being released as he deliberates.
One internal Microsoft memo from 2023 stated, "Millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions." The memo further described large AI models as products that "destroy their supply chain."
Microsoft clarified that these memos were authored by Brent Hecht, an applied science director and Northwestern University affiliate, and do not reflect the company's official stance. Hecht was not in a decision-making role and was tasked with offering diverse and contrasting viewpoints, according to the company's filings.
“Creating a Mess”
Satya Nadella testified that any content behind a paywall should be licensed by those who intend to use it, adding that had he known OpenAI was using paywalled material, he would have insisted on retraining its models. A spokesperson noted that he was addressing general principles regarding how information is accessed and consumed.
At OpenAI, a staff member informed president Greg Brockman about a method to circumvent the NYT paywall, to which Brockman responded, "ah nice."
Myriad: Who will IPO first, OpenAI or Anthropic? Make your prediction.Nick Turley, who led the ChatGPT team, remarked in June 2023 that AI poses an "existential threat" to publishers, and noted in February 2024 that AI products will increasingly replace traditional offerings as they improve. He also stated that AI products are fundamentally substitutive.
An OpenAI engineer commented in February 2023 that "no matter how prominently we show the links, users won't click," contradicting the idea that chatbots redirect traffic to publishers.
In a 2020 memo directed to Brockman and Sam Altman, former policy director Jack Clark cautioned that the company was creating systems that could replace the labor of individuals who contribute to societal culture, warning that it would symbolize Silicon Valley's careless encroachment into various life aspects, leaving chaos in its wake. Clark later co-founded Anthropic, which deferred comment requests to OpenAI.
Both Microsoft and OpenAI maintain that their training methods fall under fair use, arguing that they transform articles into new works rather than replacing the originals. Steven Lieberman, representing the New York Daily News and several other publications, stated, "The world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior."
The Times, a plaintiff in the lawsuit, declined to comment to its own reporters, while OpenAI did not respond to inquiries.
