Summary
- Researchers from ETH Zurich, MATS, and Anthropic developed an AI system that can associate anonymized accounts on Hacker News with actual LinkedIn profiles with 90% accuracy.
- This method relies solely on web searches, embeddings, and reasoning models like GPT-5.2, costing approximately $1 to $4 per target—no hacking required.
- The authors of the study have kept their code, prompts, and the identities they uncovered confidential, asserting that their research passed through ETH Zurich's ethics review board before it was published.
Many social media users are expressing alarm over a recent AI study suggesting that online anonymity may be effectively over. This research, which has resurfaced on platforms like X and Reddit, reveals that artificial intelligence can unmask individuals behind anonymous online accounts.
The research paper, titled "Large-scale online deanonymization with LLMs," was created by a team from ETH Zurich and the AI safety organization MATS, in collaboration with Nicholas Carlini from Anthropic, the company known for developing Claude.
The researchers assert that large language models, which underpin AI systems like Claude, ChatGPT, and Gemini, can analyze anonymous posts to deduce a person's real identity. This is achieved not through hacking but through regular web browsing, searching, and reasoning—tasks that these AI systems already perform daily.
Researchers have created an AI capable of revealing any "anonymous" Reddit account in under a minute for just $2.
ETH Zurich and Anthropic published a concerning paper indicating that the concept of "practical anonymity" on the internet is now obsolete.
They developed a fully autonomous AI pipeline that processes your… pic.twitter.com/XMeXifhDf3
— Superman (@thesupermannx) September 23, 2026
The paper states, “Our results show that the practical obscurity protecting pseudonymous users online no longer holds and that threat models for online privacy need to be reconsidered.”
Mechanics of the AI's Functionality
The researchers outlined a four-step process for the AI’s operation, termed Extract, Search, Reason, and Calibrate.
“Extract” involves inputting a person's posts into the AI, which then summarizes details about the individual, such as their potential location, occupation, hobbies, or even unique traits like their use of British spelling or a dog’s name mentioned briefly.
“Search” converts this summary into an embedding—a numerical representation that helps the AI quickly locate similar profiles from a vast pool of candidates.
“Reason” allows a more advanced model to analyze the top matches and determine the most accurate fit by cross-referencing specifics, akin to how a human investigator would, but at a much faster pace.
“Calibrate” enables the AI to assess its confidence level, ensuring it only makes guesses when it is quite certain and refrains from conjecturing otherwise.
In one experiment, the researchers examined 338 users from the predominantly pseudonymous Hacker News forum who had included links to their LinkedIn profiles in their bios. After removing identifiable information, they tasked an AI to determine their identities using only web searches. The AI successfully identified 226 users, achieving a 67% accuracy rate, with incorrect guesses occurring in about 10% of its attempts.
In another test involving transcripts from Anthropic's research interviews with 125 scientists, the AI accurately identified at least nine individuals based solely on their work descriptions.
Implications of the Findings
The cost of conducting these searches is a significant takeaway from the study, with expenses ranging from $1 to $4 in AI subscription fees.
Notably, this method does not require any data breach, hacking, or insider access; it merely leverages the web search and summarization capabilities inherent in standard chatbots. The researchers contend that this makes the process difficult to counteract, as there is no singular "deanonymize this person" function to disable, only a series of seemingly harmless tasks.
This isn't the first instance where minimal information has led to someone's identification. In 2008, researchers successfully identified users from Netflix's supposedly anonymous movie-rating dataset by correlating it with public reviews on IMDb. The key difference now is that AI can perform these associations on chaotic, unstructured text—such as jokes, comments, and casual mentions—rather than organized data, and it can do so autonomously.
Addressing Concerns
While the findings may seem alarming, it's important to consider the context. To gauge effectiveness, the researchers chose subjects whose identities they already knew, selecting accounts that had previously connected to LinkedIn or splitting a single person’s post history into two segments while concealing the link.
This represents a controlled scenario rather than definitive evidence that any random pseudonymous account can be easily exposed today.
Furthermore, the challenge of scale works against the study's more alarming statistics. The larger the pool of potential candidates, the more challenging it becomes to pinpoint a specific individual. Against a set of 89,000 candidates, even the most effective AI approach achieved only about 50% of correct matches at 90% precision—where precision refers to the accuracy of its guesses, and recall pertains to the total actual matches it identified.
The researchers also refrained from disclosing their tools, prompts, or any real identities they uncovered, and the study underwent an ethics review prior to publication. They are not providing a doxxing toolkit but rather documenting an existing capability in current AI models, irrespective of the paper's findings.
Significance Beyond Reddit Users
Individuals involved in cryptocurrency are acutely aware of this threat. Following the 2025 Coinbase data breach and other incidents, there has been a surge in doxxing incidents and kidnapping attempts, highlighting how quickly a leaked identity can lead to real-world consequences.
AI-enabled deanonymization presents a similar risk, operating based on publicly available information without requiring a data leak.
This discussion comes shortly after Anthropic reported that state-sponsored hackers utilized Claude to execute a significant cyber-espionage operation, underscoring that research on AI misuse is progressing faster than the safeguards intended to regulate it.
For those who use pseudonyms due to activism, abuse, sexual orientation, immigration status, or employment concerns, the accumulation of seemingly innocuous details, such as hometowns, employers, and even pet names, across years of online comments forms a distinctive fingerprint. While this was always a possibility, AI has made the process quicker and more cost-effective.
For those worried about potential doxxing, the solution lies not in panic but in changing habits: minimizing specific identifying details linked to any pseudonym and utilizing tools designed to prioritize privacy.
