Researchers from ETH Zurich, the AI safety group MATS and Anthropic described a four-stage process that turns publicly posted text into a probable real-world identity. First, an AI model ingests a user’s contributions and produces a concise profile covering possible location, occupation, hobbies and other quirks. Next, the profile is transformed into a numerical embedding that can be rapidly compared against millions of public profiles. A more powerful model then evaluates the top matches, reasoning about which details line up best. Finally, the system assigns a confidence score and only outputs a guess when it deems the match reliable.

Test results and observed limits

The team evaluated the approach on two data sets. In the first, 338 members of the Hacker News community who had voluntarily linked their LinkedIn pages were stripped of identifying information. The AI correctly matched 226 of them, representing a 67% success rate, while roughly one in ten of the generated guesses were incorrect. A separate experiment using interview transcripts from 125 scientists at Anthropic yielded at least nine correct identifications based purely on descriptive language.

When the candidate pool was expanded to 89,000 possible profiles, the strongest configuration still achieved about a 50% recall while maintaining 90% precision – meaning that when the system offered a match, it was right nine times out of ten, but it only found half of the true matches present in the set.

Cost efficiency and accessibility

Running a single identification cycle reportedly costs between $1 and $4 in AI subscription fees, as the pipeline relies on publicly available web search, embedding generation and reasoning models similar to those powering mainstream chatbots. No data breaches, hacking tools or privileged access are required; the process merely chains together benign-looking operations that are individually harmless.

Implications for privacy and online behavior

The findings suggest that practical anonymity on platforms that rely on user-generated text is increasingly fragile. While the experiments focused on accounts that already disclosed a link to a professional profile, the authors note that the technique could be adapted to any pseudonymous presence where sufficient textual material exists. For individuals who depend on anonymity for activism, personal safety, or to protect financial activities such as crypto transactions, the aggregation of seemingly innocuous details—like a hometown reference or a pet’s name—can create a distinctive fingerprint that AI can exploit.

Why it matters

The study highlights a new privacy risk that does not depend on external data leaks but on the publicly shared content itself. As large language models become more capable and affordable, the barrier to conducting large-scale deanonymization drops dramatically, prompting a need to reassess threat models for online anonymity and consider protective practices that limit the amount of personally identifying information tied to any single pseudonym.