Three sites made 215,128 “best software” pages for AI. Perplexity cites them
Perplexity AI's "search" models are heavily citing hundreds of thousands of low-quality, machine-generated "best software" pages, many from recently created sites sharing suspicious infrastructure. This deep dive exposes the alarming vulnerability of AI-powered information retrieval to large-scale content manipulation, echoing the worst aspects of traditional SEO. Hacker News is buzzing about the implications for AI models "eating their own tail" and the rapid degradation of internet information quality.
The Lowdown
Trellner Research exposes a troubling reality in AI-powered search, demonstrating how Perplexity AI's models predominantly cite low-quality, machine-generated content for "best software" recommendations, raising serious questions about information integrity.
- The research queried Perplexity's
sonarandsonar-promodels for "best software" in 380 categories, collecting 7,534 citations. - A staggering 59.8% of cited domains rank worse than #100,000 on Tranco, and 23.4% aren't in the top million, with many being very new domains.
guideflow.com, a vendor's marketing blog, surprisingly accounted for 194 citations across 96 categories, despite not being a review site.- Three specific sites (
wifitalents.com,worldmetrics.org,gitnux.org) were found to share infrastructure, register dates (Dec 2023-May 2024), and explicitly title their pages "Facts & Grounding Page"—a term for AI retrieval. - These three sites collectively generated 215,128 "best software" pages, often showing inconsistent rankings for the same category and containing unrendered template variables.
- Some recommended vendor homepages were unreachable, redirected to unrelated sites (e.g., online gambling, casinos), or were plainly incorrect.
- The study emphasizes its focus solely on Perplexity's retrieval layer, acknowledging limitations like specific prompt wording, proxy usage, and the dynamic nature of search indexes.
This investigation reveals a new front in the content wars, where AI is both the weapon and the victim, as models like Perplexity are shown to be highly susceptible to strategically engineered, machine-generated content, ultimately compromising the reliability of AI-sourced information.
The Gossip
AI-Generated Garbage: The Internet's New Trash Problem
Commenters lament the rapid proliferation of AI-generated content flooding the internet, making it increasingly difficult to find reliable information. Many note that searching for products or even basic facts has become nearly impossible due to the overwhelming volume of low-quality, AI-optimized spam. There's a shared concern that LLMs training on other LLM-generated output will create a feedback loop, amplifying misinformation and degrading the overall quality of online knowledge.
Perplexity's Plight: A Business Model on Shaky Ground
Several users, including former paying subscribers, express disillusionment with Perplexity's declining quality, questionable business practices, and overall viability. They recount experiences of diminishing accuracy, prioritization of response speed over quality, and frustrating issues with billing and customer support. Many believe Perplexity's aggressive user acquisition strategy, relying on free tiers, undermined its long-term potential by attracting non-paying users and then degrading service for paying ones.
The LLM Echo Chamber: AI Eating Its Own Tail
A central theme is the alarming prospect of Large Language Models (LLMs) consuming and reproducing AI-generated content, leading to a self-referential 'echo chamber' or 'snake eating its own tail.' Users share anecdotes of LLMs hallucinating facts based on obscure, single sources, or consistently favoring their own generated output over human-written alternatives. This raises serious questions about the future reliability of AI-driven search and information retrieval as the internet becomes increasingly saturated with machine-generated text.
Meta-Critique: Is the Messenger Also an AI?
A skeptical faction of commenters questions the legitimacy of the article itself and its source. Some accuse the prose of being "clearly a Claude artifact" or "AI slop." Others point out perceived ironies, such as the research firm's own website being labeled "nonsense" or not appearing on the Tranco list despite criticizing low-ranked domains. This critical lens highlights a broader distrust of online content in the AI era, where even reports exposing AI-generated content can be met with suspicion.