Key Takeaways:
- Cloudflare alleges Perplexity AI used deceptive methods to scrape content from websites that explicitly disallowed bots in their robots.txt files.
- The accusations claim Perplexity masked its identity using spoofed user agents and alternate IP addresses.
- Perplexity denies the behavior, calling Cloudflare’s claims a “sales pitch” and distancing itself from the implicated bot.
- This follows earlier reports from Wired and Forbes accusing Perplexity of similar scraping practices despite explicit blocks.
- Legal pressure is mounting, including demands from the BBC and lawsuits from Dow Jones, raising questions about fair use, consent, and transparency in AI data collection.
Cloudflare has accused Perplexity AI, a fast-growing AI startup known for its conversational search engine, of scraping web content in direct violation of publisher guidelines. According to a blog post from Cloudflare, the company detected Perplexity’s crawler accessing content from websites that had explicitly blocked such activity using robots.txt—a standard mechanism for indicating that bots should not access certain parts of a site.
Allegations of Deceptive Crawling
Per Cloudflare’s analysis, Perplexity’s systems didn’t just ignore these directives—they actively masked their identity. When its declared crawler (PerplexityBot) was blocked, the company allegedly reverted to using spoofed user-agent strings, such as impersonating a macOS Chrome browser, to continue accessing restricted content. Cloudflare described the behavior as an “effort to bypass intentional restrictions set by website operators.”
The company claims these masked crawlers generated “millions of requests per day” across tens of thousands of domains.
Perplexity’s Response
Jesse Dwyer, a Perplexity spokesperson, dismissed Cloudflare’s post as misleading, stating that “no content was actually accessed” and suggesting that the traffic in question did not originate from their systems. In a later clarification, Dwyer said the IPs Cloudflare referenced were not associated with Perplexity.
This response appears at odds with a growing body of complaints that have emerged in recent months. Wired and Forbes previously reported that Perplexity accessed sites using undisclosed IPs even when robots.txt directives were in place. One Wired investigation traced a specific AWS IP address back to visits on Condé Nast properties despite explicit blocks.
Building Legal Tension
The issue of consent-based scraping has begun to spill into legal territory. In June 2025, the BBC issued a formal cease-and-desist letter to Perplexity, demanding that the company delete previously scraped content, halt all future scraping, and provide compensation. The BBC claims its journalism was reproduced verbatim in some cases, without attribution or permission.
Perplexity rejected the BBC’s claims, calling them “factually inaccurate and opportunistic.” The company maintains that it aggregates public information under what it believes to be fair use and that it is not training large language models from scratch but instead indexing the web to deliver AI-generated summaries.
The firm is also the subject of a broader lawsuit filed by Dow Jones and other News Corp subsidiaries, which accuse the company of infringing on copyrighted content and republishing without license.
Technical and Legal Questions Ahead
At the heart of the dispute lies the question of whether ignoring a robots.txt file constitutes a breach of law. While robots.txt itself is not enforceable as legislation, it has been recognized in various legal interpretations as forming the basis for implied terms or contractual expectations.
If a crawler claims to respect robots.txt but circumvents those directives—particularly while also cloaking its identity—it may open itself to claims of bad faith, misrepresentation, or even computer fraud under U.S. and international laws.
This has prompted Cloudflare and others to consider new defensive tools. In May, Cloudflare introduced capabilities allowing publishers to block AI crawlers entirely or require payment to access site content—effectively creating a licensing marketplace for training data.
The Broader Implications for Generative AI
Perplexity’s case underscores a tension now surfacing across the AI landscape. As generative tools seek fresh, high-quality web content to power their responses, publishers are increasingly pushing back. Some opt to partner and license content. Others attempt to block it. The arms race between bots and websites is escalating, and the current legal and ethical frameworks are struggling to keep pace.
While Perplexity is not alone—Google, OpenAI, and others have faced similar critiques—the company’s alleged use of misdirection in scraping behavior may carry reputational and legal consequences beyond the norm.
What’s Next?
The outcome could help shape future norms around AI data collection. If courts or regulators interpret behavior like Perplexity’s as deceptive or unlawful, it may force AI firms to reevaluate how they source content—and how transparent they are with the public and publishers.
For now, Perplexity denies any wrongdoing, but scrutiny is increasing, and the company’s web practices may become a test case for broader industry accountability.
Learn how AI Agents can supercharge your company’s profits and productivity at TMC’s AI Agent Event in Sept 29-30, 2025 in DC.
Rich Tehrani serves as CEO of TMC and chairman of ITEXPO #TECHSUPERSHOW Feb 10-12, 2026 and is CEO of RT Advisors and is a Registered Representative (investment banker) with and offering securities through Four Points Capital Partners LLC (Four Points) (Member FINRA/SIPC). He handles capital/debt raises as well as M&A. RT Advisors is not owned by Four Points.
The above is not an endorsement or recommendation to buy/sell any security or sector mentioned. No companies mentioned above are current or past clients of RT Advisors.
The views and opinions expressed above are those of the participants. While believed to be reliable, the information has not been independently verified for accuracy. Any broad, general statements made herein are provided for context only and should not be construed as exhaustive or universally applicable.
Portions of this article may have been developed with the assistance of artificial intelligence, which may have contributed to ideation, content generation, factual review, or editing.






