{"id":24400,"date":"2025-08-04T12:59:32","date_gmt":"2025-08-04T16:59:32","guid":{"rendered":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/?p=24400"},"modified":"2025-08-04T15:12:39","modified_gmt":"2025-08-04T19:12:39","slug":"perplexity-ai-accused-of-scraping-sites-that-explicitly-blocked-crawlers","status":"publish","type":"post","link":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/ai\/perplexity-ai-accused-of-scraping-sites-that-explicitly-blocked-crawlers.html","title":{"rendered":"Perplexity AI Accused of Scraping Sites That Explicitly Blocked Crawlers"},"content":{"rendered":"\n<p>Key Takeaways:<\/p>\n\n\n\n<ul>\n<li>Cloudflare alleges Perplexity AI used deceptive methods to scrape content from websites that explicitly disallowed bots in their robots.txt files.<\/li>\n\n\n\n<li>The accusations claim Perplexity masked its identity using spoofed user agents and alternate IP addresses.<\/li>\n\n\n\n<li>Perplexity denies the behavior, calling Cloudflare\u2019s claims a \u201csales pitch\u201d and distancing itself from the implicated bot.<\/li>\n\n\n\n<li>This follows earlier reports from Wired and Forbes accusing Perplexity of similar scraping practices despite explicit blocks.<\/li>\n\n\n\n<li>Legal pressure is mounting, including demands from the BBC and lawsuits from Dow Jones, raising questions about fair use, consent, and transparency in AI data collection.<\/li>\n<\/ul>\n\n\n\n<p>Cloudflare has <a href=\"https:\/\/techcrunch.com\/2025\/08\/04\/perplexity-accused-of-scraping-websites-that-explicitly-blocked-ai-scraping\/\">accused<\/a> Perplexity AI, a fast-growing AI startup known for its conversational search engine, of scraping web content in direct violation of publisher guidelines. According to a blog post from Cloudflare, the company detected Perplexity\u2019s crawler accessing content from websites that had explicitly blocked such activity using robots.txt\u2014a standard mechanism for indicating that bots should not access certain parts of a site.<\/p>\n\n\n\n<p><strong>Allegations of Deceptive Crawling<\/strong><\/p>\n\n\n\n<p>Per Cloudflare\u2019s analysis, Perplexity\u2019s systems didn\u2019t just ignore these directives\u2014they actively masked their identity. When its declared crawler (PerplexityBot) was blocked, the company allegedly reverted to using spoofed user-agent strings, such as impersonating a macOS Chrome browser, to continue accessing restricted content. Cloudflare described the behavior as an \u201ceffort to bypass intentional restrictions set by website operators.\u201d<\/p>\n\n\n\n<p>The company claims these masked crawlers generated \u201cmillions of requests per day\u201d across tens of thousands of domains.<\/p>\n\n\n\n<p><strong>Perplexity\u2019s Response<\/strong><\/p>\n\n\n\n<p>Jesse Dwyer, a Perplexity spokesperson, dismissed Cloudflare\u2019s post as misleading, stating that \u201cno content was actually accessed\u201d and suggesting that the traffic in question did not originate from their systems. In a later clarification, Dwyer said the IPs Cloudflare referenced were not associated with Perplexity.<\/p>\n\n\n\n<p>This response appears at odds with a growing body of complaints that have emerged in recent months. Wired and Forbes previously reported that Perplexity accessed sites using undisclosed IPs even when robots.txt directives were in place. One Wired investigation traced a specific AWS IP address back to visits on Cond\u00e9 Nast properties despite explicit blocks.<\/p>\n\n\n\n<p><strong>Building Legal Tension<\/strong><\/p>\n\n\n\n<p>The issue of consent-based scraping has begun to spill into legal territory. In June 2025, the BBC issued a formal cease-and-desist letter to Perplexity, demanding that the company delete previously scraped content, halt all future scraping, and provide compensation. The BBC claims its journalism was reproduced verbatim in some cases, without attribution or permission.<\/p>\n\n\n\n<p>Perplexity rejected the BBC\u2019s claims, calling them \u201cfactually inaccurate and opportunistic.\u201d The company maintains that it aggregates public information under what it believes to be fair use and that it is not training large language models from scratch but instead indexing the web to deliver AI-generated summaries.<\/p>\n\n\n\n<p>The firm is also the subject of a broader lawsuit filed by Dow Jones and other News Corp subsidiaries, which accuse the company of infringing on copyrighted content and republishing without license.<\/p>\n\n\n\n<p><strong>Technical and Legal Questions Ahead<\/strong><\/p>\n\n\n\n<p>At the heart of the dispute lies the question of whether ignoring a robots.txt file constitutes a breach of law. While robots.txt itself is not enforceable as legislation, it has been recognized in various legal interpretations as forming the basis for implied terms or contractual expectations.<\/p>\n\n\n\n<p>If a crawler claims to respect robots.txt but circumvents those directives\u2014particularly while also cloaking its identity\u2014it may open itself to claims of bad faith, misrepresentation, or even computer fraud under U.S. and international laws.<\/p>\n\n\n\n<p>This has prompted Cloudflare and others to consider new defensive tools. In May, Cloudflare introduced capabilities allowing publishers to block AI crawlers entirely or require payment to access site content\u2014effectively creating a licensing marketplace for training data.<\/p>\n\n\n\n<p><strong>The Broader Implications for Generative AI<\/strong><\/p>\n\n\n\n<p>Perplexity\u2019s case underscores a tension now surfacing across the AI landscape. As generative tools seek fresh, high-quality web content to power their responses, publishers are increasingly pushing back. Some opt to partner and license content. Others attempt to block it. The arms race between bots and websites is escalating, and the current legal and ethical frameworks are struggling to keep pace.<\/p>\n\n\n\n<p>While Perplexity is not alone\u2014Google, OpenAI, and others have faced similar critiques\u2014the company\u2019s alleged use of misdirection in scraping behavior may carry reputational and legal consequences beyond the norm.<\/p>\n\n\n\n<p><strong>What\u2019s Next?<\/strong><\/p>\n\n\n\n<p>The outcome could help shape future norms around AI data collection. If courts or regulators interpret behavior like Perplexity\u2019s as deceptive or unlawful, it may force AI firms to reevaluate how they source content\u2014and how transparent they are with the public and publishers.<\/p>\n\n\n\n<p>For now, Perplexity denies any wrongdoing, but scrutiny is increasing, and the company\u2019s web practices may become a test case for broader industry accountability.<\/p>\n\n\n\n<p><strong>Le<em>arn how AI Agents can supercharge your company\u2019s profits and productivity at&nbsp;<a href=\"http:\/\/www.tmcnet.com\/\">TMC\u2019s&nbsp;<\/a><a href=\"https:\/\/www.aiagentevent.com\/\">AI Agent Event&nbsp;<\/a>in Sept 29-30, 2025 in DC.<\/em><\/strong><\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><a href=\"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-content\/uploads\/2025\/07\/AiAgent-500x600-Speaker-logos-v3.jpg\"><img loading=\"lazy\" decoding=\"async\" width=\"600\" height=\"500\" src=\"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-content\/uploads\/2025\/07\/AiAgent-500x600-Speaker-logos-v3.jpg\" alt=\"\" class=\"wp-image-23949\"\/><\/a><\/figure><\/div>\n\n\n<p><em>Rich Tehrani serves as CEO of&nbsp;<a href=\"http:\/\/www.tmcnet.com\/\">TMC<\/a>&nbsp;and chairman of&nbsp;<a href=\"http:\/\/www.itexpo.com\/\">ITEXPO<\/a>&nbsp;#TECHSUPERSHOW Feb 10-12, 2026 and is CEO of&nbsp;<a href=\"https:\/\/www.rt-advisors.com\/\">RT Advisors<\/a>&nbsp;and is&nbsp;a Registered Representative (investment banker) with and offering securities through&nbsp;<a href=\"https:\/\/www.4pointscapital.com\/\">Four Points Capital Partners LLC&nbsp;<\/a>(Four Points) (Member FINRA\/SIPC). He handles capital\/debt raises as well as M&amp;A. RT Advisors is not owned by Four Points.<\/em><\/p>\n\n\n\n<p>The above is not an endorsement or recommendation to buy\/sell any security or sector mentioned. No companies mentioned above are current or past clients of RT Advisors.<\/p>\n\n\n\n<p>The views and opinions expressed above are those of the participants. While believed to be reliable, the information has not been independently verified for accuracy. Any broad, general statements made herein are provided for context only and should not be construed as exhaustive or universally applicable.<\/p>\n\n\n\n<p><em>Portions of this article may have been developed with the assistance of artificial intelligence, which may have contributed to ideation, content generation, factual review, or editing<\/em>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Key Takeaways: Cloudflare has accused Perplexity AI, a fast-growing AI startup known for its conversational search engine, of scraping web content in direct violation of publisher guidelines. According to a blog post from Cloudflare, the company detected Perplexity\u2019s crawler accessing content from websites that had explicitly blocked such activity using robots.txt\u2014a standard mechanism for indicating<\/p>\n","protected":false},"author":44,"featured_media":24401,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[194],"tags":[],"post_mailing_queue_ids":[],"_links":{"self":[{"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/posts\/24400"}],"collection":[{"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/users\/44"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/comments?post=24400"}],"version-history":[{"count":2,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/posts\/24400\/revisions"}],"predecessor-version":[{"id":24427,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/posts\/24400\/revisions\/24427"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/media\/24401"}],"wp:attachment":[{"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/media?parent=24400"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/categories?post=24400"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.tmcnet.com\/blog\/rich-tehrani\/wp-json\/wp\/v2\/tags?post=24400"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}