New · A dedicated AI SEO channel: see conversions from ChatGPT, Perplexity, Claude & Gemini →
← Blog
Part of: How to Track ChatGPT, Perplexity & Gemini Conversions, Not Just Mentions →

PerplexityBot: User Agents, IPs, robots.txt Rules and Log Checks

Portrait of Samy ThuillierBy ··12 min read
PerplexityBot crawling a website while a robots.txt file and firewall decide which pages it can reach

PerplexityBot is Perplexity’s web crawler: it finds and indexes pages so Perplexity can surface and link them in its answers, and Perplexity says it is not used to train AI models. You control it in robots.txt with User-agent: PerplexityBot. A second agent, Perplexity-User, fetches pages when someone asks a question, and it generally ignores robots.txt.

This guide covers both agents, their user agent strings and IP lists, copy-paste robots.txt rules, firewall setup, a script that tells real Perplexity traffic from fakes, the stealth crawling dispute, the claims where the popular guides contradict the official docs, and how to find out whether letting the bot in produces any business at all.

What Is PerplexityBot?

Perplexity is an answer engine. It replies to questions with a written answer and numbered links to the pages it used. To have pages to choose from, it keeps its own index of the web, and PerplexityBot is the crawler that builds and refreshes that index.

Perplexity’s crawler documentation describes it in three statements that matter for site owners:

  • It is designed to surface and link websites in Perplexity’s search results.
  • It is not used to crawl content for AI foundation models.
  • If you want to appear in those results, Perplexity recommends allowing PerplexityBot in robots.txt and permitting requests from its published IP ranges.

It is a search crawler, not a training crawler. Blocking it removes you from the pool of pages Perplexity can cite, and it does not buy you any training protection, because Perplexity documents no separate training crawler to begin with.

PerplexityBot vs Perplexity-User

Perplexity documents two agents. Most robots.txt mistakes come from treating them as one.

PerplexityBotPerplexity-User
JobCrawls in the background to build the search indexVisits a page when a user’s question needs it, and may link it in the answer
Follows robots.txtYes, for groups addressed to PerplexityBotGenerally no: Perplexity says it ignores robots.txt because a user asked
Used for model trainingNo, per PerplexityNo, per Perplexity
IP listperplexitybot.jsonperplexity-user.json
How to restrict itrobots.txtFirewall or bot management rule

A Perplexity-User hit is the closer signal that your page was used in a live answer. A PerplexityBot hit only means your page was indexed or refreshed.

The crawler page lists only these two agents. It says nothing about how Perplexity’s browser agent identifies itself when it browses for a user, so do not assume these two tokens cover every request Perplexity products make.

User Agent Strings and IP Ranges

The full user agent strings, as published in Perplexity’s docs:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

In robots.txt you use only the product token: PerplexityBot or Perplexity-User. Matching is case-insensitive.

Perplexity publishes its IP addresses as JSON, one file per agent, at https://www.perplexity.com/perplexitybot.json and https://www.perplexity.com/perplexity-user.json. When we fetched them on October 5, 2026, the PerplexityBot file listed 8 IPv4 prefixes (created February 7, 2025) and the Perplexity-User file listed 4 (created October 17, 2025). The lists change, so read them from the URL each time instead of copying them into a config file once.

The user agent proves nothing on its own. Anyone can send PerplexityBot/1.0 in a header. A request is genuinely from PerplexityBot only when the user agent and the source IP both match.

How to Allow or Block PerplexityBot in robots.txt

Your robots.txt lives at the root of the host, for example https://example.com/robots.txt. Each subdomain needs its own file. Perplexity says each setting works independently and changes can take up to 24 hours to show.

Allow everything

User-agent: PerplexityBot Allow: /

If your file has no rule that blocks PerplexityBot, it is already allowed. An explicit group is useful when a User-agent: * group blocks bots by default. Read the next section before you add it.

Block everything

User-agent: PerplexityBot Disallow: /

This keeps the declared crawler out and, over time, your pages out of Perplexity’s index. It does not stop Perplexity-User. To keep that agent out too, you need a firewall rule (covered below).

Block some paths

User-agent: PerplexityBot Disallow: /members/ Allow: /

Under the Robots Exclusion Protocol (RFC 9309), the most specific rule wins, meaning the matching rule with the longest path. Disallow: /members/ beats Allow: / for anything under /members/, whatever order you write them in. When an Allow and a Disallow are equally long, Allow wins.

The robots.txt Mistake That Opens Your Admin Pages

Here is a common file. The site blocks private sections for every bot, then adds a group to let PerplexityBot in:

User-agent: * Disallow: /admin/ Disallow: /cart/ Disallow: /search User-agent: PerplexityBot Allow: /

It looks harmless. It is not. RFC 9309 says a crawler obeys the group that matches its product token, and falls back to the * group only when no group matches. Once a PerplexityBot group exists, PerplexityBot ignores the * group entirely. In this file PerplexityBot may crawl /admin/, /cart/ and your internal search pages, because its own group allows everything.

The fix is to repeat every rule you still want inside the specific group:

User-agent: * Disallow: /admin/ Disallow: /cart/ Disallow: /search User-agent: PerplexityBot Disallow: /admin/ Disallow: /cart/ Disallow: /search Allow: /

Two more rules from the standard that save debugging time:

  • Order of groups does not matter. Putting the PerplexityBot group above or below the * group changes nothing. Group selection is by user agent, not position.
  • Duplicate groups merge. If two groups both name PerplexityBot, the crawler combines their rules into one group. Search the whole file before you assume a rule is the only one.

The same logic applies to every crawler you name: GPTBot, ClaudeBot, Googlebot. Any time you add a named group, copy the restrictions from * into it.

Firewall and CDN Rules for Perplexity

robots.txt is a request. A web application firewall (WAF) is enforcement. Perplexity’s docs say sites behind a WAF may need to allow its agents explicitly and give the steps:

  • Cloudflare WAF: create a rule where the user agent contains PerplexityBot or Perplexity-User and the source IP is in the ranges from the two JSON files, then set the action (allow or skip).
  • AWS WAF: create IP sets from the JSON files, add string match conditions for the two user agents, create allow rules and associate them with your web ACL.

To block instead of allow, use the same conditions with a block action. A rule on the user agent alone stops the declared agents, which is all robots.txt does too. A rule that also requires the published IPs to allow traffic means a fake PerplexityBot from another network gets no special treatment.

Crawl rate. Perplexity’s docs do not mention the Crawl-delay directive, and RFC 9309 does not define it, so do not count on it. If Perplexity traffic loads your server, rate limit at the CDN or server and return 429 with a Retry-After header.

Cloudflare specifically. In its August 2025 report, Cloudflare said it had de-listed Perplexity as a verified bot. If you rely on a “verified bots” allow rule, Perplexity may not be in it. Look at the security events for requests with Perplexity user agents and see what action was applied.

Verify Real PerplexityBot Traffic in Your Logs

Google Search Console only reports Google’s own crawlers, so it shows nothing about Perplexity. Your web server access logs, or your CDN’s logs and bot dashboard, are the record. A quick look, for the common combined log format:

grep -E "PerplexityBot|Perplexity-User" access.log | awk '{print $1, $9, $7}' | sort | uniq -c | sort -rn | head -20

That prints the most frequent IP, status code and path combinations for both agents. It still trusts the user agent. To separate real Perplexity traffic from impostors, check every IP against the published CIDR ranges. This Python 3 script uses only the standard library, downloads both lists, and counts hits by agent, verdict and status code:

import ipaddress, json, re, sys, urllib.request from collections import Counter LISTS = { "PerplexityBot": "https://www.perplexity.com/perplexitybot.json", "Perplexity-User": "https://www.perplexity.com/perplexity-user.json", } def load(url): with urllib.request.urlopen(url) as r: data = json.load(r) return [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in data["prefixes"]] nets = {agent: load(url) for agent, url in LISTS.items()} # combined log format: IP - - [time] "GET /path HTTP/1.1" status bytes "referer" "user-agent" line_re = re.compile(r'^(\S+) .*?"\S+ (\S+) [^"]*" (\d{3}) .*"([^"]*)"$') results = Counter() for line in open(sys.argv[1], errors="replace"): m = line_re.match(line.strip()) if not m: continue ip, path, status, ua = m.groups() for agent, ranges in nets.items(): if agent + "/" in ua: addr = ipaddress.ip_address(ip) real = any(addr in n for n in ranges) results[(agent, "verified" if real else "SPOOFED", status)] += 1 for (agent, verdict, status), n in sorted(results.items()): print(f"{agent:16} {verdict:9} {status} {n}")

Run it with python3 check_perplexity.py /var/log/nginx/access.log. On a four-line test log it printed:

Perplexity-User verified 200 1 PerplexityBot SPOOFED 403 1 PerplexityBot verified 200 2

How to read the result:

  • Verified with 200: Perplexity reaches your pages. Check that a /robots.txt request appears before the crawl, which shows it read your rules.
  • Verified with 403, 429 or 503: your firewall, CDN or rate limiter is turning Perplexity away, whatever robots.txt says.
  • Spoofed: something is using the PerplexityBot name from an IP outside the list. Treat it like any unknown scraper.
  • No hits at all: either you are blocked upstream (CDN rules apply before your server logs) or Perplexity has not crawled you. Check the CDN’s logs before concluding anything.

If your log format differs, adjust the regular expression. The IP check is the part that matters.

When PerplexityBot Is Allowed but Still Cannot Crawl

robots.txt itself can fail in ways that change what a crawler is allowed to do. RFC 9309 sets the rules, and they surprise people:

SymptomLikely causeFix
No PerplexityBot page fetches, only /robots.txt requestsrobots.txt returns a 5xx error. Under RFC 9309 a crawler must then assume everything is disallowedMake robots.txt return 200. Check that it is not behind the same rate limit or challenge as the rest of the site
PerplexityBot crawls paths you thought were blockedrobots.txt returns a 4xx (often 403 from a WAF or 404). Crawlers may then crawl anythingServe the file with a 200 to every agent, even ones you block elsewhere
Rules seem ignored right after an editCached copy: crawlers should not keep it more than 24 hours, and Perplexity says changes take up to 24 hoursWait a day, then check the logs again
Rules ignored on one host but not anotherrobots.txt redirects too many times, or the subdomain has no file of its ownKeep redirects to a minimum (the standard asks crawlers to follow at least five) and add a file per host
Paths under a named group crawled unexpectedlyA PerplexityBot group replaced the * groupCopy the * restrictions into the named group (see above)
Verified hits get 403 or a challenge pageBot management or WAF ruleAdd an allow rule matching the user agent and the published IPs
Page fetched with 200 but content never shows up in answersMain content is loaded by JavaScript after the HTML arrivesRun curl with the PerplexityBot user agent and check the main text is in the raw HTML

The Stealth Crawling Dispute

Two public reports say Perplexity reached sites that had blocked it, using requests that did not identify as Perplexity.

June 2024. Developer Robb Knight, then Wired, reported Perplexity fetching pages on sites that disallowed PerplexityBot, from cloud IPs with a generic browser user agent. As reported at the time, Perplexity said it did not ignore robots.txt and attributed the traffic to a third-party crawling provider.

August 2025. Cloudflare’s investigation set up brand-new domains that disallowed all crawling in robots.txt and blocked both Perplexity agents with WAF rules. It then asked Perplexity about them and got detailed answers. Cloudflare reported:

  • The declared Perplexity-User agent made 20 to 25 million requests a day across its network.
  • An undeclared crawler, posing as Chrome on macOS, made 3 to 6 million a day, from IPs outside Perplexity’s published ranges and rotating across networks (ASNs).
  • Cloudflare de-listed Perplexity as a verified bot and added signatures for the undeclared crawler to its managed rule that blocks AI crawling, which is available to free plans.
  • When the stealth crawler was blocked too, Perplexity’s answers about the test sites became vaguer and lacked details from the original pages.

The practical lesson does not depend on who is right. robots.txt only binds requests that identify themselves. If content must not be read by any automated system, put it behind a login or a bot management product that fingerprints traffic, not just a robots.txt line.

Should You Allow or Block PerplexityBot?

PerplexityBot does not affect Google. Googlebot reads only its own groups and the * group, and Perplexity does not pass any ranking signal to Google. So the decision is only about Perplexity.

Your situationRecommended setup
Open content you want cited: guides, docs, product and service pagesAllow PerplexityBot; allow Perplexity-User at the firewall
Paywalled or member content on part of the siteAllow PerplexityBot with Disallow on the paid paths; firewall rule for those paths
Licensing terms forbid AI summaries of your contentDisallow PerplexityBot, plus a firewall rule for both agents and bot management
Server load is the only problemKeep it allowed and rate limit with 429 instead of blocking
You block it today and do not know whyCheck whether a default CDN setting did it; decide on purpose

If you allow it to get cited, getting crawled is only the entry ticket. What earns citations is covered in our Perplexity SEO guide.

How to Stop AI Crawlers (and What robots.txt Cannot Do)

Most AI companies run separate agents for training, for search indexing and for live user requests. Blocking one does not block the others. The main tokens, all used in robots.txt as User-agent: <token>:

CompanySearch or indexTrainingUser-triggered fetch
PerplexityPerplexityBotNone documentedPerplexity-User (generally ignores robots.txt)
OpenAIOAI-SearchBotGPTBotChatGPT-User
AnthropicClaude-SearchBotClaudeBotClaude-User
GoogleGooglebot (also feeds Search AI features)Google-Extended (a token only, no separate crawler)n/a
Common Crawln/aCCBot (open dataset used by many model builders)n/a

Vendors rename and add agents, so check each company’s crawler page before you rely on a list, including this one. Then accept the limits: robots.txt is honored only by crawlers that choose to, user-triggered fetchers are often exempt by design, and RFC 9309 states plainly that its rules are not a form of access authorization. For anything that must stay private, use authentication. For everything else, a firewall rule on user agent plus IP backs up what robots.txt asks.

Where the PerplexityBot Guides Contradict Each Other

We read the ten pages ranking for “perplexitybot” and compared them with Perplexity’s documentation and RFC 9309. Several popular claims are wrong or out of date:

Claim you will findWhat the primary source says
Blocking PerplexityBot only prevents AI trainingPerplexity says PerplexityBot is not used for model training. Blocking it removes you from its search results
PerplexityBot collects pages for model trainingSame: not used to crawl content for AI foundation models, per Perplexity
Perplexity does not publish IP rangesIt publishes two JSON files, one per agent
Perplexity-User respects robots.txtPerplexity says it generally ignores robots.txt because a user requested the fetch
Place the PerplexityBot group above User-agent: *Order does not matter in RFC 9309; a matching named group replaces * entirely
robots.txt changes take daysPerplexity says up to 24 hours
The user agent ends in perplexity.ai/bot; verify with reverse DNSThe current documented strings end in perplexity.ai/perplexitybot and perplexity.ai/perplexity-user; verification is by the published IP lists
Cloudflare reported the stealth crawling in 2024The 2024 reports came from Robb Knight and Wired; Cloudflare’s report is dated August 4, 2025

One question we could not settle from a primary source: how Perplexity responded to the 2025 report. One ranking guide says Perplexity acknowledged the issue and committed to stricter compliance; another says it disputed the findings. Neither links a Perplexity statement, so treat both as unconfirmed.

Is Allowing PerplexityBot Worth It? Measure the Outcome

Crawler hits are not visitors. A thousand PerplexityBot requests can produce zero customers. To decide whether access pays off, connect three numbers per landing page:

  1. Perplexity-User fetches from your verified logs: a rough proxy for how often a page is pulled into live answers.
  2. Sessions referred by perplexity.ai, from your analytics landing page report. Our guide to AI traffic analytics shows how to isolate them in GA4. Visits from apps that send no referrer will land in Direct, so this count is a floor.
  3. Conversions and their value on those sessions: demo requests, purchases, calls, each with a dollar value.
Worked example (illustrative numbers)

A B2B software site checks one month. Its pricing guide had 600 verified Perplexity-User fetches, 150 sessions referred by perplexity.ai, and 6 demo requests. The team values a demo at $400 (25% of demos close, average first-year deal $1,600, so 0.25 x $1,600 = $400). That page produced 6 x $400 = $2,400 of pipeline value from Perplexity, or $16 per Perplexity session.

Its glossary had 2,000 fetches, 80 sessions and 0 conversions. More crawling and citing, no business. The team keeps both pages open to PerplexityBot, but puts content work into pages like the pricing guide, and adds a demo link to the glossary to test whether those readers convert at all.

This only works if you track conversions per landing page and per AI source, not just traffic. Our AI conversion tracking guide walks through that setup. SEOConversion does it with one cookieless script: it identifies Perplexity visits from the referrer, records the conversions on each landing page and turns them into value, and leaves visits that arrive without a referrer in Direct instead of guessing.

FAQ

What is a PerplexityBot?

PerplexityBot is the web crawler Perplexity uses to find and index pages so it can surface and link them in its search answers. Perplexity says it is not used to crawl content for training AI foundation models. It identifies itself with PerplexityBot/1.0 in the user agent and follows robots.txt rules addressed to PerplexityBot.

Does blocking PerplexityBot affect my Google rankings?

No. Google ranks pages with its own crawlers and only reads robots.txt groups for its own user agents or the * group. A group for PerplexityBot does not change what Googlebot can crawl. Google Search Console also never shows PerplexityBot activity; you need server or CDN logs for that.

Is robots.txt illegal to ignore?

robots.txt is a voluntary convention, standardized in RFC 9309, and the standard itself says its rules are not a form of access authorization. Whether ignoring it creates legal exposure depends on the country, your terms of service and the case, so ask a lawyer if that matters to you. If you need enforcement rather than a request, use a firewall rule or a login.

Is the Perplexity AI app free?

Perplexity offers a free tier and paid subscriptions with extra features. That has no bearing on the crawler: PerplexityBot indexes public pages the same way whichever plan the person asking is on. Perplexity-User fetches happen when any user’s question needs a live page.

Is Cloudflare blocking AI crawlers?

Cloudflare lets each customer choose which declared AI crawlers can reach the site, and it offers a managed rule that blocks AI crawling, available on free plans too. In August 2025 it de-listed Perplexity as a verified bot and added signatures for an undeclared Perplexity crawler to that rule. Check your own dashboard settings, because a default you did not choose may be blocking PerplexityBot.

How do I stop AI from crawling my website?

Add a robots.txt group with Disallow: / for each AI crawler token you want to stop, such as GPTBot, ClaudeBot, CCBot, Google-Extended and PerplexityBot. robots.txt only works on crawlers that honor it and only for the tokens you name, so add a firewall or bot management rule for anything that must stay out. Remember that user-triggered fetchers such as Perplexity-User generally ignore robots.txt.

Find out whether Perplexity sends you customers, not just crawlers.

SEOConversion identifies visits from Perplexity and other AI assistants by referrer, records the conversions on each landing page and turns them into value with one cookieless script.

Start free