Engine deep-dive
How Perplexity selects and cites sources
Perplexity is the only AI search engine that shows its sources with every answer. That makes it the most transparent engine, and the most actionable to optimize for.
Why Perplexity matters most for B2B services
- Every answer includes clickable source links. Your website gets actual traffic, not just a mention.
- Perplexity users are research-heavy. They are comparing, evaluating, making decisions. High-intent traffic.
- Perplexity referral traffic tends to be high-intent: users arrive mid-research, already comparing options.
- Perplexity searches on every query, so a new page can be cited within days of being indexed instead of waiting for a model cycle.
How Perplexity's retrieval works
Perplexity is a retrieval-first system. Unlike ChatGPT, which can answer from training data alone, Perplexity always searches the web first and builds its answer from what it finds.
The pipeline, in outline:
- The query goes to search.
- PerplexityBot's index and live fetches supply candidate pages.
- The model writes an answer from passages in those pages.
- Numbered citations point to the specific pages it used.
This means traditional content quality signals matter here. If your page would rank well in a search engine for a query, it has a good chance of being cited by Perplexity for the same query.
The 4 signals Perplexity prioritizes
Direct query match in title and headings
Perplexity favors pages whose title or H1 closely matches the user's query. 'Best MSP for Healthcare Compliance' gets cited for that exact query far more often than a generic 'Our Services' page. Match the query intent in your page structure.
Recency and freshness
Perplexity strongly prefers recent content: a 2026 article outperforms an identical 2024 one. Date your content clearly and update existing pages with current-year references. Perplexity checks publication dates.
Structured, scannable content
Perplexity extracts specific paragraphs to cite. Clear headings, numbered lists, comparison tables and FAQ sections give it discrete chunks to quote. Long, unstructured paragraphs are harder to extract and get cited less.
Domain trust for the topic
Perplexity weights domain relevance to the query topic. For 'best MSP in Ontario,' a page on an IT services domain carries more weight than the same content on a generic blog. Your own domain is an asset for queries about your service category.
Perplexity vs ChatGPT: key differences
| Signal | Perplexity | ChatGPT |
|---|---|---|
| Sources shown | Always, with clickable links | Sometimes, depends on mode |
| Update speed | Days (searches every query) | Days, when search runs |
| Traffic attribution | Clear (referral in GA4) | Indirect (some referral) |
| Content type preferred | Structured, specific, dated | Schema-marked, entity-verified |
| White space | Large but filling faster | Massive |
What to do for Perplexity visibility
- Match query intent in your page titles. "IT Support for Dental Practices" beats "Our IT Services." Perplexity matches titles to queries literally.
- Date your content. Include "2026" in titles and headings for comparison/recommendation content. Perplexity rewards freshness.
- Structure for extraction. Use H2/H3 headings, numbered lists, comparison tables. Give Perplexity discrete chunks to quote.
- Track your Perplexity referral traffic. In GA4, filter referrals by perplexity.ai. This is the only AI engine where you can directly measure traffic from citations.
How Perplexity retrieves and cites, from Perplexity's own documents
- From the vendor PerplexityBot is the index, and it is not a training crawler. Perplexity says PerplexityBot surfaces and links websites in its search results, is not used to crawl content for AI foundation models, and recommends allowing it in robots.txt if you want to appear. Perplexity: its crawlers
- From the vendor A live fetch follows the user, not your robots.txt. Perplexity-User visits a page when a user's question needs it. Because a user requested the fetch, Perplexity states it generally ignores robots.txt rules. Blocking the index crawler does not stop a live read. Perplexity: Perplexity-User
- From the vendor A passage has to make sense on its own. Perplexity's own retrieval research warns that a chunk 'may contain the answer but lack the supporting context needed to understand or verify it'. It found 1,067 of the 2,099 test questions in its benchmark needed evidence from two or more places. Write sections that name you, the service and the market inside the passage itself. Perplexity Research
- Our read It cites the page, so the page wins. Perplexity links the specific page it used, not your domain. A sharp page on a small site beats a vague page on a big one for the query it answers.
What moves your Tofu Index and your Bofu Index on Perplexity
Your AI Visibility Index is two numbers with two jobs.
- The Tofu Index asks whether Perplexity puts you in the consideration set: the share of buying queries where it names you.
- The Bofu Index asks whether it picks you: the share of head-to-head comparisons, you against a named rival, where it chooses you.
- A free scan reports your Tofu Index on every engine we query. Fix and Dominate add the Bofu Index.
| Lever | Tofu Index: are you considered | Bofu Index: are you picked |
|---|---|---|
| Allow PerplexityBot | Out of the index without it, so you are a candidate only when a live fetch happens to reach you. | Indirect. |
| A dated page per buying query, headed with the query | Perplexity moves first: a page published this month can be cited this month. | Small on its own. |
| Self-contained sections | A passage that names you, the service and the market survives being cut out of the page and quoted. | A comparison passage that states the whole trade-off is the one quoted when the buyer asks which to pick. |
| Alternatives and versus pages | Puts you in the answers to alternatives queries. | The main lever: Perplexity quotes the page that settles the comparison. |
| Coverage and directory listings | More independent pages that name you means more candidates per query. | Independent pages that prefer you are evidence it can cite for the pick. |
What our own scans measured
In our July to August study of 46 completed scans, Perplexity named the brand in 7.4% of buying queries, roughly double the next engine, at firms that scanned while suspecting they were missing. That is not a random sample of B2B websites.
- That lead came from searching while the others answered from memory.
- Every engine we query now searches first, so we no longer rank Perplexity ahead of the rest.
- The rest of the field sat between 0.4% and 4.0% on the same suspect-selected reports.
- That gap is a mechanism, not a preference. A page you publish this month can be cited this month, while an engine answering from training data waits for a model cycle.
Grounding shows it too. In our 15-region managed IT study, 101 of the 257 firms Perplexity named could be corroborated against a real website and an independent listing, against 8 of the 114 from ChatGPT and 2 of the 129 from Claude.
It is also the argument for keeping Perplexity in a scan instead of trimming it for cost. It disagrees with the memory-based engines more often than they disagree with each other. A disagreement is usually where a winnable query is hiding.
Read the MSP AI Visibility Report 2026, which shows the method and the names we could not verify.
How we count, because the denominator does the work
These rates cover buying queries only: what a buyer types when choosing a vendor, before they know who you are.
- We deliberately leave out queries that contain the company's own name.
- Across the same 46 reports, every engine answers a name-in-the-question probe with that company 82% to 100% of the time.
- Folding those in lifts the headline to about 10% for the field and 19% for the leader.
- We published that flattering version on this page until 2026-08-12, and have corrected it.
The buyer who already knows your name is not the buyer you are missing. A number that counts them measures your own marketing back at you.
What you get, and what it costs
Perplexity is one of the 5 engines every plan queries, including the free Track plan. Engines are not plan-gated. What a plan changes is how many queries you get per scan, how often the scan runs, whether the scan adds the head-to-head comparisons behind the Bofu Index, and whether we draft the pages that close the gaps it finds.
| Plan | Price | Queries per engine | Scans | Drafts a month |
|---|---|---|---|---|
| Track | Free | 5 | 1 a month | Ideas only |
| Fix | $99/mo | 10 | 4 a month | 10 |
| Dominate | $499/mo | 25 | 4 a month | 25 |
Rule is portfolio level and sales-assisted. Full pricing · How the scan works
See if Perplexity cites your company
Check your AI visibility freeFrequently asked questions
Does Perplexity use Google search results?
No, not directly. Perplexity runs its own crawler, PerplexityBot, and fetches pages live when a question needs them. A page can rank well on Google and never appear in Perplexity, and the reverse happens too. Google rank is no proxy for Perplexity visibility: measure Perplexity itself.
How does Perplexity choose which sources to cite?
It favors pages that directly answer the query, are clearly structured, recently published or updated, and sit on domains it already trusts. It cites the specific page, not the domain, so a sharp, well-structured page on a modest site can beat a vague page on a big one. Specificity beats brand size here more than on most engines.
How do I actually get cited by Perplexity?
Answer the specific buying query on a specific, well-structured page: a heading that matches the question, a direct answer up top, concrete detail, and schema where it fits. Perplexity re-reads the live web, so it rewards fresh, precise pages. Then corroborate: it leans on sources it already trusts, so third-party mentions and reviews in your category raise the odds it reaches for you.
How fast does Perplexity pick up new or updated content?
Fast, relative to the training-based engines. Perplexity retrieves live results per query, so a newly indexed or updated page can appear within days to a couple of weeks, not model cycles. Every engine we query now searches the live web, so a content change can show up on any of them once its search finds the page. Perplexity and Google AI Mode have always worked this way.
Why does Perplexity cite my competitors and not me?
Usually their pages answer the specific query more directly than yours, or they have more third-party corroboration in your category, so Perplexity treats them as the safer source. It is a readout of your content and citations, not a fixed ranking. Find the queries where a competitor is cited and you are not, then build the page and proof that answer them better.
Does Perplexity have ads or paid placement?
Perplexity has been introducing sponsored formats, but its core answer citations are still earned: you cannot pay to be named as a recommended source inside an organic answer. As with the other engines, the durable way in is being the best, most-corroborated answer. Ad formats in AI search are evolving quickly, so watch this space.
How do I know if Perplexity actually crawled my site?
Look for PerplexityBot in your server logs or your CDN's bot analytics. That confirms the crawler reached a page, which is the precondition for citation, though a visit does not guarantee you get cited. If PerplexityBot is not reaching your key pages at all, check that your robots.txt and any bot-blocking rules are not quietly excluding it.
Can I see if Perplexity is driving traffic to my site?
Yes, and better than most. Perplexity referral traffic shows up in analytics as perplexity.ai, and because Perplexity always attaches clickable citation links to its answers, it is the most attribution-friendly AI engine. If you get cited, you can usually see the clicks, unlike ChatGPT where much of the influence is invisible.
Sources and further reading
- Perplexity Documentation: How Perplexity's retrieval-augmented generation pipeline works
- Perplexity: PerplexityBot and Perplexity-User: what each crawler does and how robots.txt applies
- Perplexity Research: contextual embedding beyond the gold passage: Technical approach to real-time web retrieval and citation
- G2 AI Search Insight Report (2026): Buyer behavior shift to AI search
- Schema.org FAQPage Specification: Structured data format for AI-parseable FAQ content