Inside Perplexity's ranking stack: chunks, hybrid retrieval and layered reranking
Perplexity is the most transparent of the major answer engines about how it works, largely because its founders keep explaining it in public. Pulling the public record together, from Aravind Srinivas's long-form interviews to engineering write-ups of the architecture, produces a clearer picture of the retrieval stack than we have for any Google surface. It also explains why Perplexity cites so differently from everything else.
It runs its own index
The first thing that matters is that Perplexity is not a wrapper around someone else's search results. It crawls with PerplexityBot and maintains its own index, which is why its citation profile diverges so sharply from Google's in the cross-engine data. If you block its crawler, you are not in the answers, and no amount of Google ranking substitutes.
It retrieves chunks, not pages
The unit of retrieval is a sub-document chunk. Architecture write-ups describe documents being split into fine-grained pieces and scored individually, so a single relevant paragraph inside a long article can surface without carrying the surrounding thousands of words into the model's context. This is the clearest engineering statement of the thesis running through most of our coverage: the passage is the unit of competition.
The hybrid step is worth dwelling on. Keyword matching catches exact terms, product names and error codes. Semantic matching catches meaning when the words differ. Fusing them means you cannot win purely on synonyms or purely on exact phrasing, which is why thin keyword-stuffed pages perform so badly here while genuinely specific technical writing performs well.
The rule that shapes everything
"The principle in Perplexity is you are not supposed to say anything that you don't retrieve."
That line, from Srinivas in conversation with Lex Fridman, is the design constraint that explains the product. Citations are not decoration added after generation; retrieval comes first and the answer is assembled from what was retrieved. It also explains Perplexity's appetite for fresh material and its unusually high reliance on community sources, documented in the cross-engine citation data where Reddit alone accounts for a very large share of its citations.
What to do this week
Confirm PerplexityBot can reach you, because everything else is moot if it cannot. Then pick your most valuable technical page and read it as chunks: if you cut it into paragraphs and showed each one alone, how many would answer a real question on their own? That number is roughly your surface area in this engine. And if you sell to developers or B2B buyers, take the community layer seriously: on the published data, Perplexity trusts a good Reddit thread more than it trusts your homepage.
