AEDaily SUBSCRIBE
Home / Playbooks / The AEO audit
PLAYBOOK · TECHNICAL Aug 12, 2026 · 7 min read

The AEO audit: 27 checks we run on every site

MK Mara Kovač Editor-in-Chief · 6:00 AM ET 𝕏 in ✉
Editorial illustration: a grid of small square checks, most complete, a few outstanding

Most AEO audits are SEO audits with a new cover page. The genuinely new work sits in three places the old checklist never covered: whether AI crawlers can reach you at all, whether your content survives being cut into passages, and whether you are described consistently everywhere an engine might look. Here is the checklist we run, grouped so you can hand each section to whoever owns it.

⚡TL;DRTwenty-seven checks in four groups: access (can AI crawlers fetch and read you), structure (does your content survive passage extraction), evidence (does each passage carry proof), and presence (are you described consistently off-site). Access failures are the most common and the most expensive, because nothing downstream matters if the crawler never renders your page.

Group one: access, checks 1 to 8

The most expensive failures happen here, and they are invisible in every SEO tool that renders JavaScript for you. Crawlability audits for AI consistently find the same handful of blockers:

ACCESS CHECKS011 to 4: crawler permissionsCheck robots.txt individually for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended. Blanket disallow rules and legacy blocks catch these constantly, and blocking the training crawler while intending to block only the search one is the classic error.025 and 6: rendering without JavaScriptFetch your key pages with curl and read the raw HTML. AI crawlers do not reliably execute JavaScript. If your content only exists after hydration, it does not exist. This is the single most common silent failure.037: firewall and bot managementCloudflare and similar layers now block AI crawlers by default in some configurations. Check the edge, not just robots.txt.048: status codes and redirectsChains, soft 404s and inconsistent canonicals confuse retrieval the same way they confuse indexing, with less tolerance.

Group two: structure, checks 9 to 16

This group asks one question in eight ways: does a paragraph of your page still make sense when it is lifted out and shown alone? Engines retrieve passages, as documented in our brief on passage retrieval, and a passage that depends on the paragraph above it dies in extraction.

CHECKWHAT PASSES
9. Self-contained passagesThe answer to each target question sits in its first two sentences with no back-references
10. Passage lengthKey answers land under about 90 words
11. Heading semanticsOne H1, descriptive H2s that match real questions
12. Extractable tablesAt least one comparison or spec table in real HTML, not an image
13. FAQ blocksPresent with matching schema on pages that answer questions
14. Lists that parseReal list markup, consistent structure per item
15. Entity consistencyThe same canonical description of what you are, everywhere on site
16. Schema coverageOrganization on the homepage, plus Article, Product or FAQ per page type
Structure checks. Format research from HubSpot and Wix Studio identifies statistics, dates, author bios and FAQ blocks as the citation-correlated on-page signals.

Groups three and four: evidence and presence, checks 17 to 27

Evidence, 17 to 22. Does each substantive claim carry a number, a source and a date? Is the last-updated date visible and honest? Is there a named author with credentials? Are your statistics first-party where possible, and cited where not? This group exists because it is the highest-leverage edit anyone has measured: the GEO benchmark recorded a 41% visibility lift from adding quotations and 37% from adding statistics. Check 22 is the one people skip: recency with genuine change, since Seer's recency study shows date-stamp edits without substantive updates do not earn the freshness premium.

Presence, 23 to 27. Where do engines look for your category, and are you there? Check your Wikipedia eligibility, your review-platform profiles, the Reddit and forum threads that rank for your buying questions, your YouTube descriptions and auto-generated transcripts (engines read both), and whether your product description is identical across all of them. This group is uncomfortable because most of it is not on your website, which is exactly why it is where most of the citation weight lives.

What to do this week

Run group one today. It takes twenty minutes, it is the cheapest fix on this list, and a surprising number of sites fail check five without knowing it: curl your three most valuable pages and read what comes back. If the content is not in the raw HTML, nothing else on this checklist can help you. Then take one revenue page through groups two and three properly before touching anything else. Twenty-seven checks across a whole site is a quarter of work. Twenty-seven checks on the page that matters most is an afternoon.

KEY TAKEAWAYS01Access failures are the most expensive and least visible. Curl your key pages and read the raw HTML.02Check each AI crawler user-agent individually. Blanket rules and edge firewalls block them silently.03Structure for extraction: self-contained answers under 90 words, real tables, honest headings.04Evidence is the measured lever: 41% from quotations, 37% from statistics in the GEO benchmark.05Most citation weight lives off your domain. Audit review platforms, forums and video descriptions too.