TechnicalSEOPlaybook

Technical SEO Checklist for AI Search (2026)

A technical SEO checklist in 2026 has to serve two kinds of visitor: search engine crawlers and AI crawlers. The fundamentals are shared, but AI crawlers add three failure points that traditional audits miss: CDN-level bot blocking, JavaScript-dependent content, and missing Bing indexation. This checklist covers both, grouped in the order you should check them.

It goes deeper than the technical pillar of our GEO checklist. If you only have an hour, do sections 1 to 3; they catch most of the problems that make a site invisible to AI engines.

1. Crawl access

  • robots.txt returns 200 and plain text at /robots.txt, not a redirect or an HTML error page.
  • No leftover Disallow: / from a staging environment. This single line is the most expensive typo in SEO.
  • Search and retrieval bots are allowed: Googlebot, Bingbot, OAI-SearchBot, ChatGPT-User, and PerplexityBot.
  • Training bots follow a deliberate policy: GPTBot, Google-Extended, CCBot, ClaudeBot, and Applebot-Extended are allowed or blocked on purpose. See robots.txt for AI crawlers and should you block AI crawlers.
  • CDN and firewall bot settings are reviewed. Cloudflare began blocking known AI crawlers by default for new domains in July 2025, and other providers offer similar switches. robots.txt can say yes while the CDN says no.
  • Rate limiting does not return 403 or 429 to legitimate crawlers during normal crawl volume.
  • Your sitemap is declared in robots.txt.

2. Indexation

  • The XML sitemap lists only canonical, indexable, 200-status URLs. No redirects, no noindexed pages, no 404s.
  • lastmod reflects real content changes, not the build date of every page.
  • Canonical tags self-reference on canonical pages and never point at a redirected or missing URL.
  • noindex appears only where intended, such as thank-you pages and internal search results.
  • Search Console’s page indexing report is reviewed, and “Crawled, currently not indexed” pages are investigated rather than ignored.
  • Bing Webmaster Tools is verified and your sitemap submitted. ChatGPT search leans on Bing’s index; a site absent from Bing is hard for ChatGPT to cite.
  • IndexNow is configured so Bing, Yandex, Seznam, and Naver hear about new and changed URLs immediately.
  • One canonical host. HTTP, www, and trailing-slash variants all 301 to a single form in one hop.

3. Rendering

  • Main content is in the initial HTML response. Fetch a page with curl or view source; if the text is not there, crawlers that do not run JavaScript never see it. See JavaScript rendering and AI crawlers.
  • Internal links are real <a href> elements, not click handlers on other elements.
  • Structured data is in the HTML, not injected only after the page hydrates.
  • Nothing essential depends on an interaction. Tabs and accordions are fine if their content is in the DOM on load.
  • Lazy loading is limited to below-the-fold media, never main copy.

4. Speed and stability

  • Core Web Vitals pass at the 75th percentile on mobile: LCP within 2.5 seconds, INP within 200 milliseconds, CLS within 0.1. See Core Web Vitals explained.
  • Server response time is fast and consistent. Slow responses reduce how much of your site any crawler fetches in a visit.
  • Images are sized, compressed, and in modern formats, with explicit dimensions to prevent layout shift.
  • Web fonts do not block rendering.

5. Structured data

  • An Organization node with a stable @id, plus WebSite and a WebPage per page, connected by reference. See Organization schema for AI search.
  • Page-type markup on the right templates: BlogPosting or Article, Service, Product, LocalBusiness, BreadcrumbList.
  • Markup matches visible content. Google’s structured data policies require it, and mismatches erode trust.
  • It validates. Use the Schema Markup Validator for everything and Google’s Rich Results Test for rich-result types.
  • Expectations are current. FAQ rich results have been limited to well-known government and health sites since August 2023, and HowTo rich results were removed the same year. Mark up for understanding, not for stars. See FAQ schema for AI answers.
  • Important pages sit within three clicks of the homepage.
  • No orphan pages. Every indexable page has at least one internal link pointing to it.
  • Anchor text describes the destination. “Click here” tells a crawler nothing.
  • Redirect chains are at most one hop, and there are no loops.
  • Removed content returns 404 or 410, or 301s to a genuine equivalent. No soft 404s that return 200 with an error message.
  • URLs are stable, and any change ships with a redirect. For site-wide changes, use the website migration SEO checklist.

7. Page-level semantics

  • One H1 per page, and heading levels in order.
  • Layout uses landmark elements: <header>, <nav>, <main>, <article> and <footer>, so a parser can separate the content from the chrome around it.
  • Tabular data uses real <table> markup with header cells, not styled divs.
  • The lang attribute is set, with hreflang if you publish in more than one language.
  • Titles are unique and under about 60 characters; meta descriptions are unique.
  • Meaningful images have descriptive alt text; purely decorative images use empty alt. See image SEO for AI search.

8. Monitoring

  • Server logs are reviewed monthly for crawler activity by user-agent. Verify genuine bots against the IP ranges their operators publish, since user-agent strings are trivial to fake.
  • Search Console crawl stats and indexing reports are checked monthly.
  • Bing Webmaster Tools crawl and AI performance reports are checked where available.
  • Uptime monitoring alerts on 5xx errors.
  • The site is re-crawled after every release with a desktop crawler to catch regressions before search engines do.

Test what each crawler actually receives

The checklist above tells you what should be true. This is how to prove it in five minutes, without a paid tool. Request a page as each user-agent and look at the status code and whether your main content is in the response:

for ua in "Googlebot" "bingbot" "GPTBot" "OAI-SearchBot" "ClaudeBot" "PerplexityBot"; do
  printf '%-14s ' "$ua"
  curl -s -o /tmp/page.html -w '%{http_code} ' -A "Mozilla/5.0 (compatible; $ua)" https://example.com/your-page/
  grep -c 'a phrase from your main content' /tmp/page.html
done

Read the output as a grid. A 403 or 503 for one bot and 200 for the others is a CDN or firewall rule, not robots.txt. A 200 with a count of 0 means the bot gets the page shell but not the content, which is the JavaScript rendering failure. A 200 with a count of 1 or more for every row is a pass.

Two caveats keep this honest. Some firewalls verify bots by IP address as well as user-agent, so a spoofed request can be treated differently from the real crawler; your server logs are the final word. And a pass here proves eligibility only. Whether the page is indexed still has to be confirmed in Search Console and Bing Webmaster Tools.

The five failures we see most often

FailureSymptomWhere to check
CDN blocking AI botsrobots.txt allows bots, logs show noneCDN bot settings, logs
JavaScript-only contentPage looks fine in a browser, source is emptyView source, curl
Missing from BingGoogle traffic, no ChatGPT citationsBing Webmaster Tools
Staging rules shippedSudden deindexing after a launchrobots.txt, meta robots
Redirect chainsSlow crawling, diluted signalsCrawl report

Where the technical checklist fits

Technical SEO is the floor, not the ceiling. Passing every item here makes you eligible to be crawled, indexed, and retrieved; being cited depends on entity clarity and content quality on top. If you would rather have this audited and fixed for you, or you are planning a rebuild, our web development service ships sites that pass it by default, and our AI SEO service handles the layer above.

Frequently asked questions

What should a technical SEO checklist include in 2026?

Crawl access, indexation, rendering, speed, structured data, site architecture, page-level semantics, and monitoring. In 2026 each area also needs an AI check: whether AI crawlers are allowed through your CDN, whether content renders without JavaScript, and whether you are indexed in Bing as well as Google.

Do AI crawlers need different technical SEO?

Mostly the same fundamentals, with three differences that matter. Many AI crawlers do not execute JavaScript, so server-side rendering matters more. ChatGPT search leans on Bing's index, so Bing indexation matters more. And AI crawlers have their own user-agents, which robots.txt and CDN rules can block without you noticing.

How do I check whether AI crawlers can access my site?

Check three layers: robots.txt rules for each AI user-agent, your CDN or firewall bot settings, and your server logs for successful fetches by those user-agents. Verify genuine bots against the IP ranges their operators publish, since user-agent strings are easy to fake.

Does structured data help with AI search?

It helps engines understand what a page and a brand are, which supports entity recognition and accurate descriptions. It does not guarantee a citation, and it must match the visible content of the page.

What is the 80/20 of technical SEO for AI search?

Four checks catch most of the damage: AI and search bots are not blocked at the CDN, the main content is present in the raw HTML without JavaScript, the page is indexed in both Google and Bing, and no staging noindex or robots rule shipped to production. Everything else on the checklist refines a site that already passes those four.

How often should I run a technical SEO audit?

Run the full checklist quarterly. After every release, re-crawl the site and re-check robots.txt, meta robots, and your top templates, because most technical regressions ship with a deploy.

OK

Olga Kunger

Founder & Lead Strategist, Ambeltek

Olga leads Ambeltek's web development, AI SEO, and GEO work — helping brands rank on Google and get cited by AI engines. More about Olga →

Want structured data that engines trust?

We audit, design and ship connected JSON-LD for your whole site, validated and monitored.