AI crawlers in robots.txt
GPTBot, Google-Extended and PerplexityBot must be allowed to fetch your pages. Blocking them hides you from AI search entirely.
AI crawlers read your raw HTML — not a rendered browser. Run a free 30-second check and see which of ten technical factors stop ChatGPT, Perplexity and Google AI Overviews from reading and citing your site.

Most local business sites score poorly because content lives behind JavaScript, AI bots are blocked in robots.txt, or there is no llms.txt and FAQ schema. This tool shows exactly what is failing — the same signals we fix with GEO andentity pages.
Functional tool: enter a website URL to score ten technical factors that affect whether AI crawlers can read and cite your site.
Inspired by how answer engines classify pages — adapted for UK trades, professional services and local SEO. Ecommerce stores: use the sister checker onloudcrowd.agency.
GPTBot, Google-Extended and PerplexityBot must be allowed to fetch your pages. Blocking them hides you from AI search entirely.
Machine-readable manifests point AI systems to your entity facts, service pages and citation URLs — the same pattern we use on /info-for-ai/ and /llms.txt.
JSON-LD Organization, Service and FAQPage schema give AI extractable facts. Semantic FAQ lists (<dl>) help Google AI Overviews and ChatGPT cite direct answers.
AI crawlers often skip JavaScript. If your H1, body copy and schema only exist after client-side rendering, you are invisible — regardless of how good the content is.
Google’s web.dev guidance covers AI agents that browse and act on your site. Our checker scores the crawlability layer those agents and answer engines share — raw HTML, robots.txt, llms.txt and schema.
See Google’s overview:Build agent-friendly websites(web.dev). Below is how that advice maps to our free checker and GEO work.
| Agent-friendly guidance | LoudCrowd checker & GEO |
|---|---|
| Readable content in raw HTML (not JS-only shells) | Readable text in HTML — H1, title and body copy visible without executing JavaScript |
| Machine-readable entity and page signals | llms.txt, ai.txt, JSON-LD Organization and FAQ schema |
| AI crawlers allowed to fetch pages | robots.txt — GPTBot, Google-Extended, PerplexityBot not blocked |
| Accessibility tree: roles, names and labelled inputs | Not scored automatically — we fix via semantic HTML, <label for>, real <button>/<a> CTAs and stable layouts in web design |
| Stable CTAs and semantic actions (not divs pretending to be buttons) | GEO + conversion-focused web design — consistent contact paths agents and humans can follow |
Technical crawlability is step one. GEO adds extractable FAQs, reviews, list inclusion and entity consistency so AI systems recommend you — not just read you.
A page can rank in Google and still be invisible in ChatGPT. Google renders JavaScript; many LLM crawlers do not. If your service pages are built in React, Wix or a heavy page builder without server-side HTML, AI systems may never see your H1, FAQs or schema.
Technical readiness is step one. Step two is GEO — entity clarity, reviews, list inclusion, extractable FAQs and consistent facts across your domain. LoudCrowd is Sheffield-based and works with UK local businesses only.
Direct answers on what this tool does and how it relates to GEO.
Page updated
Get a free marketing review — no hard sell, just an honest conversation about SEO, GEO or web design for your trade.