Paste any URL. In under 10 seconds, find out which AI engines can reach, read, and cite your site, and what's quietly blocking them.
Free · No account · Nothing stored
The problem
Even "modern" SEO setups can be silently blocking the crawlers that drive AI citations, usually because a plugin shipped a "block AI bots" toggle defaulted on. Your rankings suffer and you have no idea.
# Auto-generated by SEO plugin User-agent: * Disallow: / # ← blocks everything, incl. AI search
# Intentionally configured User-agent: OAI-SearchBot Allow: / User-agent: Claude-SearchBot Allow: / # ← AI search crawlers allowed
It's simple
No sign-up, no profile, no password. Paste a URL, get your full results immediately. We check; you learn; you go.
Scans run fresh every time. Nothing about your site is persisted between sessions. Stateless by design. Nothing persists to leak.
Score, Crawler Gate, and six sub-scores are free. Enter your email on the results page to unlock the full prioritized fix list: copy-paste robots.txt rules, JSON-LD snippets, and exact steps.
The crawler gate
Every AI platform sends specialized crawlers to index content for citations. Your robots.txt decides who gets through. Most sites haven't configured it intentionally.
› Blocking training crawlers (GPTBot, ClaudeBot) is a legitimate choice and does not affect citations. Only blocking search/retrieval bots costs you AI visibility.
What we check
The gate. robots.txt per-bot verdicts, server-level block detection (WAF/CDN comparison), sitemap reachability, and meta robots signals. If crawlers can't reach you, nothing else matters.
The signals AI weighs most when deciding who to cite. sameAs links to Wikipedia, Wikidata, LinkedIn; named author attribution; Wikidata knowledge graph presence; About page.
How confidently AI models parse and attribute your content. JSON-LD presence and coverage of high-value types: Organization, FAQPage, Article, WebSite, and more.
Can AI crawlers actually read your content? JavaScript-rendering dependence and login/paywall gating, two ways a "reachable" site is still effectively invisible.
Is your content quotable? Heading hierarchy, answer-first structure, question-style headings, lists and tables: the formats AI engines cite most readily.
Date signals, llms.txt, canonical tags, HTTPS, response speed, and sitemap lastmod. Recency and technical cleanliness matter, especially for time-sensitive citations.
Learn
Five guides that cover the actual mechanics — not another list of tips.
FAQ
ChatGPT uses OAI-SearchBot to index content for real-time citations. This is a different crawler from GPTBot (which is for training). If your robots.txt blocks OAI-SearchBot — or has a wildcard rule that catches it — ChatGPT cannot cite your site, regardless of how relevant your content is. Paste your URL above to check in under 10 seconds.
No — not directly. GPTBot is OpenAI's training crawler. Blocking it limits what GPT models learn during pre-training, but has no effect on real-time citations. The citation pipeline uses OAI-SearchBot, which is a separate user agent. Many sites that blocked GPTBot (a defensible choice) accidentally also blocked OAI-SearchBot because they used wildcard disallow rules. We check both.
The most common causes, in order of frequency:
We check all four as part of the free scan.
GEO (Generative Engine Optimization) is the practice of making sure AI search engines can find, read, and cite your content. It's related to SEO but distinct in important ways. Traditional SEO asks: can Google rank this page? GEO asks: can AI systems extract a citable answer from it? The technical requirements differ — AI crawlers often don't render JavaScript, entity authority matters more than backlinks, and content structure for quotability is different from content structure for dwell time. Our GEO guide walks through all six signals.
To be citable by each major platform, you need to allow these retrieval crawlers:
OAI-SearchBot — ChatGPT citationsClaude-SearchBot — Claude citationsPerplexityBot — Perplexity citationsGooglebot — Gemini / Google AI OverviewsBingbot — Microsoft CopilotThe training crawlers (GPTBot, ClaudeBot, Google-Extended) are separate. Blocking those is fine; it only affects model training, not citations. See our full robots.txt guide.
llms.txt is a plain-text file at yourdomain.com/llms.txt that gives AI systems a structured summary of your site: what it does, who runs it, and what the key pages are. It was proposed in 2024 and is increasingly checked by Perplexity and other AI engines during indexing. Creating one takes under 20 minutes and contributes to your freshness and hygiene score. It won't fix access or readability problems on its own, but it's worth having once the bigger issues are resolved. See our llms.txt guide.
No. letthebots.in is stateless by design. Scans run fresh every time and no site data is persisted between sessions. If you enter your email to unlock the full fix guide, your email and scan score are stored in our email platform (Kit), but the scan results themselves are not. We run 6 live checks every time you scan — nothing is cached or stored about the site.
A score of 85 or above earns the Citation-Ready badge. It means your site passes all critical GEO signals: AI crawlers can reach you, your content is readable without JavaScript rendering, you have JSON-LD structured data covering key types, your entity is identifiable (sameAs links, authorship), your content is structured for extraction, and your date signals are current. Sites at this level are well-positioned to be cited when AI engines encounter relevant queries.
The full picture
Paste any URL above. The score, Crawler Gate, and six category sub-scores are free. Enter your email on the results page to unlock the prioritized fix guide for every deduction, with copy-paste robots.txt rules and JSON-LD snippets.
Check your site →No account. Nothing stored. Score 85+ to earn the Citation-Ready badge.