Orphan pages, click depth, link graph

Find the pages your site buries

Paste a URL. Get the pages worth investigating, the evidence behind each finding, and what to do next. Add your Search Console data to see where real search demand meets a crawl problem. No signup to start.

Free 20-page sample; 50 with email. Small sitemaps (up to 150 pages) get a complete crawl.

SitemapIllustrative example
Example sitemap
//blog/products/about/contact/blog/category/products/reviews/careers/faq/faq/billing
sitemap.digital crawl log
Useful actionsbefore the detailed map
No signupto start scanning
11 bot rulessearch, fetch and training
Private CSVsprocessed in your browser
What you get

From a crawl result to a practical action

Start with what the scan can establish. Add your business priorities and search evidence when you have them.

Know what to fix first

Start with up to three priority actions: the affected pages, evidence, and a practical next step. Open the rest when you need the detail.

Add real search demand

Import a Search Console Pages CSV to put clicks and impressions beside crawl findings. Add a comparable earlier period to investigate losses. Your export stays in this browser.

Make a useful handoff

Explore the map, share the public crawl, or export a report for your developer. Re-scan after a change to see which crawl findings remain.

How the scan works

Paste a URL. Watch the map build itself.

No installs, no config. The crawl starts the moment you submit, and results stream in as each page is checked.

01 / CRAWL

Live site crawl

We read the sitemap and fetch pages, homepage first. Sitemaps listing up to 150 pages get a complete crawl. Larger sites are sampled to 20 pages, or 50 with email. Without a usable sitemap, we follow internal links to that same tier limit.

02 / CHECK

Per-page analysis

For every page, the scan reads the title tag, meta description, H1 heading, canonical link, and robots meta rules, then flags anything that is missing, duplicated, or likely to hurt how the page is understood.

03 / ACT

An evidence-led next step

See the highest-priority fixes, choose the pages that matter to your business, and optionally add private search data. Crawl evidence suggests what to investigate; it cannot promise higher rankings.

The AI-readiness score

Built for the age of AI search

example.comIllustrative example
82
AI readiness score

Category weights · maximum points

Content
50
Access
10
Structure
20
Metadata
10
Schema
10

Is your site readable by AI?

The score summarises raw-HTML content, crawler access, structure, metadata and schema checks. It helps explain what our crawler could read. It does not measure your rankings, confirm Google indexing, or predict AI citations. Use the individual findings to decide whether a change is appropriate for that page.

See how the score is calculated

Who it is for

SEOs use sitemap.digital to get a fast, visual read on any site's structure before a deeper audit: it surfaces thin sections, broken internal links, and missing metadata in one pass instead of clicking through pages one at a time. It works just as well on a competitor or client site as on your own.

Migration and redesign teams paste the old and new URLs to compare structure before and after a move, spotting orphaned pages and broken hierarchy before launch. IA, UX, and development teams use it to understand an unfamiliar codebase or content set instantly, then share the map in a single link.

Indie developers and content teams can check a public production URL before or after a change. Review crawler rules in context: allowing search access and allowing model training are separate choices. Do not submit private staging content; anyone with a scan link can read its crawl results.

Free tools and guides

Single-purpose checks and cornerstone guides

Free tools for when you need one answer fast, plus guides covering every check sitemap.digital runs on a scan.

Free tools

Learn: guides on AI crawlability

Guide

Orphan pages in SEO: what they are and how to fix them

What an orphan page is, how sitemap.digital finds one from real crawl data, and the fastest way to fix it before your next scan.

Guide

Click depth in SEO: what it is and why it matters

What click depth means in SEO, why pages more than four hops from the homepage get flagged, and how to flatten a site so more of it gets crawled.

Guide

Internal links and the internal link graph for AI crawlers

How the internal link graph controls what GPTBot, ClaudeBot, and PerplexityBot can find on your site, and the audit checklist to fix it.

Guide

llms.txt: what it is and whether it matters

The llms.txt format explained: what it is meant to do, a real worked example, and an honest, evidence based look at whether it currently changes anything.

Guide

AI crawlability score: how it is calculated

How sitemap.digital scores a page for AI crawlability out of 100: content 50, structure 20, access 10, metadata 10, schema 10.

Guide

Why your site is not showing up in ChatGPT and AI answers

Your pages are indexed in Google, yet ChatGPT and other AI answers never mention your site. Here is why that gap exists and what to check.

Guide

Best AI Search Visibility and Monitoring Tools (2026)

Compare 9 AI search visibility tools spanning free crawlability checks to brand mention trackers like Profound, Otterly and Peec. Honest pricing and limits.

Guide

How to Track Brand Mentions in AI Search (Manual + Tools)

A practical method for tracking whether ChatGPT, Perplexity and Google AI Mode mention your brand. Manual prompts, metrics, tools and a monthly cadence.

Guide

AI Crawlers: The Complete List of Major AI Bots (2026)

Every major AI crawler explained by vendor and job: GPTBot, ClaudeBot, PerplexityBot, Bytespider, CCBot and Google-Extended. Includes a robots.txt example.

Guide

ClaudeBot: Anthropic's Crawler and How to Control It

ClaudeBot is Anthropic's training crawler. Learn its user agent, how it differs from Claude-SearchBot and Claude-User, and how to block it safely.

Guide

GPTBot Explained: OpenAI's Crawlers and How to Control Them

GPTBot, OAI-SearchBot and ChatGPT-User are three separate OpenAI crawlers with different jobs. See the UA tokens and the robots.txt rules that control each one.

Guide

PerplexityBot: Perplexity's Crawler and How to Control It

Learn what PerplexityBot is, how to verify it with IP ranges, and how to block it in robots.txt without losing Perplexity's citation traffic.

Guide

Bytespider: ByteDance's AI crawler and how to block it

Bytespider is ByteDance's AI crawler with no published documentation, verification or IP ranges. Learn what it does and how to block it reliably.

FAQ

Frequently asked questions

Is it free?

Yes. Start without an account and get explained fixes, affected pages and a shareable crawl. On larger sites the free sample is 20 pages, or 50 with an email unlock. Sites with a usable sitemap listing up to 150 pages are scanned to completion. Search Console CSV imports stay in your browser and are free during the preview.

How many pages does it scan?

A usable sitemap listing up to 150 pages is scanned to completion. Larger sitemaps are sampled: homepage first, then sitemap order, up to 20 pages anonymously or 50 with an email unlock. Without a usable sitemap we follow links breadth-first to the same tier limit. Sitemap order does not prove which pages matter to your business. The report shows its coverage and withholds conclusions that require a complete crawl.

Do I need to sign up?

No. Paste a URL and the scan starts straight away, no account and no email needed. Viewing the results, sharing the link, and comparing a later scan are all free. An email is only requested if you want a bigger 50-page scan or to export the CSV or PDF report.

What is an AI-readiness score?

It is our diagnostic score for raw-HTML content, crawler access, structure, metadata and schema. It is not a search ranking, indexing check or prediction that an AI service will cite you. Use the underlying evidence and action list, not the score alone. An llms.txt file is optional and is not part of the score.

Which AI crawlers do you check?

We report robots rules for search, user-fetch and training agents separately, including OAI-SearchBot, Claude-SearchBot, PerplexityBot, GPTBot and ClaudeBot. Blocking model training is a policy choice, not automatically a growth problem. Google-Extended is a control token, not a separate crawling agent.

Find your next useful fix

No signup, no config. Start a free crawl, check its coverage, and work through the priority actions. Bring search data when you want a stronger basis for what comes next.