Dottly AI
  • Features
  • Pricing
  • Blog
  • Docs
  • About
Dottly AI
How to Find All Pages on a Website: A Defensible SEO Inventory
2026/09/09

How to Find All Pages on a Website: A Defensible SEO Inventory

Learn how to inventory a website with sitemaps, crawlers, links, logs, and Search Console without confusing discovery with indexing.

Free AI visibility check

See where AI recommends your brand

Review the buyer questions, AI answers, competitors, and available sources shaping your visibility.

Dottly AI
  • 6 buyer questions
  • Answers and available source evidence
  • No credit card to start
Run the free check

There is no single search query that reliably returns every page on a website. A defensible inventory combines several sets: URLs the CMS publishes, URLs listed in sitemaps, URLs a crawler can reach, URLs observed in logs, and URLs search engines report as discovered or indexed. The differences are the useful part of the audit because they reveal orphan pages, stale routes, redirects, and gaps in your measurement.

Define “all pages” before you count

Teams often compare numbers that describe different things. Write down the set you need before opening a tool:

  • Published pages: records the CMS or application considers public.
  • Served routes: URLs that respond from production, including utility and legacy routes.
  • Discoverable pages: URLs exposed by links, sitemaps, feeds, or other crawl paths.
  • Crawlable pages: URLs a crawler reaches under a stated starting point, user agent, and rendering configuration.
  • Indexed pages: URLs a search engine reports as eligible or indexed.
  • Content assets: PDFs, images, videos, and files that may not behave like HTML pages.

These sets overlap without being identical. A private CMS record may never be served. An orphan landing page may be served but not linked. A URL can be crawled and then excluded from an index. Google's page-discovery guidance explicitly separates what a site contains from what Google has tried to crawl and index.

Start with first-party inventory data

Export the canonical records from your CMS, repository, route manifest, or database. Include the fields needed to explain a URL later:

FieldWhy it matters
URL and localeIdentifies the exact route and language family
Publication stateSeparates draft, scheduled, private, and public records
Canonical URLShows the preferred address when aliases exist
Last meaningful changeSupports review and sitemap dates
Content type and ownerRoutes fixes to the right team
Source file or record IDMakes the inventory reproducible

Normalize host, protocol, trailing slash, case, and query parameters before deduplicating. Keep the raw export separate from the normalized view so a future audit can reproduce how a count was produced. Do not silently discard rows that fail normalization; label them for review.

Add the sitemap set

Fetch the public sitemap and every file referenced by a sitemap index. Parse the URL set, then compare it with the first-party export. The sitemap should usually contain canonical pages you want search engines to discover, not every route your application can serve. Google's sitemap overview describes it as a file that provides information about pages and other resources; it is not a complete index ledger.

Record these mismatch types:

  • published canonical URL missing from the sitemap;
  • sitemap URL not present in the publishing system;
  • sitemap URL redirecting elsewhere;
  • sitemap URL blocked or marked noindex;
  • duplicate or non-canonical locale variant; and
  • stale URL that is no longer served.

For deeper validation rules, consult the sitemap best-practices workflow. Fix the underlying generator or publication rule rather than editing generated XML by hand.

Crawl from more than the homepage

A crawler reveals what is reachable under a defined configuration. Seed the crawl with the homepage, primary navigation, XML sitemap, RSS feeds, and important category or documentation hubs. Document the run parameters: user agent, JavaScript rendering, blocked resources, authentication, depth, and URL exclusions.

Use the crawl to collect status code, canonical, robots directive, title, heading, content type, language, and incoming-link counts. A rendered crawl may discover routes that a raw HTML crawl misses, while a raw crawl can expose server-side issues hidden by client rendering. Run both when the application is heavily client-rendered.

The result is a crawl set, not a claim that every page exists. It cannot find an orphan route with no supplied seed, a private backend record, or a URL blocked before the crawler can inspect it. Pair the work with crawling tools for SEO and keep the configuration in the audit record.

Use logs to find real requests

Server or edge logs show which URLs received requests, by which user agent, and with what response. Filter bot traffic separately from humans, monitoring services, previews, and attack noise. Logs can reveal:

  • search bots requesting a URL absent from your sitemap;
  • repeated requests to obsolete parameters or redirects;
  • pages receiving internal traffic but no crawl visits; and
  • unexpected locale, host, or protocol variants.

Logs do not enumerate pages that have never been requested. Treat them as observed behavior, not a complete inventory. Retain the collection window and sampling rules so comparisons across months remain meaningful.

Add Search Console and index evidence

Search Console reports are valuable because they show a search engine's perspective, but they are not a database export of your entire site. Use URL Inspection and Page Indexing views for representative samples and important anomalies. Compare their states with your first-party, sitemap, crawl, and log sets.

A useful reconciliation table looks like this:

URL stateInterpretationNext action
Published + linked + sitemap + indexedExpected pathMonitor changes
Published + sitemap, no crawl evidenceDiscovery or priority gapCheck links, access, and release timing
Published + crawl, not indexedEligibility or quality issueInspect canonical, robots, and content
Served + no owner recordLegacy or unintended routeAssign owner, redirect, or remove
Indexed + no current source recordStale or migrated URLVerify canonical and migration handling

This model prevents “we found 500 URLs” from becoming a misleading success metric. The number is only useful with a definition, source, timestamp, and known exclusions.

Reconcile and prioritize the gaps

After merging the sets, deduplicate by normalized canonical URL and retain provenance columns such as cms, sitemap, crawl, log, and search-console. Then classify each gap:

  1. Expected difference: private, noindex, utility, or intentionally excluded.
  2. Repair candidate: published but missing a link, sitemap entry, or canonical.
  3. Migration risk: old host, locale, parameter, or redirected URL still receiving traffic.
  4. Unknown: insufficient evidence; collect another sample before changing production.

Prioritize by user and business value, not by raw URL count. A broken canonical on a core product page deserves attention before a low-value parameter family. Keep a decision log so the same exception is not re-litigated every audit.

Account for JavaScript, parameters, and hidden routes

Modern applications can expose several URL populations that a simple HTML crawl misses. Client-side navigation may create links only after rendering, faceted filters may generate large parameter spaces, and APIs may serve content that has no indexable HTML route. Run a rendered crawl when the framework requires it, but compare the result with the raw response so you can see what depends on JavaScript.

Set parameter rules before crawling. Allowing every filter combination can turn an inventory exercise into an unbounded crawl; excluding every query string can hide valuable campaign, search, or pagination behavior. Keep a separate sample of rejected URLs and document why the pattern is outside the canonical page set.

Also inspect route manifests and server configuration for pages that are difficult to discover through content links: status pages, redirects, old campaign routes, downloadable files, and framework-generated endpoints. They may not belong in the SEO inventory, but assigning them a type and owner prevents them from becoming unexplained production surface.

Make the inventory repeatable

Schedule the sources that change frequently and keep a dated snapshot of each result. At minimum, retain:

  • the normalized first-party export;
  • the fetched sitemap URLs;
  • crawl configuration and output;
  • the log window and bot filters;
  • sampled Search Console states; and
  • the merge rules and unresolved gaps.

Run the inventory after a domain, routing, CMS, locale, or navigation migration. For normal publishing, a lighter daily or release-based check can catch missing sitemap entries and broken links before they accumulate.

Connect technical inventory to visibility questions

Finding a page is not the same as proving it is visible in search or AI answers. Once the URL set is trustworthy, you can investigate whether important pages are indexed, cited, or represented in sampled answers. Dottly AI's documentation can support the measurement handoff, but keep the evidence lanes separate: crawl data explains access, while answer-level evidence explains what a model actually returned.

FAQ

Does site:example.com show every page?

No. It is a useful spot check, not an exhaustive inventory. Search-engine result pages are sampled and can omit URLs for many reasons.

Is the sitemap the complete list of pages?

No. It is a curated discovery set. It may intentionally omit private, low-value, duplicate, or non-HTML routes, and it cannot list URLs the publishing system does not know about.

Should I crawl the sitemap or the homepage first?

Use both. The homepage and navigation show link reachability; the sitemap exposes intended canonical URLs. Their difference is often the fastest way to find orphan or stale pages.

How often should I build a page inventory?

Use a release-triggered check for routing and content changes, plus a scheduled reconciliation for logs and search data. Increase frequency during migrations or incident response.

The reliable answer to “how many pages do we have?” is a documented set comparison, not a single number. When every URL has provenance and an owner, the inventory becomes a decision tool for SEO maintenance rather than a one-time crawl report.

Continue with related guides

  • What Is Generative Engine Optimization (GEO)?
  • How to Get Cited in AI Search: A Practical Source Guide
  • AI Visibility Report Metrics Explained
All Posts
Free AI visibility check

See where AI recommends your brand

Review the buyer questions, AI answers, competitors, and available sources shaping your visibility.

Dottly AI
  • 6 buyer questions
  • Answers and available source evidence
  • No credit card to start
Run the free check

Author

avatar for Dottly AI Team
Dottly AI Team

Categories

    Define “all pages” before you countStart with first-party inventory dataAdd the sitemap setCrawl from more than the homepageUse logs to find real requestsAdd Search Console and index evidenceReconcile and prioritize the gapsAccount for JavaScript, parameters, and hidden routesMake the inventory repeatableConnect technical inventory to visibility questionsFAQDoes site:example.com show every page?Is the sitemap the complete list of pages?Should I crawl the sitemap or the homepage first?How often should I build a page inventory?

    More Posts

    How to Submit Your Site to Search Engines: A Modern Workflow

    How to Submit Your Site to Search Engines: A Modern Workflow

    Learn how to submit a website to Google, Bing, and participating search engines with verification, sitemaps, and post-launch checks.

    avatar for Dottly AI Team
    Dottly AI Team
    2026/09/09
    Sitemap SEO Best Practices: An Audit and Maintenance Framework

    Sitemap SEO Best Practices: An Audit and Maintenance Framework

    Learn what belongs in an XML sitemap, how to audit it, and how to connect sitemap data with crawling and indexing evidence.

    avatar for Dottly AI Team
    Dottly AI Team
    2026/09/09
    Backlink Software: What to Compare Before You Choose
    GEO GuidesProduct Guides

    Backlink Software: What to Compare Before You Choose

    Compare backlink software by coverage, link data, workflow, and risk controls. Use a small pilot to find a tool that fits your SEO work.

    avatar for Dottly AI Team
    Dottly AI Team
    2026/09/29

    Newsletter

    Join the community

    Subscribe to our newsletter for the latest news and updates

    Dottly AI

    Monitor how ChatGPT, Gemini and Grok talk about your brand.

    Product
    • Features
    • Pricing
    • FAQ
    Resources
    • Blog
    • Documentation
    Company
    • About
    • Contact
    Legal
    • Cookie Policy
    • Privacy Policy
    • Terms of Service
    © 2026 Dottly AI. All Rights Reserved.

    DOTTLY AI