
How to Find All Pages on a Website: A Defensible SEO Inventory
Learn how to inventory a website with sitemaps, crawlers, links, logs, and Search Console without confusing discovery with indexing.
There is no single search query that reliably returns every page on a website. A defensible inventory combines several sets: URLs the CMS publishes, URLs listed in sitemaps, URLs a crawler can reach, URLs observed in logs, and URLs search engines report as discovered or indexed. The differences are the useful part of the audit because they reveal orphan pages, stale routes, redirects, and gaps in your measurement.
Define “all pages” before you count
Teams often compare numbers that describe different things. Write down the set you need before opening a tool:
- Published pages: records the CMS or application considers public.
- Served routes: URLs that respond from production, including utility and legacy routes.
- Discoverable pages: URLs exposed by links, sitemaps, feeds, or other crawl paths.
- Crawlable pages: URLs a crawler reaches under a stated starting point, user agent, and rendering configuration.
- Indexed pages: URLs a search engine reports as eligible or indexed.
- Content assets: PDFs, images, videos, and files that may not behave like HTML pages.
These sets overlap without being identical. A private CMS record may never be served. An orphan landing page may be served but not linked. A URL can be crawled and then excluded from an index. Google's page-discovery guidance explicitly separates what a site contains from what Google has tried to crawl and index.
Start with first-party inventory data
Export the canonical records from your CMS, repository, route manifest, or database. Include the fields needed to explain a URL later:
| Field | Why it matters |
|---|---|
| URL and locale | Identifies the exact route and language family |
| Publication state | Separates draft, scheduled, private, and public records |
| Canonical URL | Shows the preferred address when aliases exist |
| Last meaningful change | Supports review and sitemap dates |
| Content type and owner | Routes fixes to the right team |
| Source file or record ID | Makes the inventory reproducible |
Normalize host, protocol, trailing slash, case, and query parameters before deduplicating. Keep the raw export separate from the normalized view so a future audit can reproduce how a count was produced. Do not silently discard rows that fail normalization; label them for review.
Add the sitemap set
Fetch the public sitemap and every file referenced by a sitemap index. Parse the URL set, then compare it with the first-party export. The sitemap should usually contain canonical pages you want search engines to discover, not every route your application can serve. Google's sitemap overview describes it as a file that provides information about pages and other resources; it is not a complete index ledger.
Record these mismatch types:
- published canonical URL missing from the sitemap;
- sitemap URL not present in the publishing system;
- sitemap URL redirecting elsewhere;
- sitemap URL blocked or marked
noindex; - duplicate or non-canonical locale variant; and
- stale URL that is no longer served.
For deeper validation rules, consult the sitemap best-practices workflow. Fix the underlying generator or publication rule rather than editing generated XML by hand.
Crawl from more than the homepage
A crawler reveals what is reachable under a defined configuration. Seed the crawl with the homepage, primary navigation, XML sitemap, RSS feeds, and important category or documentation hubs. Document the run parameters: user agent, JavaScript rendering, blocked resources, authentication, depth, and URL exclusions.
Use the crawl to collect status code, canonical, robots directive, title, heading, content type, language, and incoming-link counts. A rendered crawl may discover routes that a raw HTML crawl misses, while a raw crawl can expose server-side issues hidden by client rendering. Run both when the application is heavily client-rendered.
The result is a crawl set, not a claim that every page exists. It cannot find an orphan route with no supplied seed, a private backend record, or a URL blocked before the crawler can inspect it. Pair the work with crawling tools for SEO and keep the configuration in the audit record.
Use logs to find real requests
Server or edge logs show which URLs received requests, by which user agent, and with what response. Filter bot traffic separately from humans, monitoring services, previews, and attack noise. Logs can reveal:
- search bots requesting a URL absent from your sitemap;
- repeated requests to obsolete parameters or redirects;
- pages receiving internal traffic but no crawl visits; and
- unexpected locale, host, or protocol variants.
Logs do not enumerate pages that have never been requested. Treat them as observed behavior, not a complete inventory. Retain the collection window and sampling rules so comparisons across months remain meaningful.
Add Search Console and index evidence
Search Console reports are valuable because they show a search engine's perspective, but they are not a database export of your entire site. Use URL Inspection and Page Indexing views for representative samples and important anomalies. Compare their states with your first-party, sitemap, crawl, and log sets.
A useful reconciliation table looks like this:
| URL state | Interpretation | Next action |
|---|---|---|
| Published + linked + sitemap + indexed | Expected path | Monitor changes |
| Published + sitemap, no crawl evidence | Discovery or priority gap | Check links, access, and release timing |
| Published + crawl, not indexed | Eligibility or quality issue | Inspect canonical, robots, and content |
| Served + no owner record | Legacy or unintended route | Assign owner, redirect, or remove |
| Indexed + no current source record | Stale or migrated URL | Verify canonical and migration handling |
This model prevents “we found 500 URLs” from becoming a misleading success metric. The number is only useful with a definition, source, timestamp, and known exclusions.
Reconcile and prioritize the gaps
After merging the sets, deduplicate by normalized canonical URL and retain provenance columns such as cms, sitemap, crawl, log, and search-console. Then classify each gap:
- Expected difference: private, noindex, utility, or intentionally excluded.
- Repair candidate: published but missing a link, sitemap entry, or canonical.
- Migration risk: old host, locale, parameter, or redirected URL still receiving traffic.
- Unknown: insufficient evidence; collect another sample before changing production.
Prioritize by user and business value, not by raw URL count. A broken canonical on a core product page deserves attention before a low-value parameter family. Keep a decision log so the same exception is not re-litigated every audit.
Account for JavaScript, parameters, and hidden routes
Modern applications can expose several URL populations that a simple HTML crawl misses. Client-side navigation may create links only after rendering, faceted filters may generate large parameter spaces, and APIs may serve content that has no indexable HTML route. Run a rendered crawl when the framework requires it, but compare the result with the raw response so you can see what depends on JavaScript.
Set parameter rules before crawling. Allowing every filter combination can turn an inventory exercise into an unbounded crawl; excluding every query string can hide valuable campaign, search, or pagination behavior. Keep a separate sample of rejected URLs and document why the pattern is outside the canonical page set.
Also inspect route manifests and server configuration for pages that are difficult to discover through content links: status pages, redirects, old campaign routes, downloadable files, and framework-generated endpoints. They may not belong in the SEO inventory, but assigning them a type and owner prevents them from becoming unexplained production surface.
Make the inventory repeatable
Schedule the sources that change frequently and keep a dated snapshot of each result. At minimum, retain:
- the normalized first-party export;
- the fetched sitemap URLs;
- crawl configuration and output;
- the log window and bot filters;
- sampled Search Console states; and
- the merge rules and unresolved gaps.
Run the inventory after a domain, routing, CMS, locale, or navigation migration. For normal publishing, a lighter daily or release-based check can catch missing sitemap entries and broken links before they accumulate.
Connect technical inventory to visibility questions
Finding a page is not the same as proving it is visible in search or AI answers. Once the URL set is trustworthy, you can investigate whether important pages are indexed, cited, or represented in sampled answers. Dottly AI's documentation can support the measurement handoff, but keep the evidence lanes separate: crawl data explains access, while answer-level evidence explains what a model actually returned.
FAQ
Does site:example.com show every page?
No. It is a useful spot check, not an exhaustive inventory. Search-engine result pages are sampled and can omit URLs for many reasons.
Is the sitemap the complete list of pages?
No. It is a curated discovery set. It may intentionally omit private, low-value, duplicate, or non-HTML routes, and it cannot list URLs the publishing system does not know about.
Should I crawl the sitemap or the homepage first?
Use both. The homepage and navigation show link reachability; the sitemap exposes intended canonical URLs. Their difference is often the fastest way to find orphan or stale pages.
How often should I build a page inventory?
Use a release-triggered check for routing and content changes, plus a scheduled reconciliation for logs and search data. Increase frequency during migrations or incident response.
The reliable answer to “how many pages do we have?” is a documented set comparison, not a single number. When every URL has provenance and an owner, the inventory becomes a decision tool for SEO maintenance rather than a one-time crawl report.
Continue with related guides
Author

Categories
site:example.com show every page?Is the sitemap the complete list of pages?Should I crawl the sitemap or the homepage first?How often should I build a page inventory?More Posts

How to Submit Your Site to Search Engines: A Modern Workflow
Learn how to submit a website to Google, Bing, and participating search engines with verification, sitemaps, and post-launch checks.


Sitemap SEO Best Practices: An Audit and Maintenance Framework
Learn what belongs in an XML sitemap, how to audit it, and how to connect sitemap data with crawling and indexing evidence.


Backlink Software: What to Compare Before You Choose
Compare backlink software by coverage, link data, workflow, and risk controls. Use a small pilot to find a tool that fits your SEO work.

Newsletter
Join the community
Subscribe to our newsletter for the latest news and updates
