How Google Crawls Websites: From Discovery to Indexing

Santaji GadeTechnical SEOSEO2 weeks ago36 Views

How Google Crawls Websites

Google indexes only 30-60% of a typical site's pages. Learn the full discovery-to-indexing pipeline, the two-wave rendering process, and where pages get lost.

Technical SEO Crawling and Indexing Googlebot Search Console

Google indexes only 30 to 60% of a typical website's pages, according to Google's own John Mueller. The rest sit in some in-between state, crawled but excluded, discovered but never fetched, or fetched but never rendered properly. Understanding the actual pipeline between discovery and indexing explains exactly where most of that other 40 to 70% quietly gets lost.

Google crawls websites through a multi-stage pipeline: discovering URLs through links, sitemaps, and prior crawl history, fetching their content, rendering JavaScript-heavy pages in a headless browser, and finally evaluating that content for inclusion in the index. Crawling and indexing are related but genuinely separate processes.

Advertisement
Advertisement

We covered why Google sometimes skips pages entirely in our crawl budget guide. This article covers the full pipeline that budget gets spent moving through.

30-60%

of a typical site's pages actually make it into Google's index, per John Mueller

2

separate waves handle indexing: a quick HTML pass, then a deferred rendering pass

0

guarantee that crawling a page results in it actually being indexed

The Full Pipeline, Start to Finish

1Discovery
2Crawling / Fetching
3Rendering
4Indexing

The pipeline isn't strictly linear: new links found during rendering can trigger additional crawling, looping back to Stage 1.

Stage 1: How Google Discovers a URL

According to Fudugo's technical deep dive into how Google does indexing, discovery starts from seed URLs: previously crawled pages, known backlinks, submitted XML sitemaps, and pages manually submitted through Search Console.

According to Incremys' guide to understanding Google's site exploration, a sitemap is a discovery and prioritization signal, not an "index now" button. Submitting one never forces crawling or indexing directly.

Stage 2 and 3: Fetching, Then Rendering

Modern Googlebot processes pages in two distinct waves.

First Wave: Quick Crawl

Googlebot extracts what it can directly from the raw HTML, links, metadata, and visible text, without executing any JavaScript yet.

Second Wave: Rendering

A separate, more resource-intensive pass executes JavaScript in a headless browser to see the page the way a real visitor would.

According to SEO Hacker's step-by-step guide to Google crawling, rendering requires significant computational resources, which is why Google prioritizes which pages get rendered based on page importance and freshness requirements.

Stage 4: What Happens During Indexing

According to Incremys' 2026 guide to managing Googlebot, indexing analyzes the fully rendered content to determine its topic and decide whether it earns a place in the index at all.

Why Crawling Doesn't Guarantee Indexing

Search Console StatusWhat Actually Happened
Crawled – currently not indexedPassed discovery and fetch, failed the quality evaluation stage
Discovered – currently not indexedGoogle knows the URL exists but hasn't crawled it yet
Excluded by noindex tagFetched successfully, then explicitly told not to index
Alternate page with proper canonical tagCrawled, but a different URL was chosen as canonical
Page with redirectCrawled, but authority passed to a different destination URL

Crawl-to-Index Diagnostic

Select the Search Console status you're seeing to identify which pipeline stage is the actual problem.

Crawl-to-Index Diagnostic

Matches a Search Console status to the pipeline stage and likely fix

Google knows this URL exists but hasn't crawled it yet. Usually a crawl budget or priority issue. Add strong internal links pointing to it.
Advertisement
Advertisement

Why Understanding Crawling and Indexing Together Matters

According to SEO Kreativ's guide to crawling and indexing explained simply, treating these as one combined process is a common mistake. Every page you want visible must be crawlable, renderable, and indexable, three distinct hurdles rather than one.

According to a detailed explainer on how Google Search works, crawling and indexing together form the foundation the entire ranking system depends on, powering everything from traditional results to AI Overviews and AI Mode.

Internal Linking's Role in the Discovery Stage

According to Incremys' guide to SEO crawling referenced above, internal linking plays a dual role in both crawling and indexing outcomes: it aids discovery of new pages and clarifies the site's overall hierarchy for Google to interpret.

The deeper a page sits, meaning the more clicks required from an entry point to reach it, the harder it becomes to discover and the less frequently it gets revisited once it is crawled and indexed for the first time.

What Slows This Pipeline Down

According to Viacon's beginner's guide to Google crawl versus index, excessive JavaScript reliance without server-side rendering delays the second wave significantly, since the rendering queue is inherently slower than the initial HTML pass.

According to Digital Strategy Force's guide to how Google crawls and indexes, architectural clarity, every URL reachable, renderable, and unambiguous in its canonical identity, matters more to full indexing than content quality alone.

Advertisement
Advertisement

FAQs on How Google Crawls Websites

Does submitting a sitemap guarantee Google will index my pages?
No. A sitemap is a discovery and prioritization signal for crawling and indexing, helping Google find URLs faster, but it never forces either process directly.
Why would Google crawl a page but not index it?
This typically means the page passed discovery and fetching but failed the quality evaluation stage, often due to thin content, duplication, or a mismatch with search intent.
What is the two-wave indexing process?
Google first extracts what it can from raw HTML without executing JavaScript, then runs a separate, more resource-intensive rendering pass later to see the fully executed page.
How long does crawling and indexing take for a new page?
It varies widely, from hours for established, high-authority sites to several weeks for new sites without existing authority or backlinks.
Can blocking CSS or JavaScript in robots.txt hurt crawling and indexing?
Yes. If critical resources needed to render the page correctly are blocked, Google may fail to see the page's actual content, which can degrade both evaluation and indexing.
What's the fastest way to diagnose a crawling and indexing problem?
Check the specific exclusion reason in Search Console's Coverage or Pages report. Each status maps to a distinct stage in the pipeline, pointing directly at where the process broke down.

What We Learn Today

Only 30-60% of a typical site's pages actually get indexed

Discovery starts from sitemaps, links, and Search Console submissions

A sitemap signals priority; it never forces indexing directly

Indexing runs in two waves: quick HTML pass, then deferred rendering

Crawling and indexing are separate; one doesn't guarantee the other

Search Console's exclusion reasons map directly to the pipeline stage

0 Votes: 0 Upvotes, 0 Downvotes (0 Points)

Leave a reply

Loading Next Post...
Search
Popular Now
Loading

Signing-in 3 seconds...

Signing-up 3 seconds...