Technical SEO
Everything in SEO presumes a pipeline most people never look at: Googlebot finds a URL, fetches it, renders it, decides whether it earns a place in the index — and only then does any ranking conversation begin. When pages mysteriously don't rank, the boring explanation is usually here: they were never crawled, never rendered, or never indexed. This guide walks the actual pipeline, stage by stage, with the diagnostics for each.
Stage 1: Discovery — how Google learns a URL exists
Google discovers URLs from links (following the web's graph — the primary path), sitemaps (your declared inventory), and history (URLs it has seen before get revisited). Practical consequences: a page nothing links to relies entirely on the sitemap (fragile), which is why internal linking is a discovery system before it's an equity system — and why new content on well-crawled sites gets found in hours while orphaned pages wait weeks.
Stage 2: Crawling — the fetch
Googlebot requests the URL (mobile user-agent by default, per mobile-first), subject to two gates: permission (robots.txt checked first — disallowed URLs aren't fetched at all) and capacity (crawl rate adapts to your server's health; slow or erroring servers get crawled less — one of several reasons speed compounds). How often any URL gets recrawled tracks its importance signals — links, freshness patterns, sitemap hints — which is why changes to well-linked pages register in days and forgotten corners take months. At the scale where these constraints genuinely bind, that's the crawl budget conversation.
Stage 3: Rendering — the modern extra step
Googlebot renders pages with a real Chromium, executing JavaScript — but rendering is queued separately and can lag the initial crawl. The classic failure: content or links that exist only after client-side JS runs get discovered late or flakily. The working rules: critical content and internal links present in the served HTML (server-side rendering or hydration-friendly frameworks), and the URL Inspection tool's rendered-HTML view as the truth of what Google actually sees — not your browser.
Stage 4: Indexing — the selection
Fetched and rendered is not indexed: Google then decides whether the page merits a slot. It canonicalises (picking one version among duplicates, hopefully agreeing with your canonical tags), applies quality thresholds (thin, boilerplate or near-duplicate pages get "Crawled — currently not indexed", the coverage report's most-asked-about row), and honours directives (noindex removes; so does an accidental one — the second classic self-inflicted wound after robots Disallow). Indexing is also continuous: pages get re-evaluated at recrawl, which is how improvements and devaluations both take effect.
The diagnostic ladder
For any "why isn't this page ranking?" mystery, walk the pipeline in order via URL Inspection in Search Console:
- Known to Google? No → discovery problem: add internal links, check the sitemap includes it.
- Crawled? No, or long ago → check robots.txt, server errors, and whether anything signals the page matters (links again).
- Rendered correctly? View the rendered HTML — missing content means a JS delivery problem.
- Indexed? "Crawled — not indexed" → quality/duplication question: strengthen or consolidate per the cannibalization playbook; "Excluded by noindex/canonical" → verify the directive is intended.
- Indexed but invisible? Now — and only now — it's a ranking question: content, intent, and the authority fee.
Frequently asked questions
How do I get a new page indexed faster?
Internal links from crawled pages (the real accelerator), sitemap inclusion, and URL Inspection's "Request indexing" for individual priority URLs. Beyond that, sitewide crawl frequency is earned — by the site's overall authority and freshness record, not requested.
Why does Google index some pages and refuse others on the same site?
Per-URL quality selection is working as designed — the index is curated, not archival. Persistent "crawled — not indexed" on pages you care about is feedback: differentiate, deepen or consolidate them.
Does crawling frequency affect rankings?
Not directly — but recrawl speed sets how fast changes count, which during migrations, cleanups and recoveries is the difference between weeks and months. Keep the pipeline healthy before you need it fast; the audit checklist is the maintenance schedule, and the inputs that make Google care remain the usual two — content worth indexing and links that vouch for it (we do the second).