Technical SEO
Crawl budget is the amount of attention Googlebot spends on your site — how many URLs it fetches, how often it returns. It's also the most misapplied concept in technical SEO: genuinely decisive for sites with hundreds of thousands of URLs, and almost entirely irrelevant below ten thousand, where Google's own guidance says not to worry. This closing guide of the technical series explains what crawl budget actually is, who needs to manage it, and the optimisation playbook for sites that do — plus the useful habit hiding inside the topic for everyone else.
What crawl budget is made of
Two components, per Google's own definition. Crawl capacity — how hard Googlebot is willing to hit your server without degrading it: healthy, fast servers earn higher parallel fetch rates; slow responses and 5xx errors throttle it down (another compounding return on speed work). Crawl demand — how much Google wants your URLs: popular, well-linked, frequently-updated pages get recrawled often; stale, orphaned, duplicate-looking URLs drift to monthly-or-never. Budget is the product of the two, set algorithmically, earned rather than requested — the same currency as recrawl speed throughout this series.
Do you actually have a crawl budget problem?
The qualifying test: are important pages going undiscovered or stale-in-index because Googlebot's attention runs out? Diagnose in Search Console's crawl stats report plus server logs. Symptoms that qualify: new products taking weeks to appear while the crawler churns through parameter permutations; log files showing 60% of Googlebot hits on faceted-filter URLs; "Discovered — currently not indexed" piling up on legitimate pages at large scale. If your site has 800 URLs and everything indexes within days — you don't have a crawl budget problem, and "optimising" it further is effort better spent literally anywhere else. What small sites should take from this topic: the waste patterns below are worth preventing on principle, because they're duplication and quality problems wearing a crawl costume.
Where budget actually goes to die
- Faceted navigation — the classic: filter/sort/paginate combinations generating millions of URL permutations of the same inventory. The fix stack: robots-block the infinite parameter spaces, canonicalise the crawlable variants, and internally link only the canonical forms.
- Redirect chains — every hop is a spent fetch, per the redirect hygiene rules; flatten them.
- Soft 404s and error pages — pages returning 200 with "no results found" get recrawled as if they were content; return real 404/410s so the crawler learns.
- Infinite spaces — calendar widgets with "next month" links to the year 3000, session-ID URLs, search-results pages: robots-block them all.
- Duplicate and thin sprawl — every near-identical variant fetched is a real page not fetched; consolidation is crawl optimisation.
- Slow responses and 5xx spikes — capacity throttling: the crawler backs off exactly when you deploy something broken. Server health is crawl strategy.
The optimisation playbook (large sites)
In order of leverage: fix server speed and error rates (raises capacity); block the infinite spaces in robots.txt (stops the bleeding — log analysis tells you which patterns first); flatten redirects and correct soft 404s; keep sitemaps clean and lastmod honest (steers demand toward what changed); strengthen internal linking to priority sections (demand follows links); and prune or noindex the sprawl that shouldn't compete for attention. Then re-read the logs a month later — crawl distribution shifting toward money pages is the success metric, ahead of any indexing count.
Frequently asked questions
Can I increase my crawl budget directly?
No dial exists. You raise capacity by being fast and stable, and demand by being linked, fresh and non-duplicative. Requesting individual crawls (URL Inspection) is a per-page nudge, not a budget change — the sitewide allocation is earned by the site's authority and health.
Does blocking pages in robots.txt free up budget?
Yes — blocked paths aren't fetched, and that attention redistributes. It's the right tool for infinite, valueless spaces; the wrong one for pages needing canonical or noindex treatment, which must be crawled to be processed. Match the tool to the intent, per the whole series.
My 500-page site — should I do any of this?
The prevention, yes (chains flattened, parameters canonicalised, sitemap honest — it's all just hygiene); the log-analysis campaign, no. At your scale the constraint on rankings was never crawl attention — it's content strength and the authority fee. That closes the technical series where every guide in it ends, deliberately: the plumbing earns you a fair hearing, and links win the argument (talk to us about those).