Technical SEO

XML Sitemaps: How to Create, Submit, and Maintain Them

Your declared URL inventory: the format, the golden rule of only canonical live URLs, submission, and the diagnostics that make sitemaps an instrument panel.

XML Sitemaps: How to Create, Submit, and Maintain Them

An XML sitemap is your site's declared inventory: a machine-readable list of the URLs you want search engines to know about, with optional freshness hints. It doesn't make anything rank — it makes things findable, faster and more completely than link-following alone, which matters most exactly when sites are new, large, or freshly restructured. Here's how to create one properly, submit it, and — the part everyone skips — keep it telling the truth.

What a sitemap does (and doesn't)

It feeds the discovery stage: URLs in the sitemap get found without waiting for link-crawling, and lastmod hints help recrawl prioritisation. It does not guarantee indexing (quality selection still applies), doesn't override robots or noindex, and passes no equity — inclusion is information, not endorsement. Small, well-linked sites barely need one; new, large, or deep sites lean on it heavily. Everyone should have one anyway, because it costs nothing and its diagnostics (below) are valuable on every site.

The format, in one example

<?xml version="1.0" encoding="UTF-8"?> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"> <url> <loc>https://example.com/blogs/link-building</loc> <lastmod>2020-03-11</lastmod> </url> </urlset>

Notes that matter: loc must be the absolute, canonical URL; lastmod should reflect real content changes (Google has said it uses lastmod when it's trustworthy — auto-stamping every URL daily teaches it to ignore yours); priority and changefreq are largely ignored — skip them. Limits: 50,000 URLs / 50MB per file; beyond that, split into multiple sitemaps under a sitemap index file — and splitting by section (posts, products, categories) is worth doing far below the limit, because per-sitemap indexing stats are the diagnostic payoff.

The golden rule: only canonical, indexable, live URLs

A sitemap should contain exactly the URLs you want indexed and nothing else: 200-status, canonical versions, no redirects, no 404s, no noindexed pages, no canonicalised-away variants. Every junk entry wastes crawl attention and — worse — muddies the diagnostics: "why is Google indexing 400 of my 500 sitemap URLs?" is a useful question only when the 500 are all genuine. Sitemap hygiene is index hygiene made visible.

Creation and submission

  1. Generate dynamically — every serious CMS and framework builds sitemaps from the live database (plugins for WordPress-class platforms; a route in custom apps). Static hand-built sitemaps rot within a week; if yours has a "generated on" date months back, it's fiction.
  2. Reference it in robots.txt (Sitemap: line) — the passive declaration every crawler reads.
  3. Submit in Search Console (Sitemaps report) — the active handshake, and the unlock for the reporting: per-sitemap discovered/indexed counts, the fastest visibility into systematic indexing problems that exists.

Maintenance: the sitemap as instrument panel

The ongoing value is diagnostic. In Search Console, the gap between "URLs in sitemap" and "indexed" is your quality/duplication backlog, itemised by reason; a section sitemap whose indexing rate suddenly drops is an early alarm (template problem, quality issue, accidental noindex) that beats waiting for traffic to fall. Fold the checks into the quarterly audit: sitemap fetches cleanly, counts match the CMS's reality, no redirects/404s inside (a crawler run over the sitemap URLs verifies in minutes), and lastmod moving only when content does — the same honesty this site's own refresh discipline enforces editorially.

Frequently asked questions

Do images and videos need their own sitemaps?

Large visual catalogues benefit from image/video sitemap extensions (or separate sitemaps) for media discovery, feeding the image search channel. Ordinary blogs don't need them — in-page markup suffices.

Why does Google index pages that aren't in my sitemap?

Sitemaps are additive, not exclusive — link discovery still finds everything reachable. Unwanted indexed URLs are a noindex/canonical job; the sitemap can't unlist what links reveal.

My sitemap says 2,000 URLs; Google indexed 900. Emergency?

A prompt, not an alarm: read the excluded reasons. Duplicates-with-canonicals and intentional noindex are fine; "crawled — not indexed" at scale means a quality selection conversation, and discovery-only exclusions on a new site mean patience plus internal links. The sitemap didn't cause any of it — it's the instrument that let you see it. Ranking what's indexed remains the usual business: on-page, coverage, and authority (yes, us).

Put this into practice

Every site on BacklinksMedia is verified, priced upfront and ready to order.

Explore marketplace
Technical SEO xml sitemap sitemap.xml create sitemap submit sitemap google
RG
Rajiv Gupta

Growth engineer at BacklinksMedia, working on outreach analytics and the verified link marketplace.