WordPress SEO
WordPress manufactures duplicate content by design: the same post reachable through its permalink, category archives, tag archives, date archives, author archives, paginated feeds and — until you stop it — its own attachment pages. None of it is penalty material (the standing correction), all of it dilutes: crawl attention spread across near-copies, archive pages competing weakly with the posts they excerpt, and the site's quality aggregate carrying a long tail of thin URLs. Here's WordPress's native duplication catalogued, and the fix per mechanism — most of them toggles you've met across this silo, assembled here as one sweep.
The mechanisms and their fixes
- Archive proliferation (the big one): category, tag, date, author and format archives all excerpt the same posts. The fix is the indexation map: categories curated and indexed as real pages (per the taxonomy doctrine), tags noindexed by default, date/format archives noindexed or disabled, author archives disabled on single-author sites — one deliberate pass in the SEO plugin.
- Attachment pages: every upload minting a page containing one image — the thin-page flood, ended by the redirect-to-file toggle, per the standing rule.
- The URL-variant layer: http/https, www/bare, trailing-slash, ?replytocom= comment parameters, ?print= and friends — handled by the canonical stack: server-level 301s to the one form (the address consistency check), the SEO plugin's automatic self-canonicals absorbing the parameter noise (verify they render — one canonical per page, pointing at the clean URL), and replytocom specifically defused by the plugins' defaults.
- Excerpt-vs-full-content archives: archives showing full post content are literal duplication — theme setting to excerpts, which also improves the archives as pages.
- Pagination of comments and archives: comment pagination off (the discussion settings), archive pagination left as ordinary self-canonical pages per the pagination rules.
- Plugin-manufactured duplicates: print pages, AMP variants (the AMP decision's leftovers), builder revisions exposed at URLs, staging copies indexed (the lock) — the quarterly crawl finds them; the fix is per-plugin (canonical, noindex, or deletion of the feature).
- Cross-site duplication: your RSS-scraped copies (originals nearly always win — no action per the scraper rules) and your own syndication (canonical or delay, per the same guide).
The verification sweep
After the toggles: a crawl filtered for the patterns above (parameter URLs, attachment pages, archive depth — each should now be redirecting, canonicalised or noindexed as designed), the site:-search spot check for a distinctive sentence (multiple results = a mechanism missed), and Search Console's duplicate-cluster reports as the ongoing monitor — "duplicate without user-selected canonical" trending down is the sweep working, per the audit's reading.
Frequently asked questions
Which of these actually costs rankings if left alone?
Ranked by observed damage: indexed staging/AMP/print copies (real competition with yourself), attachment-page floods at scale (quality aggregate), full-content archives (occasional archive-outranks-post weirdness), then the parameter noise (mostly absorbed by canonicals anyway). The sweep's hour covers all tiers; the first two justify it alone.
Should I use canonicals or noindex for the archives?
The decision grid: archives aren't duplicates of any single post (they're collections) — so it's index-or-noindex by archive value (the taxonomy doctrine), not canonicalise-to-somewhere. Canonicals handle the true-variant layer (parameters, print views); noindex handles the thin-collection layer.
My scraped content outranks me occasionally. Escalate?
Check the original-wins conditions first — it usually self-corrects on recrawl; persistent losses signal the site's authority deficit more than a duplication failure, and the fix is the general one: earned links making your domain the obvious original (our fix), with DMCA reserved for the rare commercial-scale theft.