Technical SEO
Duplicate content
The same or substantially similar content appearing at more than one URL, which splits ranking signals and wastes crawl budget.
By Shimon Carroll, Founder, SEO for AI Agents · Last updated
Duplicate content is content that is identical or near-identical across multiple URLs. It is rarely a penalty in the punitive sense; the real cost is dilution. When the same content lives at several URLs, ranking signals (links, engagement) are split across them, no single version is as strong as it could be, and the engine has to spend crawl budget deciding which one to show.
Most duplication is accidental and technical: http and https versions, www and non-www, trailing-slash variants, URL parameters for tracking or sorting, printer-friendly pages, and pagination. Some is structural, like manufacturer product descriptions copied across thousands of retailer pages, or boilerplate city pages that differ only by a place name. The latter is also a quality problem, not just a technical one.
The fixes are canonical tags to nominate the preferred URL, consistent internal linking and redirects so all signals agree, and parameter handling so tracking variations do not generate index entries. For structurally thin duplicates like boilerplate location pages, the fix is editorial: make each page genuinely unique, or consolidate. We flag near-duplicate clusters and check that canonicals resolve them rather than leaving the engine to guess.
Related terms
Primary sources
- Google Search Central, consolidate duplicate URLs
Google Search Central