Two URLs that return the same body are two pages to a machine that fetches once. A search engine has an index, a cluster of known duplicates and years of heuristics for choosing between them. An agent handed a URL has the URL, the response, and whatever the response says about itself. That last part is the job of rel=canonical, and it is a smaller job than most people assume. This post is about the duplicates that live inside one host: tracking parameters, trailing slashes, letter case, pagination and print views. The host-level duplicates, www against the apex and http against https, are covered in canonical hygiene, and this post assumes you have read it.
Where the copies come from
None of these is a mistake anyone made on purpose. Each is a URL your server answers with a 200 and the same content as some other URL.
| Source | Example | Why it exists |
|---|---|---|
| Tracking parameters | /pricing?utm_source=newsletter |
Every campaign link carries a different one |
| Trailing slash | /pricing and /pricing/ |
Both routes match and neither redirects |
| Letter case | /Pricing, /PRICING |
A case-insensitive filesystem or router |
| Index files | /docs/ and /docs/index.html |
Static hosting exposes both |
| Sort and filter parameters | /shoes?sort=price, /shoes?colour=black |
The same list in a different order |
| Session and cache-busting parameters | /about?v=3, /about?sid=… |
Appended by a framework or a tag |
| Print views | /guide/print, /guide?print=1 |
A stripped template around the same text |
| Pagination | /blog?page=2 |
Different content sharing a title and a template |
The last row is different in kind and gets its own section below; the rest are one page wearing different addresses, and a machine cannot know that unless the page tells it.
What a single fetch can and cannot know
An agent is handed https://example.com/pricing/?utm_source=newsletter, because that was the link in an email or a search result. It makes one request, receives a 200 and a body, and reads it. It has not fetched /pricing, does not know /pricing exists, and has no store of earlier fetches to compare against. If it later cites what it read, the obvious thing to cite is the URL it fetched, parameters and all. The two fetches covers what such a client does with the body; this is what it does with the address.
A search engine is in a different position. It has crawled both URLs, compared both bodies, noticed they match, and applied a set of signals to pick one, which is what Google's page on consolidating duplicate URLs describes. That machinery is a property of the index, not of the page. A client doing a plain fetch has none of it.
So the page has to carry its own identity. That is all rel=canonical is: a statement, inside the document, of which URL this content belongs to.
<link rel="canonical" href="https://example.com/pricing">
A consumer that reads it can record the canonical URL rather than the one it happened to fetch, treat a later fetch of another variant as the same document, and cite one address for one page. A consumer that ignores it sees nothing different. The element does not redirect, does not alter the response and does not stop the variant being fetched, and Google's documentation describes it as a strong signal it weighs alongside others rather than an instruction it must follow. It is an assertion, useful only when it is present and true.
How the scanner reads it
Check B7, canonical and duplication hygiene, is worth three points and rated Low effort. Two of those points depend on every crawled page carrying a rel=canonical whose path matches the path of the page it was fetched from. The comparison strips a trailing slash from both sides and ignores the query string, so a page fetched at /pricing/?utm_source=newsletter with a canonical of /pricing is self-consistent. It does not ignore case: a page fetched at /Pricing with a canonical of /pricing is a mismatch.
The crawl itself also has a view on duplicates. When the scanner collects links from the homepage to choose which five interior pages to fetch, it keys each candidate on its path with any trailing slash removed and keeps only the first URL for each key, shortest first. Links to /pricing and /pricing/ occupy one slot, and so do /pricing and /pricing?ref=nav. That is the scanner keeping its six-page budget on distinct pages, not a sign that the duplicates are harmless.
Parameters, slashes and case
For the rows in the table that are genuinely the same page, there are two correct answers, and you want both.
The first is the canonical. Emit it from the template, computed from the route rather than from the request. If the template builds the canonical from the request URL, a fetch of /pricing/?utm_source=x yields a canonical of /pricing/?utm_source=x, which passes the scanner's path comparison and is still wrong, because it declares the tracking variant to be its own page. Build it from the route's path and the site's canonical origin, and there is nothing to strip at request time.
The second is a redirect for the variants that can be redirected without losing anything. A trailing slash, a capitalised path and an index.html carry no information, so a permanent redirect to the canonical form costs nothing and removes the duplicate from existence.
GET /pricing/ -> 301 Location: /pricing
GET /Pricing -> 301 Location: /pricing
GET /docs/index.html -> 301 Location: /docs/
GET /pricing?utm_source=x -> 200 canonical: https://example.com/pricing
Tracking parameters are the exception, which is why the last line is a 200. Analytics needs the parameter to reach the page, so the response keeps the URL and the canonical does the work. Most frameworks expose a single trailing-slash setting that redirects the other form on every route; Next.js calls its one trailingSlash.
Sort and filter parameters
A list sorted by price is the same list. A list filtered to one colour is a subset, closer to a distinct page but rarely one you want treated as such. The usual approach is to canonicalise every sorted and filtered view to the unfiltered list, since that is the page you would want cited, and to make sure the unfiltered list is reachable at a clean URL with no parameter at all. The scanner's path comparison ignores the parameters either way; the reason to get this right is the consumer that arrives from elsewhere. A shopping agent handed /shoes?colour=black&sort=price learns from the canonical, and from nothing else, that the page to remember is /shoes.
Pagination is not duplication
Page two of a listing shares a title, a description and a template with page one, and has different content. Pointing its canonical at page one is a common mistake and a false statement: it tells a consumer that the second page's items are to be found on the first, and a consumer that believes it will never read them. The scanner fails it too, because /blog?page=2 with a canonical of /blog is a path mismatch under B7 however the path is written.
Each page in a series should carry a canonical naming itself, page parameter included, and a title that differs from the others so that check B4's uniqueness test does not see six copies of the same one. If the series exists only to be walked, listing every item in the sitemap is the better map for a machine.
Print views
A print view serves the same text with the chrome stripped, and it is the one duplicate on the list that is arguably the better page for a machine: less navigation, less script, a higher clean text ratio. It still should not be its own page. Give it a canonical naming the full version, keep it out of the sitemap, and let the full version earn the citation. If the print view is the only version that renders the whole article without JavaScript, that is a finding about the full version.
Check your own site
Run the free scan and open B7 on the report; the evidence records selfConsistentCanonicals against pagesCrawled, alongside the host probes. The check is defined on the methodology page under understanding, and every page on this site emits a canonical computed from its route, which you can see on our own report.