Synthetic Industry

Troubleshooting guide · updated 2026-10-11

Search reports show duplicates, blocked or redirected pages? Make canonical tags, sitemap and robots rules agree

Canonical tags, the sitemap and robots.txt each tell search engines something different. When they contradict each other, pages are skipped or grouped wrongly.

Three systems, three different messages

A site tells search engines which addresses matter in three separate places. Canonical tags say which address is the preferred version of a page. The sitemap lists addresses you would like to appear in results. The robots.txt file says which paths crawlers may fetch. They are maintained by different people at different times, so they drift apart. The consequences are specific: a page may be reported as a duplicate of another, a sitemap may list addresses that redirect or do not exist, and a page you wanted removed may remain visible.

  • Make a list of what each system says about ten representative pages.
  • Look for any page where two of them disagree.

Canonical tags: one clear signal per page

Google describes redirects and rel="canonical" as strong signals and sitemap inclusion as a weak one, and calls all of them signals, not commands. It advises absolute addresses, a self-referential canonical on the preferred page itself, and one method per page, because telling Google different things by different techniques, for example a sitemap address that differs from the page's canonical tag, creates conflicting signals. Linking internally to the canonical address and avoiding HTTP versions when HTTPS is intended also helps. It specifically warns against using robots.txt for canonicalisation: blocked addresses can still be indexed without their content.

  • A canonical that points to a redirecting or non-existent address is a fault to fix.
  • JavaScript that changes a canonical set in the HTML source can cause disagreement.
  • Pick the HTML tag or the HTTP header for a page, not both.

The sitemap: only addresses you want found

Google says a sitemap should use fully qualified absolute addresses and include the canonical addresses you want in results, listing only the preferred one where the same content lives at several addresses. A single file is limited to 50,000 addresses or 50MB uncompressed. Google uses lastmod only when it is consistently and verifiably accurate, and it ignores priority and changefreq. Submission is described as a hint, not a guarantee. A sitemap padded with redirecting, blocked, error or noindex addresses wastes that hint and generates warnings.

  • Remove redirecting, error and noindex addresses from the sitemap.
  • Do not set lastmod to the build time of every page.
  • Split large sets into several files with an index.

The robots and noindex trap

robots.txt controls crawling, not indexing. Google states it is not a mechanism for keeping a page out of Google, and that a blocked address may still be indexed, without a description, if other sites link to it. To keep a page out, use noindex or a password. But noindex has its own trap: Google can only act on it if the crawler can fetch the page, so a page that is blocked in robots.txt and marked noindex keeps the block and never shows the rule. Teams add both "for safety" and get the opposite of what they wanted.

  • If a page must stay out of results, allow crawling and add noindex, then wait for a recrawl.
  • If a page should be found, make sure no rule blocks it and no noindex remains from a staging copy.
  • After a relaunch, check that a staging noindex or blocking rule was not copied to production.

What the paid job covers and how it is accepted

These checks are part of our address job (posted test price from £395, untested, quoted after we see your list or a crawl; up to 300 agreed addresses on one site). It is accepted when each indexable page on the list carries one absolute self-referencing canonical tag (or a deliberate exception), the sitemap lists every indexable successful canonical address on the list and none that redirect, error, carry noindex or are blocked, and a sample of addresses outside the list is unchanged. It makes no promise about indexing or ranking. Send the site and the exact search-report messages as text in your first enquiry, not logins.

Sources and limits

  • Google Search Central: Consolidate duplicate URLs Checked 2026-10-11.
    • Redirects and rel=canonical are strong canonical signals and sitemap inclusion is a weak one; use absolute URLs; add a self-referential canonical on the canonical page; avoid giving a page different canonical URLs through different techniques; do not use robots.txt for canonicalisation; link internally to the canonical URL consistently.
  • Google Search Central: Build and submit a sitemap Checked 2026-10-11.
    • A sitemap file is limited to 50MB uncompressed or 50,000 URLs; use fully qualified absolute URLs; include canonical URLs you want in results; lastmod is used only if consistently and verifiably accurate; Google ignores priority and changefreq; submission is a hint.
  • Google Search Central: Introduction to robots.txt Checked 2026-10-11.
    • robots.txt is for crawl management and is not a mechanism for keeping a page out of Google; a blocked URL can still be indexed without a description if other sites link to it.
  • Google Search Central: Block search indexing with noindex Checked 2026-10-11.
    • noindex only works if the page is not blocked by robots.txt, because the crawler must fetch the page to see the rule.