HTML with three conflicting canonical tags (deliberately invalid)
A deliberately invalid page carrying two canonical tags in the head and a third in the body. Search engines ignore body-level canonicals entirely and treat multiple head-level ones as a conflict. For testing that an extractor reports position as well as value.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Gloves | Example</title>
<link rel="canonical" href="https://www.example.com/catalog/gloves">
<link rel="canonical" href="https://www.example.com/catalog/gloves?colour=black">
<meta name="robots" content="index,follow">
</head>
<body>
<h1>Gloves</h1>
<link rel="canonical" href="https://www.example.com/catalog/">
<p>Sample page with a third canonical illegally placed in the body.</p>
</body>
</html>
Specifications
- Seed
- 70400
- Site
- example.com (fictional)
- Format
- HTML5 head fixture
- Encoding
- UTF-8
- Canonical Tags
- 3
- In Head
- 2
- In Body
- 1
- Deliberately Invalid
- true
Testing contract
Expected to fail- Scenario
- Extract canonical tags and decide which one, if any, applies.
- Expected result
- Three link elements are found but only the two in <head> are eligible; the extractor reports a conflict rather than silently taking the first, and the body-level tag is explicitly discarded.
What is a .html file?
HTML (HyperText Markup Language) is the structural markup language of the web, using nested tags to define document content, semantics, and links. It is typically paired with CSS for presentation and JavaScript for behavior. It is the foundational format rendered by browsers.
How to use this file
Use an example HTML file to test parsers, DOM construction, sanitization of untrusted markup, and rendering or scraping pipelines.
How to use this file for testing
“HTML with three conflicting canonical tags (deliberately invalid)” is a deterministic Novus Examples fixture for Web assets, HTML parsing, Metadata testing. Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources — for testing web tooling, crawlers, PWA installers, and asset pipelines.
Documented properties for this file: seed 70400 · UTF-8 · HTML5 head fixture. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Web-platform fixtures are standards-compliant samples against fictional example.com data. Test crawlers, PWA installers and manifest validators, favicon/icon pipelines, service-worker registration, or .well-known parsers against the documented structure.
Related files
- htmlHTML hreflang cluster, English pageThe English member of a correct three-locale hreflang cluster: it lists itself, both translations and the x-default, and its canonical agrees. Ships with the German page so reciprocity can be verified across real files.

- htmlHTML hreflang cluster, German pageThe German member of the same cluster, with lang="de" on the html element and the identical alternate block. For testing that an audit compares the annotation set rather than the document language, and that the two agree.

- htmlHTML with deliberately invalid hreflang annotationsA deliberately invalid annotation block: `en-UK` is not a region code, `de_DE` uses an underscore, `zz` is not a language, one alternate has no hreflang attribute at all, one drops to http, there is no x-default, and the canonical points outside the cluster. For testing that every defect is reported, not just the first.

- txtdnt-policy.txt (Do Not Track compliance statement)A sample machine-discoverable Do Not Track policy of the kind served at /.well-known/dnt-policy.txt, stating retention windows, exceptions and a contact for a fictional site. For testing crawlers and privacy scanners that look for the document and read its version header.

- xmlhost-meta (XRD, XML form)Host metadata in its original XRD form: a subject, an alias, an expiry, one property, and three Link elements including an lrdd template with a {uri} placeholder. Paired with the JRD twin that carries exactly the same data.

- jsonhost-meta.json (JRD, JSON form)The JSON Resource Descriptor twin of the XRD host-meta: identical subject, alias, expiry, property and three links, expressed with lower-case JSON member names. For testing XRD-to-JRD conversion against a known answer.

Generated by generation/web_p7.py. Free for any use, no attribution required — license.