robots.txt that is actually an HTML 404 page
The single most common broken robots.txt in the wild: a server that answers /robots.txt with its HTML 404 page instead of a 404 status. For testing that a crawler treats unparseable markup as 'no robots.txt' and allows the site rather than inventing rules from tag names.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>404 - Not Found</title>
</head>
<body>
<h1>404 - Not Found</h1>
<p>The requested URL /robots.txt was not found on this server.</p>
<hr>
<address>example.com sample fixture</address>
</body>
</html>
Specifications
- Seed
- 70400
- Site
- example.com (fictional)
- Format
- robots.txt (RFC 9309)
- Actual Content
- text/html error page
- Directives
- 0
- Deliberately Invalid
- true
Testing contract
Expected to fail- Scenario
- Feed an HTML error page to a robots.txt parser.
- Expected result
- Zero groups and zero rules are extracted, `<html>` and `<h1>` are never mistaken for fields, and the crawler falls back to allow-all as it would for a missing file.
What is a .txt file?
TXT is a plain-text file containing unformatted character data with no styling or structure beyond line breaks. Its interpretation depends on character encoding, most commonly UTF-8, and on line-ending convention. It is the most universal and portable text container.
How to use this file
Use an example TXT to test encoding detection, line-ending (LF versus CRLF) handling, and any tool that reads or streams raw text input.
How to use this file for testing
“robots.txt that is actually an HTML 404 page” is a deterministic Novus Examples fixture for Web assets, Web scraping, Editor testing. Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources — for testing web tooling, crawlers, PWA installers, and asset pipelines.
Documented properties for this file: seed 70400 · robots.txt (RFC 9309). Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Web-platform fixtures are standards-compliant samples against fictional example.com data. Test crawlers, PWA installers and manifest validators, favicon/icon pipelines, service-worker registration, or .well-known parsers against the documented structure.
Related files
- txtPlain-text sitemap (one URL per line)The plain-text sitemap format the sitemaps.org protocol also accepts: one absolute URL per line, no markup, UTF-8 encoded. For testing that a crawler supports the text form as well as XML.

- xmlSitemap index over three sitemapsA sitemap index listing three child sitemaps, one dated with a plain date, one with a full W3C datetime, and one with no lastmod at all. Two of the children ship alongside it, so a crawler can be walked from index to URL.

- xmlSitemap shard 1 (core pages)The first child of the sitemap index: six core URLs with a deliberately uneven mix of lastmod, changefreq and priority, including one entry that carries nothing but a loc. For testing that optional sitemap fields really are optional.

- xmlSitemap shard 2 (blog and legal)The second child of the sitemap index: six URLs dated with W3C datetimes in two different timezone offsets, including a changefreq of never and a priority of 0.1. For testing datetime normalisation and priority ordering.

- xmlSitemap with a deliberately invalid hreflang clusterA deliberately invalid hreflang cluster: an `en-UK` region that does not exist, an underscore locale, a German page that never links back, a Spanish self-reference dropped to http, and no x-default anywhere. For testing that an international audit reports each defect rather than the first one.

- xmlSitemap with a reciprocal hreflang clusterA sitemap whose three URLs form a complete hreflang cluster: every page lists every locale including itself and a shared x-default. This is the shape an international audit should pass, and the reference twin of the broken cluster.

Generated by generation/web_p7.py. Free for any use, no attribution required — license.