Gzipped sitemap (500 URLs, .xml.gz)
A 500-URL sitemap served the way large sites serve them - gzip-compressed as sitemap-large.xml.gz. The gzip header carries mtime 0 and no embedded filename, so the bytes are stable across regenerations. For testing that a crawler decompresses .xml.gz sitemaps before parsing.
gz
Single-file gzip stream
- Contains
- sitemap-large.xml
- Codec
- gzip
- Uncompressed
- 67 KB
- Compressed
- 1.6 KB
A container-free gzip stream — decompress it to recover sitemap-large.xml. The bytes are the compressed data, so there is no inline content preview.
Specifications
- Seed
- 70400
- Site
- example.com (fictional)
- Format
- urlset 0.9, gzip-wrapped
- Urls
- 500
- Codec
- gzip
- Original Name
- sitemap-large.xml
- Original Bytes
- 68610
- Gzip Mtime
- 0
- Line Endings
- LF
Testing contract
Expected to pass- Scenario
- Fetch a sitemap whose URL ends in .xml.gz and index its contents.
- Expected result
- The gzip stream decompresses to a well-formed urlset of exactly 500 <url> elements; a crawler that parses the compressed bytes directly finds zero URLs.
What is a .gz file?
GZ is a file compressed with gzip, using the DEFLATE algorithm within a simple single-stream container that stores one compressed file plus a small header and checksum. It compresses a single stream rather than bundling multiple files. It is ubiquitous for compressing logs, tarballs, and HTTP responses.
How to use this file
Use an example GZ to test gzip decompression, streaming inflate, checksum verification, and pipelines that transparently handle gzip-encoded content.
How to use this file for testing
“Gzipped sitemap (500 URLs, .xml.gz)” is a deterministic Novus Examples fixture for Web assets, Web scraping, Compression testing. Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources — for testing web tooling, crawlers, PWA installers, and asset pipelines.
Documented properties for this file: seed 70400 · LF · urlset 0.9, gzip-wrapped. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Web-platform fixtures are standards-compliant samples against fictional example.com data. Test crawlers, PWA installers and manifest validators, favicon/icon pipelines, service-worker registration, or .well-known parsers against the documented structure.
Code examples
gunzip -k sitemap-large.xml.gz # keep original
zcat sitemap-large.xml.gz | headRelated files
- txtrobots.txt - block everything except one pathA robots.txt that disallows the entire site for every crawler while allowing one media path for a single image agent. For testing full-block handling and the common misconception that a Disallow removes a URL from a search index.

- txtrobots.txt - Crawl-delay, Request-rate and Visit-timeA robots.txt carrying three different crawl-rate hints - an integer Crawl-delay, a fractional one alongside Request-rate and Visit-time, and a very large one - none of which are part of RFC 9309. For testing that a crawler reads or ignores rate hints without dropping the Disallow rules that share the group.

- txtrobots.txt - field case, tabs and indentationA robots.txt using upper-case, mixed-case and indented field names, a tab-indented rule, a value with no space after the colon, and both trailing and full-line comments. For testing that field names are treated case-insensitively while path values stay case-sensitive.

- txtrobots.txt - five Sitemap directivesA robots.txt declaring five sitemaps - before the first group, inside two different groups, in lower case, gzipped, and on another host. For testing that a discovery crawler collects Sitemap as a file-global field instead of scoping it to the group it sits in.

- txtrobots.txt - merged groups and Allow/Disallow precedenceA robots.txt where Googlebot is named by two separate groups and every Disallow has an equal-length Allow competing with it. For testing group merging, case-insensitive product tokens, and the rule that the most specific match wins with Allow breaking ties.

- txtrobots.txt - unknown and vendor-specific fieldsA robots.txt in which real Disallow rules are surrounded by Noindex, Host, Clean-param, Nofollow and two invented fields, several with inline comments. For testing that a parser skips fields it does not implement instead of aborting or mis-binding the rules that follow.

Generated by generation/web_p7.py. Free for any use, no attribution required — license.