WARC — Multi-Record SAMPLE
Multi-record WARC 1.0 SAMPLE (warcinfo + responses + resource) for web-archive tooling.
WARC/1.0
WARC-Type: warcinfo
WARC-Target-URI: urn:warc:novus-sample
WARC-Date: 2026-01-15T00:00:00Z
WARC-Record-ID: <urn:uuid:novus-wave-g-00254227>
Content-Length: 33
Content-Type: application/http; msgtype=response
software: Novus Examples Wave G
WARC/1.0
WARC-Type: response
WARC-Target-URI: https://sample.example/
WARC-Date: 2026-01-15T00:00:00Z
WARC-Record-ID: <urn:uuid:novus-wave-g-13701281>
Content-Length: 78
Content-Type: application/http; msgtype=response
HTTP/1.1 200 OK
Content-Type: text/html
<html><body>SAMPLE</body></html>
WARC/1.0
WARC-Type: response
WARC-Target-URI: https://sample.example/about
WARC-Date: 2026-01-15T00:00:00Z
WARC-Record-ID: <urn:uuid:novus-wave-g-09425874>
Content-Length: 84
Content-Type: application/http; msgtype=response
HTTP/1.1 200 OK
Content-Type: text/html
<html><body>About SAMPLE</body></html>
WARC/1.0
WARC-Type: resource
WARC-Target-URI: https://sample.example/style.css
WARC-Date: 2026-01-15T00:00:00Z
WARC-Record-ID: <urn:uuid:novus-wave-g-97379015>
Content-Length: 33
Content-Type: application/http; msgtype=response
body { font-family: system-ui; }
Specifications
- Records
- 4
- Version
- 1.0
- Wave
- G
What is a .warc file?
A WARC (Web ARChive) file stores one or more web resources — requests, responses, and metadata — concatenated as records with headers, in the standard format used by web crawlers and archives such as the Wayback Machine.
How to use this file
Use an example WARC to test WARC parsing, record extraction, and replay or conversion of archived web content.
How to use this file for testing
“WARC — Multi-Record SAMPLE” is a deterministic Novus Examples fixture for Conversion testing. The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.
Documented properties for this file: 4 records. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Archive fixtures are documented down to their member list (shown on this page) and are deliberately safe — no zip bombs, executables, or hidden payloads. Test extractors against nested, unicode, empty, and encrypted variants.
Related files
- 7z7z - LZMA2 ArchiveA 7-Zip archive using LZMA2 compression, containing a small documented file tree. For testing 7z extraction and conversion.

- 7z7z - Password ProtectedA password-protected 7z archive (AES-256, encrypted header). The password is “novus-example” — printed here on purpose so you can test encrypted-archive extraction. Contains only harmless sample text.

- bz2BZ2 - Bzip2 Single FileA single text file compressed with bzip2 — the container-free codec on its own, for testing decompression and codec detection.

- cpioCPIO Archive (newc/SVR4)A cpio archive in the newc (SVR4) format with four members — the Unix archive format used by initramfs and RPM. Hand-written to spec and documented down to its member list; for testing cpio extractors and converters.

- zipDataset Bundle (ZIP)A dataset bundle ZIP — a CSV with its JSON Schema, a README, and a LICENSE — the way datasets are commonly distributed. For testing unpack-and-validate pipelines and dataset importers.

- zipGDPR Export — JSON Only (ZIP, SAMPLE)SAMPLE subject-access ZIP with JSON profile/preferences only — fictional PII.

Generated by generation/breadth_wave_g.py. Free for any use, no attribution required — license.