Crawler File — Humans Extended
SAMPLE crawler/policy text file (humans-extended) for bot and ads.txt parsers.
/* TEAM */
Chef: Ada SAMPLE
Contact: ada@sample.example
/* SITE */
Standards: HTML5, CSS3
Specifications
- Wave
- I
- Role
- crawler-file
Testing contract
Expected to pass- Scenario
- Exercise Crawler File — Humans Extended in its crawlers workflow. SAMPLE crawler/policy text file (humans-extended) for bot and ads.txt parsers.
- Expected result
- 6 text lines, decoded as UTF-8; first nonempty line is '/* TEAM */'. Declared feature checks: role=crawler-file.
What is a .txt file?
TXT is a plain-text file containing unformatted character data with no styling or structure beyond line breaks. Its interpretation depends on character encoding, most commonly UTF-8, and on line-ending convention. It is the most universal and portable text container.
How to use this file
Use an example TXT to test encoding detection, line-ending (LF versus CRLF) handling, and any tool that reads or streams raw text input.
How to use this file for testing
“Crawler File — Humans Extended” is a deterministic Novus Examples fixture for Web assets. Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources, for testing web tooling, crawlers, PWA installers, and asset pipelines.
Documented properties for this file: crawler-file. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.
Web-platform fixtures are standards-compliant samples against fictional example.com data. Test crawlers, PWA installers and manifest validators, favicon/icon pipelines, service-worker registration, or .well-known parsers against the documented structure.
Related files
- txtads.txtAn IAB ads.txt listing authorised digital sellers with account IDs and relationships (sample data) — for testing ads.txt parsers and ad-fraud tooling.

- gzGzipped sitemap (500 URLs, .xml.gz)A 500-URL sitemap served the way large sites serve them - gzip-compressed as sitemap-large.xml.gz. The gzip header carries mtime 0 and no embedded filename, so the bytes are stable across regenerations. For testing that a crawler decompresses .xml.gz sitemaps before parsing.

- txthumans.txtA humans.txt crediting the people and stack behind a site, in the conventional TEAM/SITE block format — for testing plain-text metadata parsers.

- txtPlain-text sitemap (one URL per line)The plain-text sitemap format the sitemaps.org protocol also accepts: one absolute URL per line, no markup, UTF-8 encoded. For testing that a crawler supports the text form as well as XML.

- txtrobots.txtA robots.txt with wildcard and per-agent rules, a crawl-delay, and a sitemap reference — for testing robots parsers and crawler policy handling.

- txtrobots.txt - block everything except one pathA robots.txt that disallows the entire site for every crawler while allowing one media path for a single image agent. For testing full-block handling and the common misconception that a Disallow removes a URL from a search index.

Generated by generation/pad_wave_i.py. Free for any use, no attribution required, license.