robots.txt - Crawl-delay, Request-rate and Visit-time
A robots.txt carrying three different crawl-rate hints - an integer Crawl-delay, a fractional one alongside Request-rate and Visit-time, and a very large one - none of which are part of RFC 9309. For testing that a crawler reads or ignores rate hints without dropping the Disallow rules that share the group.
# robots.txt - crawl-rate directives. None of these are in RFC 9309: Crawl-delay
# is a de-facto extension, Request-rate and Visit-time are older conventions.
User-agent: *
Crawl-delay: 10
Disallow: /cart/
Disallow: /checkout/
User-agent: SlowBot
Crawl-delay: 0.5
Request-rate: 1/10s
Visit-time: 0200-0600
Disallow:
User-agent: BurstBot
Crawl-delay: 120
Disallow: /api/
Specifications
- Seed
- 70400
- Site
- example.com (fictional)
- Format
- robots.txt (RFC 9309)
- Groups
- 3
- Crawl Delays
- 10, 0.5, 120 seconds
- Standardised
- false
Testing contract
Reference control- Scenario
- Parse a file whose groups mix standard Disallow rules with non-standard rate hints.
- Expected result
- Crawl-delay parses as 10, 0.5 and 120 for the three groups (or is ignored as unsupported), and either way /cart/ stays disallowed for the default agent and the empty `Disallow:` in the SlowBot group allows everything.
What is a .txt file?
TXT is a plain-text file containing unformatted character data with no styling or structure beyond line breaks. Its interpretation depends on character encoding, most commonly UTF-8, and on line-ending convention. It is the most universal and portable text container.
How to use this file
Use an example TXT to test encoding detection, line-ending (LF versus CRLF) handling, and any tool that reads or streams raw text input.
How to use this file for testing
“robots.txt - Crawl-delay, Request-rate and Visit-time” is a deterministic Novus Examples fixture for Web assets, Web scraping, Editor testing. Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources — for testing web tooling, crawlers, PWA installers, and asset pipelines.
Documented properties for this file: seed 70400 · robots.txt (RFC 9309). Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Web-platform fixtures are standards-compliant samples against fictional example.com data. Test crawlers, PWA installers and manifest validators, favicon/icon pipelines, service-worker registration, or .well-known parsers against the documented structure.
Related files
- txtPlain-text sitemap (one URL per line)The plain-text sitemap format the sitemaps.org protocol also accepts: one absolute URL per line, no markup, UTF-8 encoded. For testing that a crawler supports the text form as well as XML.

- xmlSitemap index over three sitemapsA sitemap index listing three child sitemaps, one dated with a plain date, one with a full W3C datetime, and one with no lastmod at all. Two of the children ship alongside it, so a crawler can be walked from index to URL.

- xmlSitemap shard 1 (core pages)The first child of the sitemap index: six core URLs with a deliberately uneven mix of lastmod, changefreq and priority, including one entry that carries nothing but a loc. For testing that optional sitemap fields really are optional.

- xmlSitemap shard 2 (blog and legal)The second child of the sitemap index: six URLs dated with W3C datetimes in two different timezone offsets, including a changefreq of never and a priority of 0.1. For testing datetime normalisation and priority ordering.

- xmlSitemap with a deliberately invalid hreflang clusterA deliberately invalid hreflang cluster: an `en-UK` region that does not exist, an underscore locale, a German page that never links back, a Spanish self-reference dropped to http, and no x-default anywhere. For testing that an international audit reports each defect rather than the first one.

- xmlSitemap with a reciprocal hreflang clusterA sitemap whose three URLs form a complete hreflang cluster: every page lists every locale including itself and a shared x-default. This is the shape an international audit should pass, and the reference twin of the broken cluster.

Generated by generation/web_p7.py. Free for any use, no attribution required — license.