Skip to content
Novus Examples
txt506 B

robots.txt - merged groups and Allow/Disallow precedence

A robots.txt where Googlebot is named by two separate groups and every Disallow has an equal-length Allow competing with it. For testing group merging, case-insensitive product tokens, and the rule that the most specific match wins with Allow breaking ties.

Preview — first 19 linestxt
# robots.txt - group merging and Allow/Disallow precedence (RFC 9309)

User-agent: Googlebot
User-agent: Bingbot
Disallow: /catalog/
Allow: /catalog/

User-agent: googlebot
Disallow: /beta/
Allow: /beta/preview/

User-agent: *
Disallow: /internal/
Disallow: /catalog/

# Two groups name Googlebot (product tokens are case-insensitive). A conforming
# parser merges them into ONE group of four rules rather than honouring only the
# first, and never falls back to the * group for an agent that has its own.

Specifications

Seed
70400
Site
example.com (fictional)
Format
robots.txt (RFC 9309)
Groups
3
Agents Named
3
Conflicting Rules
2

Testing contract

Expected to pass
Scenario
Resolve the effective rule set for Googlebot, Bingbot and an unnamed crawler.
Expected result
Googlebot's group merges to four rules and never inherits the * group; /catalog/shoes is ALLOWED for Googlebot (equal-length Allow wins) but disallowed for an unnamed crawler; /beta/preview/x is allowed while /beta/x is not.

What is a .txt file?

TXT is a plain-text file containing unformatted character data with no styling or structure beyond line breaks. Its interpretation depends on character encoding, most commonly UTF-8, and on line-ending convention. It is the most universal and portable text container.

How to use this file

Use an example TXT to test encoding detection, line-ending (LF versus CRLF) handling, and any tool that reads or streams raw text input.

How to use this file for testing

“robots.txt - merged groups and Allow/Disallow precedence” is a deterministic Novus Examples fixture for Web assets, Web scraping, Editor testing. Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources — for testing web tooling, crawlers, PWA installers, and asset pipelines.

Documented properties for this file: seed 70400 · robots.txt (RFC 9309). Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Web-platform fixtures are standards-compliant samples against fictional example.com data. Test crawlers, PWA installers and manifest validators, favicon/icon pipelines, service-worker registration, or .well-known parsers against the documented structure.

Generated by generation/web_p7.py. Free for any use, no attribution required — license.