EML — Encoded-Word Longer Than the 75-Character Limit
One unbroken encoded-word of 148 characters, well past the 75-character ceiling RFC 2047 sets. Real senders emit these, and a decoder should still recover the whole word rather than truncating it at 75 or rejecting the header.
From: Maya Chen <maya@brightside.example>
To: Sam Rivera <sam@meridiansupply.example>
Subject: =?UTF-8?B?U2VociBsYW5nZSBCZXRyZWZmemVpbGUgbWl0IFVtbGF1dGVuIMOkw7bDvMOfIGZ1ZXIgZGVuIFNBTVBMRSBCZWxhc3R1bmdzdGVzdCBkZXIgS29wZnplaWxlbmRla29kaWVydW5n?=
Date: Mon, 09 Feb 2026 09:09:00 -0800
Message-ID: <p7-ew-overlong@brightside.example>
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
This message exists for its Subject header, not its body.
The decoded subject is recorded in the file's specs, so a header decoder
can be checked against a published expected string.
Specifications
- Wave
- p7
- Seed
- 20260807
- Encoded Word Chars
- 148
- Limit
- 75
- Charset
- UTF-8
- Encoding
- B (base64)
- Conformant
- false
- Decoded Subject
- Sehr lange Betreffzeile mit Umlauten äöüß fuer den SAMPLE Belastungstest der Kopfzeilendekodierung
- Line Endings
- CRLF (RFC 5322)
Testing contract
Expected to recover- Scenario
- Decode a single encoded-word that exceeds the RFC 2047 length limit instead of being split into several words.
- Expected result
- The full 98-character subject is recovered, with no truncation at the 75-character boundary.
What is a .eml file?
An EML file is a single email message stored in the RFC 822 / MIME format — plain-text headers (From, To, Subject, Date, Message-ID) followed by the body, which may be plain text, HTML, or a multipart structure with alternative bodies and file attachments encoded in base64.
How to use this file
Use an example EML to test email header parsing, MIME decoding, HTML-part handling, attachment extraction, and EML-to-other-format conversion.
How to use this file for testing
“EML — Encoded-Word Longer Than the 75-Character Limit” is a deterministic Novus Examples fixture for Email parsing, Encoding detection, Internationalization. Standards-compliant RFC 822 messages — plain, multipart text+HTML, and with an attachment — plus an MBOX mailbox, for testing header parsing, MIME decoding, attachment extraction, and mailbox splitting.
Documented properties for this file: seed 20260807 · B (base64) · CRLF (RFC 5322). Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Email fixtures use fixed dates, message IDs, and MIME boundaries so runs are reproducible, and every address is fictional. Test header parsing, MIME decoding, attachment extraction, and EML/MBOX conversion against the documented structure.
Code examples
from email import policy
from email.parser import BytesParser
msg = BytesParser(policy=policy.default).parse(open("ew-overlong-word.eml", "rb"))
print(msg["subject"], msg["from"])Related files
- emlEML — UTF-8 Encoded SubjectAn email whose Subject uses RFC 2047 UTF-8 encoded-words (café / crème) — for header decoding tests.

- emlEML — RFC 2231 Continuation Filename With size and creation-dateThe conformant answer to non-ASCII filenames: a filename split into three numbered RFC 2231 segments with a charset and a language tag, plus size and creation-date parameters. Parsers commonly handle filename*= but not the numbered continuation form.

- emlEML — UTF-8 Filename AttachmentAn attachment whose filename uses RFC 2231 UTF-8 encoding (résumé) — for MIME filename decoders.

- mboxMBOX — Four Charsets and Both Transfer EncodingsOne mailbox whose four messages each use a different charset and alternate between quoted-printable and base64, with RFC 2047 subjects to match. The file itself is LF-stored while the encoded payloads decode to CRLF text — the split every importer has to handle.

- emlEML — Base64: Accented UTF-8 BodyA base64 text/plain message in utf-8 carrying the identical body to the quoted-printable twin beside it. Base64 is opaque to whitespace and line-start characters, so it is the reference side of the pair when a QP decoder is suspect.

- emlEML — Base64: From-Line and Lone DotA base64 text/plain message in utf-8 carrying the identical body to the quoted-printable twin beside it. Base64 is opaque to whitespace and line-start characters, so it is the reference side of the pair when a QP decoder is suspect.

Generated by generation/email_p7.py. Free for any use, no attribution required — license.