EPUB 3 - UTF-8 BOM on Every XML Part
The package document, the navigation document and the chapter each begin with a UTF-8 byte order mark before the XML declaration. Legal, and the classic reason a strict parser reports content before the prolog on a file that looks perfectly fine in an editor.
- mimetype
- container.xml
- content.opf
- nav.xhtml
- style.css
- chapter1.xhtml
Specifications
- Seed
- 20260807
- Epub Version
- 3.0
- Bom Bytes
- EF BB BF
- Parts With Bom
- content.opf, nav.xhtml, chapter1.xhtml
- Declared Encoding
- UTF-8
- Container Xml Has Bom
- false
Testing contract
Expected to recover- Scenario
- Parse each XML part in the container with a strict XML parser.
- Expected result
- The BOM is consumed as an encoding signature rather than as document content; no part reports a stray character before the XML declaration, and the title reads correctly.
What is a .epub file?
EPUB is the open e-book standard, a ZIP archive containing XHTML content documents, CSS, images, and a package manifest describing reading order and metadata. It supports reflowable text that adapts to screen size and is supported by most e-readers except Kindle natively. It is the dominant format for distributable e-books.
How to use this file
Use an example EPUB to test e-book parsing, ZIP-package and manifest handling, reading-order and metadata extraction, and rendering in e-reader applications.
How to use this file for testing
“EPUB 3 - UTF-8 BOM on Every XML Part” is a deterministic Novus Examples fixture for Encoding detection, Metadata testing, Error handling. UTF-8, UTF-8-BOM, UTF-16, and Latin-1 files with documented encodings and line endings for testing charset detection.
Documented properties for this file: seed 20260807. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
E-book fixtures are structured archives documented down to their spine and manifest. Test readers and converters, and where a valid↔intentionally-invalid pair exists, assert graceful handling of the malformed container.
Code examples
ebook-convert bom-on-every-xml-part.epub out.pdf # Calibre
unzip -l bom-on-every-xml-part.epub # EPUB is a ZIPRelated files
- jsonCycloneDX SBOM With Unicode Metadata (Edge Case)A CycloneDX SBOM whose author names and properties contain accented Latin, CJK, right-to-left Hebrew, emoji and a deliberately long property value — for testing encoding handling and field-width assumptions. Every package, version, hash and licence is fictional — the tree describes nothing real.

- xmlJUnit XML — Escaped Characters in Test NamesTest names containing &, <, >, quotes and an apostrophe, escaped as XML requires. Round-tripping this report through a converter is the fastest way to find double-escaping bugs that turn & into &amp; one hop at a time.

- xmlJUnit XML — Locale-Formatted and Exponent DurationsDurations written four different ways in one file: scientific notation, a comma decimal separator from a German-locale JVM, a bare integer, and a comma that could be either a decimal point or a thousands separator. A parser using a locale-sensitive number reader gets a different answer depending on where it runs.

- nmeaNMEA 0183 — LF Line Endings Instead of CRLFThe same sentences terminated with a bare LF rather than the CR LF the standard mandates — what you get after a log passes through a text editor or a Unix pipeline. Parsers that strip only "\r\n" leave a stray character on every line and then fail the checksum.

- cpgShapefile — Code Page Sidecar (.cpg, UTF-8)A one-line sidecar naming the encoding of the .dbf. Without it a reader has to guess — usually cp1252 — and every non-ASCII attribute value comes back as mojibake, which is the single most common shapefile data-loss bug.

- cpgShapefile Code Page — ISO-8859-1 DeclarationThe sidecar that declares `attributes-cp1252.dbf` as ISO-8859-1. Note that the DBF also carries a language-driver byte in its header, and the two can disagree — which is exactly the ambiguity a reader has to resolve and document.

Generated by generation/ebooks_p7.py. Free for any use, no attribution required — license.