EPUB 3 - UTF-16 Package Document
The package document is encoded as UTF-16 little-endian with a byte order mark while the content documents stay UTF-8 - a combination EPUB permits and almost nothing tests. For finding the parser that opens the OPF as UTF-8 and sees a file full of null bytes.
- mimetype
- container.xml
- content.opf
- nav.xhtml
- style.css
- chapter1.xhtml
Specifications
- Seed
- 20260807
- Epub Version
- 3.0
- Opf Encoding
- UTF-16LE with BOM
- Xml Declaration
- encoding="UTF-16"
- Content Document Encoding
- UTF-8
- Note
- valid XML; EPUB requires UTF-8 or UTF-16, and UTF-16 is the rarely-tested half
Testing contract
Expected to recover- Scenario
- Parse the package document and read the title and manifest.
- Expected result
- The XML declaration and BOM are honoured, the metadata and manifest parse correctly, and the mixed encodings across files in one container are handled per file rather than per book.
What is a .epub file?
EPUB is the open e-book standard, a ZIP archive containing XHTML content documents, CSS, images, and a package manifest describing reading order and metadata. It supports reflowable text that adapts to screen size and is supported by most e-readers except Kindle natively. It is the dominant format for distributable e-books.
How to use this file
Use an example EPUB to test e-book parsing, ZIP-package and manifest handling, reading-order and metadata extraction, and rendering in e-reader applications.
How to use this file for testing
“EPUB 3 - UTF-16 Package Document” is a deterministic Novus Examples fixture for Encoding detection, Metadata testing, Error handling. UTF-8, UTF-8-BOM, UTF-16, and Latin-1 files with documented encodings and line endings for testing charset detection.
Documented properties for this file: seed 20260807. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
E-book fixtures are structured archives documented down to their spine and manifest. Test readers and converters, and where a valid↔intentionally-invalid pair exists, assert graceful handling of the malformed container.
Code examples
ebook-convert utf16-package-document.epub out.pdf # Calibre
unzip -l utf16-package-document.epub # EPUB is a ZIPRelated files
- jsonCycloneDX SBOM With Unicode Metadata (Edge Case)A CycloneDX SBOM whose author names and properties contain accented Latin, CJK, right-to-left Hebrew, emoji and a deliberately long property value — for testing encoding handling and field-width assumptions. Every package, version, hash and licence is fictional — the tree describes nothing real.

- xmlJUnit XML — Escaped Characters in Test NamesTest names containing &, <, >, quotes and an apostrophe, escaped as XML requires. Round-tripping this report through a converter is the fastest way to find double-escaping bugs that turn & into &amp; one hop at a time.

- xmlJUnit XML — Locale-Formatted and Exponent DurationsDurations written four different ways in one file: scientific notation, a comma decimal separator from a German-locale JVM, a bare integer, and a comma that could be either a decimal point or a thousands separator. A parser using a locale-sensitive number reader gets a different answer depending on where it runs.

- nmeaNMEA 0183 — LF Line Endings Instead of CRLFThe same sentences terminated with a bare LF rather than the CR LF the standard mandates — what you get after a log passes through a text editor or a Unix pipeline. Parsers that strip only "\r\n" leave a stray character on every line and then fail the checksum.

- cpgShapefile — Code Page Sidecar (.cpg, UTF-8)A one-line sidecar naming the encoding of the .dbf. Without it a reader has to guess — usually cp1252 — and every non-ASCII attribute value comes back as mojibake, which is the single most common shapefile data-loss bug.

- cpgShapefile Code Page — ISO-8859-1 DeclarationThe sidecar that declares `attributes-cp1252.dbf` as ISO-8859-1. Note that the DBF also carries a language-driver byte in its header, and the two can disagree — which is exactly the ambiguity a reader has to resolve and document.

Generated by generation/ebooks_p7.py. Free for any use, no attribution required — license.