Plain DOCX
A minimal Word document of plain paragraphs with no styling — the simplest valid DOCX for baseline parser testing.

Rendered preview of the docx file (35.9 KB). Download above for the original.
Specifications
- Headings
- 0
- Styling
- none
- Paragraphs
- 3
What is a .docx file?
DOCX is the default Microsoft Word format, an Office Open XML document packaged as a ZIP archive of XML parts, media, and relationships. It stores structured text, styles, tables, images, and metadata. It is the dominant format for editable word-processing documents.
How to use this file
Use an example DOCX to test Office Open XML parsing, text and style extraction, ZIP-package handling, and converters that render or transform Word documents.
How to use this file for testing
“Plain DOCX” is a deterministic Novus Examples fixture for Conversion testing. The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.
Documented properties for this file: DOCX · 36,730 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Document fixtures list their internal structure (pages, fields, tracked changes, embedded objects) in the spec table. Test extractors, converters, and OCR against that known structure, and compare searchable↔scanned or format-twin companions when present.
Code examples
from docx import Document # pip install python-docx
doc = Document("plain.docx")
for p in doc.paragraphs[:10]:
print(p.text)Related files
- docDOC — Legacy Word 97The formatted Word document saved as legacy binary .doc (Word 97-2003, OLE compound file) via LibreOffice. For testing legacy-Office parsers and DOC→DOCX conversion.

- docxDOCX with CommentsA Word document with two anchored reviewer comments — for testing comment extraction and whether converters preserve or drop review annotations.

- docxDOCX with Core and Custom PropertiesA valid DOCX package with fixed core properties and two custom properties surrounding an unchanged word/document.xml part.

- docxDOCX with Metadata StrippedThe same DOCX document part without core or custom property parts, suitable for checking OOXML metadata removal without content loss.

- docxDOCX with Tracked ChangesA Word document with real tracked changes — insertions and deletions attributed to two reviewers with timestamps — for testing how tools read, accept, reject, or preserve revisions.

- docxFormatted DOCXA Word document with styled headings, bold/italic runs, a table, and an embedded image — for testing DOCX parsers and converters against a rich document.

Generated by generation/docs_office.py. Free for any use, no attribution required — license.