Tagged PDF SAMPLE — Pdfua Note
SAMPLE PDF documenting tagged/accessibility intent (pdfua-note) in specs — bookmark outline present; not a full PDF/UA export.

Rendered preview of the pdf file (2 KB). Download above for the original.
Specifications
- Pages
- 1
- Tagged Note
- SAMPLE accessibility/structure intent documented in specs; not a certified tagged/PDF/UA file
- Structure Hint
- pdfua-note
- Lang
- und
- Suite
- wave-e
What is a .pdf file?
PDF (Portable Document Format) is a page-oriented document format that preserves fixed layout, fonts, vector and raster graphics, and text across platforms. It can also embed forms, annotations, attachments, and digital signatures. It is the de facto standard for finished, print-ready documents.
How to use this file
Use an example PDF to test text extraction, rendering, page-count and metadata parsing, form-field handling, and conversion or OCR pipelines.
How to use this file for testing
“Tagged PDF SAMPLE — Pdfua Note” is a deterministic Novus Examples fixture for PDF editor testing, Editor testing. Form PDFs, bookmarked documents, scanned pairs, and deliberately corrupt files for exercising PDF editors, parsers, and fillers.
Documented properties for this file: 1 pages. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Document fixtures list their internal structure (pages, fields, tracked changes, embedded objects) in the spec table. Test extractors, converters, and OCR against that known structure, and compare searchable↔scanned or format-twin companions when present.
Code examples
import pdfplumber # pip install pdfplumber
with pdfplumber.open("pdfua-note.pdf") as pdf:
print(len(pdf.pages), "pages")
print(pdf.pages[0].extract_text())Related files
- pdf10-Page PDF with Bookmarked TOCA ten-page PDF with a table of contents and a full bookmark outline (10 entries) — for testing PDF navigation, outline parsing, and page extraction.

- pdf50-Page PDF (text-light)A 50-page, text-light PDF — for testing page-count handling, pagination, and large-document navigation without a large file.

- pdfAirline Boarding Pass (PDF)An airline boarding pass laid out as a landscape PDF — passenger, flight, gate, seat, PNR, and a decorative barcode strip. A fixture for testing document-layout parsing and field extraction. Synthetic; the barcode is decorative.

- adocAsciiDoc Guide (ADOC source)An AsciiDoc getting-started guide with source blocks, an admonition, and a table — for testing AsciiDoc renderers (Asciidoctor), editors, and conversion to HTML/PDF.

- assASS Subtitles (Advanced SSA)An Advanced SubStation Alpha (.ass) subtitle with a styles section and three timed Dialogue events — the styled format libass renders — carrying the same captions as the SRT/VTT twins.

- pdfBank Statement (PDF)A one-page monthly bank statement — account summary and a dated transaction table with running balances. Synthetic data for testing PDF statement parsers, table extraction, and OCR.

Generated by generation/docs_wave_e.py. Free for any use, no attribution required — license.