OCR Domain — Handwritten-print Form (Searchable PDF)
A searchable PDF handwritten-print form with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

Rendered preview of the pdf file (1.8 KB). Download above for the original.
Specifications
- Role
- searchable text PDF
- Domain
- handwritten
- Has Text
- true
- Page Size
- US Letter
What is a .pdf file?
PDF (Portable Document Format) is a page-oriented document format that preserves fixed layout, fonts, vector and raster graphics, and text across platforms. It can also embed forms, annotations, attachments, and digital signatures. It is the de facto standard for finished, print-ready documents.
How to use this file
Use an example PDF to test text extraction, rendering, page-count and metadata parsing, form-field handling, and conversion or OCR pipelines.
How to use this file for testing
“OCR Domain — Handwritten-print Form (Searchable PDF)” is a deterministic Novus Examples fixture for OCR testing, PDF editor testing. Image-only 'scanned' documents paired with their text source, for measuring OCR accuracy against a known ground truth.
Documented properties for this file: searchable text PDF. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Document fixtures list their internal structure (pages, fields, tracked changes, embedded objects) in the spec table. Test extractors, converters, and OCR against that known structure, and compare searchable↔scanned or format-twin companions when present.
Run OCR on the scanned or image twin and score the output against the searchable or text ground truth on this page; the documented rotation, grain, and text content are the reference.
Code examples
import pdfplumber # pip install pdfplumber
with pdfplumber.open("handwritten-searchable.pdf") as pdf:
print(len(pdf.pages), "pages")
print(pdf.pages[0].extract_text())Related files
- pdfOCR Domain — ID Card Mock (Searchable PDF)A searchable PDF id card mock with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfOCR Domain — Invoice (Searchable PDF)A searchable PDF invoice with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfOCR Domain — Multi-column Article (Searchable PDF)A searchable PDF multi-column article with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfOCR Domain — POS Receipt (Searchable PDF)A searchable PDF pos receipt with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfOCR Domain — Utility Bill (Searchable PDF)A searchable PDF utility bill with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfAirline Boarding Pass (PDF)An airline boarding pass laid out as a landscape PDF — passenger, flight, gate, seat, PNR, and a decorative barcode strip. A fixture for testing document-layout parsing and field extraction. Synthetic; the barcode is decorative.

Generated by generation/docs_ocr_wave_b.py. Free for any use, no attribution required — license.