Redaction — SAMPLE PII Unredacted (PDF)
A PDF with clearly labelled fictional SAMPLE PII fields — the unredacted twin for testing redaction tooling.

Rendered preview of the pdf file (1.7 KB). Download above for the original.
Specifications
- Role
- unredacted source
- Note
- fictional SAMPLE PII only
Testing contract
Expected to pass- Scenario
- Exercise Redaction — SAMPLE PII Unredacted (PDF) in its redaction workflow. A PDF with clearly labelled fictional SAMPLE PII fields — the unredacted twin for testing redaction tooling.
- Expected result
- PDF contains 1 pages; first-page text begins `SAMPLE PII Document (UNREDACTED) Name: Jordan Rivera Email: jordan.rivera@example.com Phone: (555) …`. Declared feature checks: role=unredacted source. The declared comparison counterpart is doc-redact-pii-black; preserve the stated difference instead of expecting the container bytes to match.
What is a .pdf file?
PDF (Portable Document Format) is a page-oriented format that fixes layout, fonts, and vector and raster graphics so a page renders identically anywhere. A `%PDF-` header is followed by numbered objects, a cross-reference table mapping each to a byte offset, and a trailer; edits append incremental updates rather than rewrite the file. Page content is a stream of drawing operators, so a PDF holds no words or paragraphs, only positioned glyph runs. Adobe released it in 1993 and gave it to ISO as ISO 32000-1 in 2008.
How to use this file
Use an example PDF to test text extraction, rendering, metadata parsing, AcroForm handling, and OCR pipelines: checking that extraction reconstructs reading order from glyph positions, that a scanned page yields no text, and that an incremental update leaves earlier revisions in the file.
How to use this file for testing
“Redaction — SAMPLE PII Unredacted (PDF)” is a deterministic Novus Examples fixture for PDF editor testing, OCR testing. Form PDFs, bookmarked documents, scanned pairs, and deliberately corrupt files for exercising PDF editors, parsers, and fillers.
Documented properties for this file: unredacted source. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.
Document fixtures list their internal structure (pages, fields, tracked changes, embedded objects) in the spec table. Test extractors, converters, and OCR against that known structure, and compare searchable↔scanned or format-twin companions when present.
Code examples
import pdfplumber # pip install pdfplumber
with pdfplumber.open("pii-unredacted.pdf") as pdf:
print(len(pdf.pages), "pages")
print(pdf.pages[0].extract_text())Related files
- jpgRedaction — SAMPLE PII Unredacted (JPEG)JPEG twin of the unredacted SAMPLE PII page — for image-based redaction and OCR pipelines.

- pdfAirline Boarding Pass (PDF)An airline boarding pass laid out as a landscape PDF — passenger, flight, gate, seat, PNR, and a decorative barcode strip. A fixture for testing document-layout parsing and field extraction. Synthetic; the barcode is decorative.

- pdfImage-Only 'Scanned' PDF (OCR twin)An image-only PDF containing a rasterised 'scan' of the simple document, with no text layer. Paired with the text version so you can score OCR output against a known ground truth.

- pdfOCR Domain — Handwritten-print Form (Searchable PDF)A searchable PDF handwritten-print form with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfOCR Domain — ID Card Mock (Searchable PDF)A searchable PDF id card mock with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

- pdfOCR Domain — Invoice (Searchable PDF)A searchable PDF invoice with selectable text — the OCR ground-truth twin of the scanned JPEG. Fictional SAMPLE content for measuring OCR accuracy.

Generated by generation/docs_ocr_wave_b.py. Free for any use, no attribution required, license.