Skip to content
Novus Examples
pdf2.5 KB

PDF with Mixed Page Sizes

A three-page PDF mixing US Letter, A4, and landscape Letter page boxes — for testing viewers and converters that assume uniform page size.

Rendered preview of PDF with Mixed Page Sizes

Rendered preview of the pdf file (2.5 KB). Download above for the original.

Specifications

Pages
3
Sizes
Letter, A4, Landscape Letter
Suite
doc-advanced

What is a .pdf file?

PDF (Portable Document Format) is a page-oriented document format that preserves fixed layout, fonts, vector and raster graphics, and text across platforms. It can also embed forms, annotations, attachments, and digital signatures. It is the de facto standard for finished, print-ready documents.

How to use this file

Use an example PDF to test text extraction, rendering, page-count and metadata parsing, form-field handling, and conversion or OCR pipelines.

How to use this file for testing

“PDF with Mixed Page Sizes” is a deterministic Novus Examples fixture for PDF editor testing, Conversion testing. Form PDFs, bookmarked documents, scanned pairs, and deliberately corrupt files for exercising PDF editors, parsers, and fillers.

Documented properties for this file: 3 pages. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Document fixtures list their internal structure (pages, fields, tracked changes, embedded objects) in the spec table. Test extractors, converters, and OCR against that known structure, and compare searchable↔scanned or format-twin companions when present.

Code examples

import pdfplumber  # pip install pdfplumber

with pdfplumber.open("mixed-page-sizes.pdf") as pdf:
    print(len(pdf.pages), "pages")
    print(pdf.pages[0].extract_text())

Generated by generation/docs_b2_extras.py. Free for any use, no attribution required — license.