Skip to content
Novus Examples
pdf2.8 KB

Juniper Workspace Cloud: Mid-cycle upgrade invoice PDF

Mid-cycle upgrade invoice PDF for Juniper Workspace Cloud. A one-page invoice for INV-2609-02 showing all three lines of a mid-cycle upgrade: the 49.00 Team recurring charge, the (24.50) credit for the unused days and the 49.50 Studio charge, totalling 74.00.

Rendered preview of Juniper Workspace Cloud: Mid-cycle upgrade invoice PDF

Rendered preview of the pdf file (2.8 KB). Download above for the original.

Specifications

Document Set
saas
Industry
software
Source Kit
saas-subscriptions
Synthetic
true
As Of
2026-09-08
Pages
1
Lines
3
Invoice Id
INV-2609-02
Total
74.00

Testing contract

Expected to pass
Scenario
Extract the invoice lines and recompute the proration from the plan rates on plan-catalog.csv.
Expected result
The three line amounts are 49.00, (24.50) and 49.50, they sum to the stated invoice total 74.00, and each proration amount equals its plan rate times 15/30 rounded half up.

What is a .pdf file?

PDF (Portable Document Format) is a page-oriented format that fixes layout, fonts, and vector and raster graphics so a page renders identically anywhere. A `%PDF-` header is followed by numbered objects, a cross-reference table mapping each to a byte offset, and a trailer; edits append incremental updates rather than rewrite the file. Page content is a stream of drawing operators, so a PDF holds no words or paragraphs, only positioned glyph runs. Adobe released it in 1993 and gave it to ISO as ISO 32000-1 in 2008.

How to use this file

Use an example PDF to test text extraction, rendering, metadata parsing, AcroForm handling, and OCR pipelines: checking that extraction reconstructs reading order from glyph positions, that a scanned page yields no text, and that an incremental update leaves earlier revisions in the file.

How to use this file for testing

“Juniper Workspace Cloud: Mid-cycle upgrade invoice PDF” is a deterministic Novus Examples fixture for Data import. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.

Documented properties for this file: 1 pages. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.

Document fixtures list their internal structure (pages, fields, tracked changes, embedded objects) in the spec table. Test extractors, converters, and OCR against that known structure, and compare searchable↔scanned or format-twin companions when present.

Code examples

import pdfplumber  # pip install pdfplumber

with pdfplumber.open("invoice-INV-2609-02-proration.pdf") as pdf:
    print(len(pdf.pages), "pages")
    print(pdf.pages[0].extract_text())

Generated by generation/industry_documents.py. Free for any use, no attribution required, license.