
PDF with Metadata Stripped
The same one-page PDF content stream with the Info dictionary and XMP stream omitted, for metadata redaction and document-ingestion tests.
- File
- PDF · PDF
- Use case
- Metadata testingConversion testing· Paired fixture
Search files, editable visual templates, and live browser targets from one registry-backed directory. Filtered query views stay crawlable for links but are deliberately noindex; the stable taxonomy pages below remain the canonical search surfaces.
Page 5 of 7; 24 results per page.

The same one-page PDF content stream with the Info dictionary and XMP stream omitted, for metadata redaction and document-ingestion tests.

A three-page PDF mixing US Letter, A4, and landscape Letter page boxes — for testing viewers and converters that assume uniform page size.

A minimal Word document of plain paragraphs with no styling — the simplest valid DOCX for baseline parser testing.

A single-sheet workbook of plain tabular values with no formulas — a baseline fixture for spreadsheet importers.

A pitch-deck slide set saved as legacy binary .ppt (PowerPoint 97-2003, OLE compound file) via LibreOffice. For testing legacy-Office parsers and PPT→PPTX conversion.

A four-slide presentation where every slide carries speaker notes in the notes pane — for testing notes extraction and whether converters preserve the notes alongside the slides.

JPEG twin with SAMPLE PII fields blacked out. Pair with the clear JPEG to score redaction tools.

The same SAMPLE PII document with email, phone, SSN, and address black-boxed. Pair with the unredacted twin to verify redaction completeness.

JPEG twin of the unredacted SAMPLE PII page — for image-based redaction and OCR pipelines.

A PDF with clearly labelled fictional SAMPLE PII fields — the unredacted twin for testing redaction tooling.

A reStructuredText document with directives, a code block, a note admonition, and a table — for testing docutils/Sphinx rendering, reST parsers, and editors.

A UTF-8 file mixing left-to-right and right-to-left scripts (Arabic and Hebrew alongside English) — for testing bidirectional text handling, reordering, and rendering.

A Rich Text Format document exercising bold, italic, underline, colour, headings, and a bullet list — all as plain control words. For testing RTF parsers, text extraction, and RTF→DOCX/PDF conversion.

The simplest possible RTF: one monospace font and a couple of paragraphs, no styling. A minimal baseline for RTF parsers and conversion tools.

A YouTube SBV caption file — start,end timestamp lines followed by caption text — carrying the same three captions, for testing SBV parsers and YouTube caption import.

A Shift-JIS encoded Japanese text file — a multi-byte East-Asian encoding, for testing CJK charset detection and Shift-JIS→UTF-8 conversion.

A one-page agreement with a signature line and fillable AcroForm fields for the signer's name and date. A fixture for testing form-field detection, filling, and signature-workflow tooling.

A single-page PDF with a title and body text — the simplest valid document for testing PDF viewers, parsers, and text extraction. Paired with an image-only scanned twin for OCR testing.

A SubStation Alpha (.ssa) subtitle — the predecessor of ASS — with a V4 styles section and timed Dialogue events, carrying the same captions, for testing SSA parsers and SSA-to-ASS upgrades.

The same café receipt as an image-only 'scan' — no text layer, with a slight rotation and grain so it reads like a photographed receipt. Paired with the searchable version as OCR ground truth.

A realistic café receipt as a searchable PDF with a real text layer — the OCR ground truth paired with an image-only scanned twin, so you can score OCR output against known text.

A SubRip (SRT) subtitle track with three timed cues — the most common subtitle format. Paired with the WebVTT twin for testing subtitle parsers and SRT↔VTT conversion.

A PDF containing a 12-row, 5-column table — for testing table extraction and layout parsing.

SAMPLE PDF documenting tagged/accessibility intent (alt-text-block) in specs — bookmark outline present; not a full PDF/UA export.
We use Google Analytics and show ads via Adsterra. Non-essential cookies and ad scripts run only after you allow the matching categories. See our cookie policy.