Skip to content
Novus Examples

Browse by purpose

Every file is tagged with what it's good for testing. Pick a use case to see its fixtures — or filter the full library.

Images & vision

Images & vision

Background removal

Product-on-background compositions, alpha PNGs over transparency, and hard-case edges — fur, mesh, glass refraction, colour-matched backgrounds — for scoring removers and recomposite workflows.

Images & vision

Chart recognition

Bar, line, pie, and scatter charts rendered as images with labelled axes, legends, and known values — for testing chart-extraction and chart-to-data tools, vision models, and OCR of embedded text.

Images & vision

Colour management

Tagged versus untagged images, ICC-profiled samples, and CMYK/wide-gamut cases — for testing colour-management pipelines, profile handling, and conversions that must not assume sRGB.

Images & vision

Colourization

True-colour ground truths paired with greyscale and sepia inputs for measuring photo colourization models and filters against known colour references.

Images & vision

Compression testing

The same source saved at quality 10/50/90 for comparing compression artefacts and quality settings.

Images & vision

ControlNet conditioning

Scenes shipped with the structure-control inputs a guided-generation model consumes — Canny, depth, normal, segmentation, scribble, line-art, and pose — each derived from the same colour ground truth so outputs can be scored against their source.

Images & vision

Deblur testing

Clean references with gaussian, motion, and camera-shake blurred variants — documented blur type and strength — for measuring deblur and sharpening accuracy with PSNR/SSIM against ground truth.

Images & vision

Denoise testing

Noisy images paired with their exact clean reference, with documented noise type, sigma, and seed — everything you need to measure how well a denoise filter recovers the original.

Images & vision

Depth estimation testing

Synthetic depth maps with a known near/far layout and matching colour reference, for depth-conditioned generation and monocular depth-estimation regression.

Images & vision

Visual diff / regression

A/B image twins with controlled pixel changes, anti-alias variants, and UI chrome crops for screenshot diff and visual regression tools.

Images & vision

Edge detection testing

Binary Canny-style edge maps paired with the exact image they were extracted from — measure an edge detector or edge-conditioned generator against a documented reference.

Images & vision

Image enhancement

Clean references with underexposed, overexposed, low-contrast, cast, haze, and heavy JPEG-artefact variants for auto-enhance, dehaze, and artefact-repair tools.

Images & vision

EXIF & orientation testing

Images with embedded EXIF orientation flags for testing whether your tool respects rotation metadata.

Images & vision

Face restore

Synthetic face illustrations (not real people) with JPEG crush, blur, and downscale damage plus masks — for testing face restore and enhancement tools against clean ground truth.

Images & vision

Greyscale testing

Colour originals paired with true-greyscale conversions, grey wedges, and gradient ramps for testing desaturation and tone handling.

Images & vision

Linear 16-bit / HDR-ish

16-bit linear TIFF ramps and twins (not PQ/HDR10 display masters) for deep-colour loaders, truncation, and tone-map pipelines.

Images & vision

ICC / colour profile

Orientation ladders, SAMPLE GPS EXIF, and sRGB vs Adobe RGB tagged images for colour-managed pipelines and EXIF readers.

Images & vision

Image matting

Subject images paired with soft-alpha mattes and trimaps (fur, mesh, glass edges) — for testing image-matting models and alpha extraction against a known ground truth.

Images & vision

Image pipeline QA

Deterministic images with known properties for regression-testing resize, crop, convert, and filter pipelines.

Images & vision

Image segmentation

RGB scenes paired with class or instance masks and documented class IDs and bounding boxes — for scoring semantic and instance segmentation against a known ground truth.

Images & vision

Inpainting

Images with cut-out regions (rects, circles, corners, strips, irregular tears) and matching masks for testing inpainting, generative fill, content-aware fill, and object-removal pipelines.

Images & vision

Line-art extraction testing

Clean line-art extractions (black lines on white) paired with their colour source — for line-art-conditioned generation and line-extraction quality tests.

Images & vision

Image matting

Hair/fur-like soft alpha composites and trimap-style inputs for matting and refined background-removal tools.

Images & vision

Mesh decimation testing

A dense ground-truth mesh plus a low-poly simplification of the same shape, with triangle counts recorded — score a decimation / LOD tool by comparing the two.

Images & vision

Mesh processing QA

Deterministic GLB meshes with known geometry and triangle counts for regression-testing mesh import/export, decimation, repair, and validation pipelines.

Images & vision

Mesh repair testing

A watertight mesh paired with a deliberately broken twin (holes, missing faces, open ends) — the input and intended result for a mesh-repair / hole-filling tool.

Images & vision

Normal map testing

Tangent-space surface-normal maps encoded from a known height/depth field — inputs for normal-conditioned generation, relighting, and normal-estimation checks.

Images & vision

OCR testing

Image-only 'scanned' documents paired with their text source, for measuring OCR accuracy against a known ground truth.

Images & vision

Outpainting

Full-frame references with empty-border inputs and expand masks for testing outpainting, generative expand, and content-aware extend tools.

Images & vision

PBR material testing

Complete PBR material sets — base colour, normal, roughness, metallic, ambient occlusion, and height — for testing material-generation models and physically-based renderers channel by channel.

Images & vision

Photogrammetry testing

Surface scans shipped as clean ground truth plus a degraded capture, for scoring photogrammetry and surface-reconstruction pipelines against a known target.

Images & vision

Point cloud testing

A dense, noise-free point cloud paired with a sparse, noisy scan of the same surface — a reconstruction / denoise target for point-cloud tooling.

Images & vision

Pose estimation testing

OpenPose-style keypoint skeletons for figure subjects, paired with the rendered figure — a pose-control input for guided generation and a target for pose estimators.

Images & vision

Resolution testing

Images from tiny to 4K, plus vertical and extreme aspect ratios, for testing scaling and layout.

Images & vision

Photo restoration

Clean references paired with damaged inputs and damage masks — creases, scratches, dust, water stains, torn corners, and cuts — for measuring photo restoration and damage-repair tools against ground truth.

Images & vision

Scribble conditioning

Loose, hand-drawn-style scribble maps derived from a scene's edges — the sparse control input for scribble-guided image generation.

Images & vision

Segmentation

Multi-class RGB scenes with indexed masks and class maps for semantic/instance segmentation loaders and eval harnesses.

Images & vision

Semantic segmentation testing

Semantic and instance segmentation maps paired with their source scene — the label reference for segmentation-conditioned generation and mask-quality scoring.

Images & vision

Style transfer

Content photos paired with posterize, line-art, paint-like, comic, and geometric style references for neural and filter style-transfer pipelines.

Images & vision

Subject detection

Images with isolated foreground subjects, cluttered backgrounds, and exact masks for testing saliency, subject selection, cropping, and background-removal pipelines.

Images & vision

Super-resolution

Large clean references with deliberately downscaled (and optionally JPEG-degraded) inputs for testing upscalers, super-resolution models, and enlarge-without-blur pipelines.

Images & vision

Texture map testing

Tileable texture channels with documented roles and resolution, for validating texture pipelines, atlas packers, and material import/export.

Images & vision

Thumbnail generation

A resolution ladder and varied aspect ratios for exercising thumbnail and preview generators.

Images & vision

Watermark removal

Clean images paired with synthetic SAMPLE watermarks (logo, text, tile, translucent) and masks — for scoring watermark removers without using real brand marks.

Audio & video

Audio & video

Ad insertion

VAST ad responses and SCTE-35 splice descriptors covering linear break signalling, wrappers, and companion creatives — for testing ad-decision integration and manifest-manipulation logic without calling a real ad server.

Audio & video

ASR testing

Short synthetic digit/tone utterances with transcript JSON and clean↔noise pairs — for testing ASR loaders, WER harnesses, and audio preprocessing.

Audio & video

Audio analysis

Pure tones, sweeps, and noise with known frequencies and levels for testing spectrum, level, and waveform tools.

Audio & video

Audio editing

Synthetic foley, ambience, dialogue, loops, stems, cue sheets, and edge cases with documented timing and signal properties for repeatable editing checks.

Audio & video

Auto-trim testing

Tones with precisely documented leading and trailing silence timestamps — direct fixtures for auto-trim tools.

Audio & video

Caption encoding

The same cue list written with a UTF-8 BOM, without one, in UTF-16, with CRLF and with bare LF, plus non-Latin scripts and right-to-left text — for finding the parser that assumed captions are always LF-terminated ASCII.

Audio & video

Loudness / LUFS

Documented loudness ladders, clipped vs clean twins, and sample-rate conversion sets for loudness meters and normalizers.

Audio & video

Media accessibility

Containers carrying several subtitle and audio tracks at once, with real language tags and the default, forced, hearing-impaired and visual-impaired flags actually set — for testing track pickers, accessibility menus, and the flag handling that only shows up when more than one track exists.

Audio & video

Media metadata

Video files with embedded chapters, full container tags, attached cover art and a SMPTE start timecode — each paired with a deliberately stripped twin, so you can tell a metadata reader that found nothing from one that was never given anything.

Audio & video

Playlist & feed parsing

RSS, Atom, and JSON Feed syndication, an OPML subscription list, and M3U / HLS M3U8 media playlists — for testing feed readers, podcast apps, and media players.

Audio & video

Sound design

Small synthetic sound-effect hits, room tones, layers, and loopable material with exact duration and level contracts for sound-design and media-pipeline QA.

Audio & video

Streaming manifests

HLS playlists and DASH MPDs covering multi-bitrate ladders, byte-range addressing, and segment templates. Served inline with permissive CORS under /files/stream/, so a player can load them from any origin without re-hosting.

Audio & video

Subtitle parsing

The same captions written out across SubRip, WebVTT, ASS/SSA, SBV, TTML and a synced LRC lyric file, with known timings — for exercising subtitle parsers, format converters and burn-in tools against every serialisation.

Audio & video

Subtitle testing

SubRip and WebVTT sidecars plus multi-track containers — forced-narrative, SDH and styled tracks in one file — for testing track selection, default flags and caption rendering in players.

Audio & video

Timeline testing

Equivalent deterministic edit decisions and project structures across common timeline formats, with expected clip order, timing, transitions, and failure behavior.

Audio & video

Video codecs

One source encoded as H.264, HEVC, AV1, VP9, MJPEG, and a ProRes-compatible mezzanine, across 8/10/12-bit and 4:2:0/4:2:2/4:4:4 — for testing decoder support, transcode pipelines, and container remuxing.

Audio & video

Video colourisation

Greyscale, sepia and partially-desaturated clips shipped beside the full-colour original they were derived from, including a 24-patch chart with documented values — so a colourisation model's output can be scored per patch rather than judged by eye.

Audio & video

Video deblur

Motion-blurred and defocused clips produced from a sharp source with a recorded kernel, so deblurring output can be measured against the original rather than compared to another blurred frame.

Audio & video

Video denoise

Clips degraded with documented noise type, sigma, and seed, each shipped beside the exact high-quality source it came from — so a denoiser's output can be scored, not eyeballed.

Audio & video

Depth from video

Layered scenes whose planes translate at documented, fixed ratios, so relative depth is defined by measurable motion rather than inferred from pictorial cues — giving depth-from-video output an objective reference.

Audio & video

Video editing

Short playable clips plus overlays, captions, timelines, colour targets, sync references, and controlled failures for testing browser and desktop editing workflows.

Audio & video

Frame interpolation

Clips with frames removed on a known pattern, paired with the full-rate source — for scoring frame-interpolation and slow-motion tooling against a real reference.

Audio & video

Video matting

Moving subjects over known backgrounds, shipped with the exact alpha matte used to composite them — for scoring background removal and green-screen keying frame by frame.

Audio & video

Video QA

Short 270p clips with known scene cuts, freezes, black frames, flash frames, and burned-in captions — for detection and QC tooling.

Audio & video

Compression artefacts

The same footage at a ladder of quantiser settings from near-lossless down to visibly broken, each paired with its high-quality source — for testing artefact-removal models and for calibrating quality metrics against known encoder settings.

Audio & video

Video segmentation

Hard-edged and anti-aliased subjects in motion, each shipped with the exact per-frame mask used to create it — for scoring video object segmentation and rotoscoping frame by frame instead of spot-checking.

Audio & video

Video stabilisation

Camera shake applied with a recorded motion path over a static original, so a stabiliser's residual motion can be measured rather than estimated.

Audio & video

Video upscaling

Low-resolution clips produced from a documented high-quality source by a recorded filter, so super-resolution output can be measured against the original rather than judged by eye.

Forms & documents

Forms & documents

AcroForm

Fillable PDF forms with text, choice, checkbox, and radio fields for AcroForm tooling.

Forms & documents

Autofill testing

Realistic HTML forms with labelled fields (tax, KYC, shipping, insurance, legal, finance) for exercising browser autofill and form-field detectors.

Forms & documents

Editor testing

Text-based files you can open, edit, and download directly in the browser editor.

Forms & documents

Form parsing

PDF AcroForms with documented fields and live HTML forms for testing form parsers and fillers.

Forms & documents

Form testing

HTML and PDF form files for testing scrapers, autofill, validators, and PDF form fillers across real-life use cases.

Forms & documents

PDF editor testing

Form PDFs, bookmarked documents, scanned pairs, and deliberately corrupt files for exercising PDF editors, parsers, and fillers.

Forms & documents

PDF form filling

AcroForm PDFs and FDF/XFDF data twins for testing PDF form fillers and form-data import.

Forms & documents

Spreadsheet testing

Multi-sheet workbooks with documented formulas and plain sheets for testing spreadsheet parsers and importers.

Data & code

Data & code

API testing

OpenAPI/Swagger specs, GraphQL SDL, JSON Schema, paginated and problem+json error payloads, and webhook samples — for testing API clients, mock servers, contract tests, and schema validators.

Data & code

Calendar parsing

iCalendar (.ics) files with events, timezones, and recurrence — for testing calendar imports and parsers.

Data & code

Code parsing

Short, known-correct source files in many languages with classes, functions, generics, enums, and error handling — for testing parsers, linters, formatters, language detection, and diff viewers.

Data & code

Config parsing

TOML and INI configuration files with nested sections and typed values — for testing config parsers and loaders.

Data & code

Config testing

TOML, INI, YAML, .env, and dotfile configuration samples with nested sections and typed values — for testing config parsers, loaders, and environment tooling.

Data & code

Contact parsing

vCard (.vcf) contact files with names, emails, phones, and addresses — for testing contact importers.

Data & code

CSV parsing

Clean and deliberately messy CSVs — quoted commas, embedded newlines, ragged rows, odd delimiters, and encodings.

Data & code

Data engineering

Columnar (Parquet/ORC/Feather), row (Avro), and messaging (MessagePack/CBOR/Protobuf) formats plus star-schema and log data — for testing ETL, data-lake ingestion, and warehouse loaders.

Data & code

Data import

Realistic faker-generated datasets with documented schemas for testing import and ETL flows.

Data & code

Email parsing

Standards-compliant RFC 822 messages — plain, multipart text+HTML, and with an attachment — plus an MBOX mailbox, for testing header parsing, MIME decoding, attachment extraction, and mailbox splitting.

Data & code

Encoding detection

UTF-8, UTF-8-BOM, UTF-16, and Latin-1 files with documented encodings and line endings for testing charset detection.

Data & code

Feed parsing

Well-formed RSS 2.0 and Atom feeds with multiple entries — for testing feed readers and parsers.

Data & code

Geospatial

GeoJSON, GPX, and KML files with points, lines, polygons, and tracks — for testing map tools, route parsers, and geo importers.

Data & code

Graph data

Node/edge datasets in GraphML and GEXF (directed and undirected, with attributes and weights) — for testing network importers, layout tools, and graph converters.

Data & code

HTML parsing

Standalone documents, email markup, forms, malformed edge cases, and semantic page structures for testing HTML parsers, extractors, sanitizers, and DOM-import pipelines.

Data & code

JSON parsing

Flat, deeply nested, JSON Lines, and intentionally invalid JSON for testing parsers and error handling.

Data & code

Log parsing

Access logs and JSON-lines application logs — for testing log parsers, tailers, and ingestion pipelines.

Data & code

Observability

Structured and plain-text telemetry with known timestamps, levels, request identifiers, and error states for testing log ingestion, correlation, dashboards, and alert pipelines.

Data & code

Performance testing

Documented size, row-count, duration, and resolution ladders for measuring parser, renderer, converter, and upload performance without relying on private production data.

Data & code

Schema / OpenAPI testing

Valid and intentionally invalid OpenAPI/JSON Schema documents plus request/response examples for schema validators and API tooling.

Data & code

Schema validation

JSON Schema documents describing a data shape — for testing validators and schema-aware tooling.

Data & code

Scientific data

Citation catalogs (BibTeX, RIS), chemistry structures (MDL Molfile, PDB), and gridded binary data (NetCDF, FITS) — for testing reference managers, molecule viewers, and scientific-data loaders.

Data & code

Serialization testing

Known records represented as JSON, MessagePack, CBOR, BSON, Protobuf, Arrow, and related formats for testing round trips, schema handling, type preservation, and binary decoders.

Data & code

SQL import

Postgres- and MySQL-flavour SQL dumps with CREATE TABLE and INSERT statements — for testing SQL importers and migrations.

Data & code

Syntax highlighting

Idiomatic hello-world programs and realistic snippets across sixteen languages, exercising comments, string escapes, interpolation, numeric literals, and language keywords — for testing syntax highlighters, editor themes, and tree-sitter grammars.

Data & code

Time-series data

Sensor, market, event, and telemetry series with documented intervals, gaps, duplicates, and timezone behavior for testing importers, resampling, charts, and anomaly pipelines.

Data & code

Time-series data

Irregular timestamps, DST gaps, and duplicate keys for time-series importers and charting libraries.

Data & code

Web assets

Favicons, web app manifests, service workers, robots and sitemap files, Open Graph images, and .well-known resources — for testing web tooling, crawlers, PWA installers, and asset pipelines.

Data & code

Web scraping

Semantic pages, tables, forms, feeds, metadata, and downloadable fixtures for testing extraction and browser automation against stable, purpose-built targets.

Conversion & robustness

Conversion & robustness

Conversion testing

The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.

1,398 filesExplore use case

Conversion & robustness

Error handling

Deliberately corrupt and invalid files, clearly labelled, for testing how your tool fails.

Conversion & robustness

Media pipeline testing

Prompt plans, input media, expected reference outputs, manifests, hashes, and controlled failures for exercising ingest, transform, validate, and export stages together.

AI / ML

AI / ML

Agent workflows

Ordered creator briefs, prompt plans, tool-shaped messages, media inputs, reference outputs, and controlled edge cases joined into complete six-artifact workflow families.

AI / ML

Computer vision

A rendered detection scene annotated in COCO, YOLO, and Pascal-VOC formats — for testing annotation loaders, format converters, and vision pipelines against a known image.

AI / ML

Embeddings

The same texts as a normalised vector set in JSON, Parquet, and NumPy .npy — for testing vector stores, similarity search, and embedding loaders.

AI / ML

ML training data

Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.

AI / ML

Model evaluation

Benchmark results, confusion matrices, ROC curves, and classification reports in CSV and JSON — for testing eval dashboards, metric parsers, and leaderboard importers.

AI / ML

Model inference testing

Small, validated ONNX graphs paired with named JSON inputs and expected tensor outputs. Use them to test runtime loading, dtype and shape handling, broadcasting, dynamic batches, reduction behavior, and numeric tolerances without downloading a production-sized model.

AI / ML

NLP datasets

Sentiment, NER, chat, instruction-tuning, QA, summarization, and translation data in JSON Lines and JSON — for testing NLP loaders, tokenizers, and format converters.

AI / ML

Object detection

Synthetic scenes with exact class labels, bounding boxes, and matching COCO, YOLO, or Pascal VOC annotations for testing detection loaders, converters, and evaluation code.

AI / ML

Prompt engineering

A templated prompt library in JSON Lines with tasks, tags, and placeholders — for testing prompt-management tools and JSONL parsers.

AI / ML

Prompt testing

Vendor-neutral plans and provider-compatible message fixtures paired with explicit inputs, assertions, failure cases, and expected results for repeatable prompt regression testing.

AI / ML

Tool calling / function calling

JSONL traces of function/tool calls with arguments and results — for testing agent harnesses and tool routers.

Security & privacy

Security & privacy

Certificate & key testing

Self-signed X.509 certificates (PEM, CRT, DER), a CSR, RSA and Ed25519 keys, an SSH public key, a PKCS#12 bundle, and an htpasswd file — all published sample-only material, for testing certificate parsers, TLS tooling, keystore importers, and PEM/DER decoders.

Security & privacy

JWT / JWKS testing

Unsigned and SAMPLE-signed JWT variants plus JWKS documents — published sample material only, for auth parser tests.

Security & privacy

Metadata testing

Images, audio, video, and documents with documented metadata paired with deliberately stripped versions, for verifying extraction, preservation, redaction, and privacy-scrubbing behavior.

Security & privacy

Privacy export testing

Sample subject-access ZIP packages, DPA/consent templates, and fictional PII — for testing privacy-export parsers and consent UX tooling.

Localization & templates

Localization & templates

Content creation

Compact prompt, audio, and video artifacts for testing a creator pipeline from brief and planning through editing, evaluation, and final export.

Localization & templates

Localization catalogs

Gettext PO/POT, XLIFF, Apple .strings, Flutter .arb, Android strings.xml, .NET .resx, and i18next JSON — with RTL and CJK variants — for testing localization pipelines, translation-memory tools, and catalog converters.

Localization & templates

Internationalization

Parallel text and accented, multi-script content — for testing translation pipelines, Unicode handling, and localization tooling.

Localization & templates

Templates

Professional, neutral starting-point documents you can download and adapt — invoices, resumes, budgets, and more.