Depth Ground Truth — Layer Depth Map
The depth map for the parallax plate, encoded brighter-is-nearer across three layers. These values are not an estimate — they are derived from the layer velocities the generator used, so the depth ordering is exact by construction. Relative depth is what matters here: the layers move at 0.6, 2.4 and 6.0 pixels per frame, a 1 : 4 : 10 ratio.
Browser-playable test video · 2 seconds
Specifications
- Resolution
- 640x360
- Fps
- 24
- Encoding
- brighter = nearer
- Layers
- 3
- Far Value
- 30
- Mid Value
- 120
- Near Value
- 230
- Derived From
- known layer velocities, not estimated
- Role
- ground truth
- Crf
- 14
- Base Plate
- parallax
Testing contract
Reference control- Scenario
- Use this clip as the depth reference when scoring depth-from-video output for the scene in this group.
- Expected result
- Depth is encoded as brightness - brighter is nearer - across 3 layers at fixed values: far 30, mid 120, near 230. Crucially these were derived from the known layer velocities rather than estimated, so the reference is exact rather than another model's opinion, and relative depth can be scored as an ordering even by a model that predicts only up to scale.
What is a .mp4 file?
MP4 (MPEG-4 Part 14) is the dominant container for digital video, holding video, audio, subtitle, and metadata tracks in a tree of typed boxes. `ftyp` declares the brand, `moov` carries the sample tables that make seeking possible, and `mdat` holds the media; when an encoder writes `moov` last, playback cannot begin until the file has fully downloaded, which a faststart remux fixes. It derives from Apple's QuickTime format, generalized by ISO into the ISOBMFF base that MOV, 3GP, and HEIF share.
How to use this file
Use an example MP4 to test box parsing and track detection, seeking, range-request streaming, and transcode or thumbnail pipelines: verifying that a file with a trailing `moov` box is still handled, and that codec support is checked per track rather than inferred from the extension.
How to use this file for testing
“Depth Ground Truth — Layer Depth Map” is a deterministic Novus Examples fixture for Depth from video, Video QA. Layered scenes whose planes translate at documented, fixed ratios, so relative depth is defined by measurable motion rather than inferred from pictorial cues, giving depth-from-video output an objective reference.
Documented properties for this file: ground truth · 24 fps · brighter = nearer. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.
Media fixtures are short and synthetic by design. Prefer waveform or transcript ground truth in the same group when measuring ASR, trim, upscale, or sync tools; do not assume broadcast-quality masters.
Code examples
<video controls preload="metadata" width="640" src="depth-map.mp4"></video>Related files
- webmAlpha Channel — Opaque VP9 TwinThe same codec and container with NO alpha plane, as the control for the alpha clip in this group. Useful for checking that alpha detection reads the pixel format rather than assuming every WebM is transparent — and for confirming a compositing bug is in the alpha handling rather than in the player.

- webmAlpha Channel — VP9 in WebMVideo with a real per-pixel alpha channel: VP9 yuva420p in WebM, the one alpha path browsers decode natively. Composite it over a page background and the transparency is genuine, not a chroma key. Roughly 90% of each frame is fully transparent. Two traps this file exists to expose. First, auto-alt-ref must be disabled at encode time or libvpx silently drops the alpha plane, producing a valid file with no transparency and no error. Second, WebM stores VP9 alpha in BlockAdditional and signals it with AlphaMode=1, so FFmpeg's NATIVE vp9 decoder reports pix_fmt yuv420p and decodes fully opaque — you must force `-c:v libvpx-vp9` to see the alpha at all. Probing this file with default settings and concluding it has no alpha is the expected mistake.

- assASS — Karaoke Timing and Override TagsAdvanced SubStation Alpha using the features that distinguish it from SubRip: per-syllable \k karaoke timings in centiseconds, a full V4+ style definition with primary and secondary colours, and an \an override that repositions a line. Converting this to SRT necessarily loses all of it, which makes it a good test of whether a converter warns about that or drops it silently.

- mp4Base Plate — Colour — 24-Patch Chart in MotionA 24-patch chart drifting slowly so it is genuinely moving footage rather than a still. Patch values are fixed and documented, so colourisation and colour-management output can be measured per patch. The values are fictional and not a reproduction of any licensed reference chart. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 — well above the house CRF 30 — because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

- mp4Base Plate — Depth — Three Parallax PlanesFar, mid and near layers translating at 0.6, 2.4 and 6.0 pixels per frame. Relative depth is defined by the motion ratio rather than guessed from cues, which gives depth-from-video and optical-flow output something objective to be scored against. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 — well above the house CRF 30 — because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

- mp4Base Plate — Detail — Siemens Star and Frequency WedgesA 36-spoke Siemens star under a slow zoom, plus bar-pair wedges from 16 pixels down to 2. Detail runs right down to the Nyquist limit, which is exactly where super-resolution and denoise either recover structure or invent it. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 — well above the house CRF 30 — because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

Generated by generation/video_ai_suites.py. Free for any use, no attribution required, license.