Depth Input — Three-Layer Parallax Scene
Three planes translating at 0.6, 2.4 and 6.0 pixels per frame — a 1 : 4 : 10 velocity ratio that defines their relative depth by motion alone. There are no pictorial depth cues to fall back on: no perspective convergence, no shading, no familiar object sizes. A monocular depth model that has learned pictorial cues rather than motion parallax will do poorly here, which is precisely what makes the fixture informative.
Specifications
- Resolution
- 640x360
- Fps
- 24
- Layer Velocities
- 0.6, 2.4, 6.0 px/frame (far, mid, near)
- Velocity Ratio
- 1 : 4 : 10
- Role
- depth-estimation input
- Reference
- vid-depth-parallax-gt
- Crf
- 19
- Base Plate
- parallax
What is a .mp4 file?
MP4 (MPEG-4 Part 14) is a widely supported multimedia container based on ISOBMFF that holds video, audio, subtitles, and metadata, most commonly H.264 or H.265 video with AAC audio. It supports streaming, chapters, and multiple tracks. It is the dominant format for distributing and playing digital video.
How to use this file
Use an example MP4 to test container demuxing, track and codec detection, seeking and streaming, and transcoding or thumbnail-extraction pipelines.
How to use this file for testing
“Depth Input — Three-Layer Parallax Scene” is a deterministic Novus Examples fixture for Depth from video, Frame interpolation, Video QA. Layered scenes whose planes translate at documented, fixed ratios, so relative depth is defined by measurable motion rather than inferred from pictorial cues — giving depth-from-video output an objective reference.
Documented properties for this file: depth-estimation input · 24 fps. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Media fixtures are short and synthetic by design. Prefer waveform or transcript ground truth in the same group when measuring ASR, trim, upscale, or sync tools; do not assume broadcast-quality masters.
Code examples
<video controls preload="metadata" width="640" src="parallax-scene.mp4"></video>Related files
- mp4Base Plate — Depth — Three Parallax PlanesFar, mid and near layers translating at 0.6, 2.4 and 6.0 pixels per frame. Relative depth is defined by the motion ratio rather than guessed from cues, which gives depth-from-video and optical-flow output something objective to be scored against. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 — well above the house CRF 30 — because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

- mp4Base Plate — Pan — City SkylineA skyline scrolling at a constant 3.5 pixels per frame: pure horizontal translation with no rotation or scale change. The lit windows give sparse high-contrast features to track, and the constant velocity means the correct answer for interpolation and optical flow is known exactly. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 — well above the house CRF 30 — because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

- mp4Interpolation Ground Truth — depth-layers at 24 fpsThe full-rate 24 fps reference for the depth-layers interpolation set. The decimated clips in this group drop frames from exactly this sequence, so every frame an interpolator is asked to synthesise has a true original to be scored against — which is the only way to tell invention from reconstruction.

- mp4Interpolation Ground Truth — orbit-solid at 24 fpsThe full-rate 24 fps reference for the orbit-solid interpolation set. The decimated clips in this group drop frames from exactly this sequence, so every frame an interpolator is asked to synthesise has a true original to be scored against — which is the only way to tell invention from reconstruction.

- mp4Interpolation Ground Truth — pan-city at 24 fpsThe full-rate 24 fps reference for the pan-city interpolation set. The decimated clips in this group drop frames from exactly this sequence, so every frame an interpolator is asked to synthesise has a true original to be scored against — which is the only way to tell invention from reconstruction.

- mp4Interpolation Input — depth-layers at 12 fpsThe depth-layers plate decimated to 12 fps by keeping every 2th frame, so 24 of the original 48 frames are missing. Interpolate back to 24 fps and each synthesised frame has an exact counterpart in the ground truth. Larger gaps need genuine motion understanding rather than blending.

Generated by generation/video_ai_suites.py. Free for any use, no attribution required — license.