JSON — Timed-Text Suite Index
A machine-readable index of all 67 timed-text, streaming-manifest and ad-signalling fixtures in this wave, with id, format, path and byte size for each. Useful as a work-list when running a parser across the whole suite, and as a manifest to diff against after regenerating.
{
"suite": "video-timed-text",
"generator": "generation/video_timedtext.py",
"addedIn": "v0",
"count": 67,
"note": "Machine-readable index of the timed-text, streaming-manifest and ad-signalling fixtures. Every file listed here is deterministic plain text; re-running the generator reproduces them byte for byte.",
"files": [
{
"id": "vid-tt-enc-utf8-lf",
"format": "srt",
"subcategory": "captions-encoding",
"path": "/files/video/captions-encoding/cues-utf8-lf.srt",
"bytes": 352,
"title": "SRT — UTF-8, LF"
},
{
"id": "vid-tt-enc-utf8-bom-lf",
"format": "srt",
"subcategory": "captions-encoding",
"path": "/files/video/captions-encoding/cues-utf8-bom-lf.srt",
"bytes": 355,
"title": "SRT — UTF-8 with BOM, LF"
},
{
"id": "vid-tt-enc-utf8-crlf",
"format": "srt",
"subcategory": "captions-encoding",
"path": "/files/video/captions-encoding/cues-utf8-crlf.srt",
"bytes": 372,
"title": "SRT — UTF-8, CRLF"
},
{
"id": "vid-tt-enc-utf8-bom-crlf",
"format": "srt",
"subcategory": "captions-encoding",
"path": "/files/video/captions-encoding/cues-utf8-bom-crlf.srt",
"bytes": 375,
"title": "SRT — UTF-8 with BOM, CRLF"
},
{
"id": "vid-tt-enc-utf16le",
"format": "srt",
"subcategory": "captions-encoding",
"path": "/files/video/captions-encoding/cues-utf16le.srt",
"bytes": 686,
"title": "SRT — UTF-16 LE with BOM"
},
{
"id": "vid-tt-enc-utf16be",
"format": "srt",Specifications
- Entries
- 67
- Suite
- video-timed-text
- Deterministic
- yes
What is a .json file?
JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.
How to use this file
Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.
How to use this file for testing
“JSON — Timed-Text Suite Index” is a deterministic Novus Examples fixture for Video QA, Subtitle parsing, Conversion testing. Short 270p clips with known scene cuts, freezes, black frames, flash frames, and burned-in captions — for detection and QC tooling.
Documented properties for this file: 67 entries. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Media fixtures are short and synthetic by design. Prefer waveform or transcript ground truth in the same group when measuring ASR, trim, upscale, or sync tools; do not assume broadcast-quality masters.
Code examples
import json
with open("timed-text-index.json") as f:
data = json.load(f)
print(type(data), len(data))Related files
- csvCSV — Cue Timing ReportPer-cue timings and text metrics for the shared cue list, as a caption QC tool would export them: start, end, duration, line count and character count. Useful as the expected output when testing a caption analyser, and as a quick way to check reading-rate calculations against known values.

- assASS — Karaoke Timing and Override TagsAdvanced SubStation Alpha using the features that distinguish it from SubRip: per-syllable \k karaoke timings in centiseconds, a full V4+ style definition with primary and secondary colours, and an \an override that repositions a line. Converting this to SRT necessarily loses all of it, which makes it a good test of whether a converter warns about that or drops it silently.

- dfxpDFXP — TTML Under Its Other ExtensionTimed text as DFXP — byte-for-byte a TTML document, just under the extension Netflix, Adobe and older captioning tool-chains still use. Same tt root, same head/body/div/p structure, same styling and region attributes. It exists to check that a caption parser dispatches on document content rather than on the file extension: a parser that accepts .ttml but rejects an identical .dfxp is exactly the bug this catches.

- jsonJSON — Caption Cue ListThe same five cues as a plain JSON array with float second timings — the shape most caption pipelines use internally between parsing one format and writing another. Handy as the expected intermediate when testing a converter, since it removes timestamp-formatting differences from the comparison.

- lrcLRC — Enhanced Word-Level TimingEnhanced LRC with per-word timings in angle brackets alongside the usual per-line timestamps — the format karaoke and lyric-sync apps consume. Includes the standard metadata tags, an offset field, and a final empty timestamp that clears the display. Simple LRC parsers read only the line timings and silently render the word markers as visible text.

- sccSCC — Broadcast CEA-608 CaptionsBroadcast closed captions as Scenarist SCC: SMPTE timecodes followed by hexadecimal CEA-608 byte pairs, with odd parity set on every byte as the standard requires. Uses pop-on mode — RCL to load, ENM to clear non-displayed memory, a preamble address code for row 15, the character pairs, then EOC to display. This is the fixture that exposes decoders which treat captions as text: the parity bits, the two-byte control codes and the frame-accurate timecodes all have to be handled.

Generated by generation/video_timedtext.py. Free for any use, no attribution required — license.