VTT — UTF-8 With BOM Before the Signature
WebVTT with a byte-order mark ahead of the mandatory WEBVTT signature. The spec explicitly allows this, and browsers accept it — but a naive check that the file literally starts with the six bytes 'WEBVTT' rejects a perfectly valid file. The mirror image of the missing-header corrupt fixture: one is valid and often refused, the other is invalid and often accepted.
WEBVTT
1
00:00:01.000 --> 00:00:03.000
Novus Examples — timed-text fixture
2
00:00:03.500 --> 00:00:06.000
Second cue, two lines:
this is the second line
3
00:00:06.500 --> 00:00:09.000
Third cue with <i>italic</i> markup
4
00:00:09.500 --> 00:00:12.000
Fourth cue — em dash, curly “quotes”, ellipsis…
5
00:00:12.500 --> 00:00:15.000
Final cue.
Specifications
- Encoding
- UTF-8
- Byte Order Mark
- yes
- Signature
- WEBVTT, preceded by U+FEFF
- Cues
- 5
- Spec Compliant
- yes — the WebVTT spec permits a leading BOM
What is a .vtt file?
WebVTT (VTT) is the W3C subtitle and caption format used by the HTML5 <track> element for timed text on the web. It extends the SubRip model with cue settings, positioning, styling, and metadata, and requires a WEBVTT header. It is the standard format for browser-based captions.
How to use this file
Use an example VTT to test HTML5 <track> caption rendering, cue-setting and positioning parsers, and converters between WebVTT and SRT.
How to use this file for testing
“VTT — UTF-8 With BOM Before the Signature” is a deterministic Novus Examples fixture for Caption encoding, Subtitle parsing, Video QA. The same cue list written with a UTF-8 BOM, without one, in UTF-16, with CRLF and with bare LF, plus non-Latin scripts and right-to-left text — for finding the parser that assumed captions are always LF-terminated ASCII.
Documented properties for this file: 5 cues · UTF-8. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Media fixtures are short and synthetic by design. Prefer waveform or transcript ground truth in the same group when measuring ASR, trim, upscale, or sync tools; do not assume broadcast-quality masters.
Code examples
<video controls src="clip.mp4">
<track kind="captions" srclang="en" label="English" src="cues-utf8-bom.vtt" default>
</video>Related files
- sccSCC — Broadcast CEA-608 CaptionsBroadcast closed captions as Scenarist SCC: SMPTE timecodes followed by hexadecimal CEA-608 byte pairs, with odd parity set on every byte as the standard requires. Uses pop-on mode — RCL to load, ENM to clear non-displayed memory, a preamble address code for row 15, the character pairs, then EOC to display. This is the fixture that exposes decoders which treat captions as text: the parity bits, the two-byte control codes and the frame-accurate timecodes all have to be handled.

- vttVTT — Positioned and Aligned CuesWebVTT cues carrying the full positioning grammar — line, position, align and size — plus named cue identifiers instead of numbers. Captions are placed at four different points in the frame, which is what real subtitles do to avoid covering burned-in text. Parsers that treat everything after the timestamp as cue text will fold these settings into the visible caption.

- vttVTT — STYLE Blocks, Regions and Inline MarkupWebVTT exercising the parts of the format beyond plain text: a STYLE block with ::cue selectors, a named REGION, voice spans, cue classes, and inline bold/italic/underline/ruby markup. A parser that handles only timestamps and text will render the tag names as visible characters.

- assASS — Karaoke Timing and Override TagsAdvanced SubStation Alpha using the features that distinguish it from SubRip: per-syllable \k karaoke timings in centiseconds, a full V4+ style definition with primary and secondary colours, and an \an override that repositions a line. Converting this to SRT necessarily loses all of it, which makes it a good test of whether a converter warns about that or drops it silently.

- csvCSV — Cue Timing ReportPer-cue timings and text metrics for the shared cue list, as a caption QC tool would export them: start, end, duration, line count and character count. Useful as the expected output when testing a caption analyser, and as a quick way to check reading-rate calculations against known values.

- dfxpDFXP — TTML Under Its Other ExtensionTimed text as DFXP — byte-for-byte a TTML document, just under the extension Netflix, Adobe and older captioning tool-chains still use. Same tt root, same head/body/div/p structure, same styling and region attributes. It exists to check that a caption parser dispatches on document content rather than on the file extension: a parser that accepts .ttml but rejects an identical .dfxp is exactly the bug this catches.

Generated by generation/video_timedtext.py. Free for any use, no attribution required — license.