Skip to content
Novus Examples
srt355 B

SRT — UTF-8 with BOM, LF

UTF-8 with a leading U+FEFF byte-order mark. Windows caption tools emit this constantly, and a parser that does not strip it sees the BOM as part of the first cue number and fails to match its index. Identical cue content across the whole set, so any difference in a parser's output is an encoding bug and nothing else.

Preview — first 21 linessrt
1
00:00:01,000 --> 00:00:03,000
Novus Examples — timed-text fixture

2
00:00:03,500 --> 00:00:06,000
Second cue, two lines:
this is the second line

3
00:00:06,500 --> 00:00:09,000
Third cue with <i>italic</i> markup

4
00:00:09,500 --> 00:00:12,000
Fourth cue — em dash, curly “quotes”, ellipsis…

5
00:00:12,500 --> 00:00:15,000
Final cue.

Specifications

Encoding
UTF-8
Byte Order Mark
yes
Line Endings
LF
Cues
5
Duration
15s

What is a .srt file?

SRT (SubRip) is a plain-text subtitle format listing numbered cues, each with a start and end timecode and one or more lines of text. It is simple, human-readable, and extremely widely supported by players. It carries no styling metadata beyond basic inline tags.

How to use this file

Use an example SRT to test subtitle parsing, timecode handling, and converters that translate between SRT and WebVTT or other caption formats.

How to use this file for testing

“SRT — UTF-8 with BOM, LF” is a deterministic Novus Examples fixture for Caption encoding, Subtitle parsing, Video QA. The same cue list written with a UTF-8 BOM, without one, in UTF-16, with CRLF and with bare LF, plus non-Latin scripts and right-to-left text — for finding the parser that assumed captions are always LF-terminated ASCII.

Documented properties for this file: 5 cues · 15s · UTF-8 · LF. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Media fixtures are short and synthetic by design. Prefer waveform or transcript ground truth in the same group when measuring ASR, trim, upscale, or sync tools; do not assume broadcast-quality masters.

Code examples

<video controls src="clip.mp4">
  <track kind="captions" srclang="en" label="English" src="cues-utf8-bom-lf.srt" default>
</video>

Generated by generation/video_timedtext.py. Free for any use, no attribution required — license.