
ASR — 0123 Clean (WAV)
Synthetic clean tone sequence encoding digits 0123 (zero one two three). Pair with the noisy twin and transcript JSON for ASR evaluation.
- File
- WAV · Asr · 16000
- Use case
- ASR testingAudio analysis· Paired fixture
Audio test files usually arrive as someone's music clip. These are engineered. Silence sets document leading and trailing trim points; ASR suites pair synthetic digit utterances with transcripts and noise twins; loudness ladders, peak-clip pairs, stereo/mono, sample-rate conversion, and DTMF sequences exercise meters and harnesses. Every file is short to stay within budget and ships with documented sample rate, bit depth, and duration.
Filter audio on Browse · 156 files · 17 subcategories

Synthetic clean tone sequence encoding digits 0123 (zero one two three). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 0123 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 0123 ASR utterance pair — expected text: “zero one two three”.

Synthetic clean tone sequence encoding digits 1357 (one three five seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 1357 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 1357 ASR utterance pair — expected text: “one three five seven”.

Synthetic clean tone sequence encoding digits 24680 (two four six eight zero). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 24680 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 24680 ASR utterance pair — expected text: “two four six eight zero”.

Synthetic clean tone sequence encoding digits 4567 (four five six seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 4567 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 4567 ASR utterance pair — expected text: “four five six seven”.

Synthetic clean tone sequence encoding digits 89 (eight nine). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 89 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 89 ASR utterance pair — expected text: “eight nine”.

Synthetic clean tone sequence encoding digits 987654 (nine eight seven six five four). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 987654 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 987654 ASR utterance pair — expected text: “nine eight seven six five four”.

Short synthetic cue tone for “one” — clean reference for wake-word / digit ASR smoke tests.

Noisy twin of the “one” cue tone (~5 dB SNR).

Short synthetic cue tone for “zero” — clean reference for wake-word / digit ASR smoke tests.

Noisy twin of the “zero” cue tone (~5 dB SNR).

Extra synthetic digit utterance 0000 (zero zero zero zero) — clean twin.

Noisy twin of digit sequence 0000.

Ground-truth transcript for extra ASR utterance 0000.

Extra synthetic digit utterance 1111 (one one one one) — clean twin.

Noisy twin of digit sequence 1111.

Ground-truth transcript for extra ASR utterance 1111.

Extra synthetic digit utterance 2222 (two two two two) — clean twin.

Noisy twin of digit sequence 2222.

Ground-truth transcript for extra ASR utterance 2222.

Extra synthetic digit utterance 3333 (three three three three) — clean twin.

Noisy twin of digit sequence 3333.

Ground-truth transcript for extra ASR utterance 3333.

Extra synthetic digit utterance 4444 (four four four four) — clean twin.

Noisy twin of digit sequence 4444.

Ground-truth transcript for extra ASR utterance 4444.

Extra synthetic digit utterance 5555 (five five five five) — clean twin.

Noisy twin of digit sequence 5555.

Ground-truth transcript for extra ASR utterance 5555.

Extra synthetic digit utterance 6666 (six six six six) — clean twin.

Noisy twin of digit sequence 6666.

Ground-truth transcript for extra ASR utterance 6666.

Extra synthetic digit utterance 7777 (seven seven seven seven) — clean twin.

Noisy twin of digit sequence 7777.

Ground-truth transcript for extra ASR utterance 7777.

Extra synthetic digit utterance 8888 (eight eight eight eight) — clean twin.

Noisy twin of digit sequence 8888.

Ground-truth transcript for extra ASR utterance 8888.

Extra synthetic digit utterance 9999 (nine nine nine nine) — clean twin.

Noisy twin of digit sequence 9999.

Ground-truth transcript for extra ASR utterance 9999.

A 440 Hz sine tone stored as 16-bit PCM — part of a bit-depth ladder (8/16/24-bit) of the identical tone, for testing bit-depth conversion, dithering, and quantisation-noise handling.

A 440 Hz sine tone stored as 24-bit PCM — part of a bit-depth ladder (8/16/24-bit) of the identical tone, for testing bit-depth conversion, dithering, and quantisation-noise handling.

A 440 Hz sine tone stored as 8-bit PCM — part of a bit-depth ladder (8/16/24-bit) of the identical tone, for testing bit-depth conversion, dithering, and quantisation-noise handling.

A six-channel 5.1 surround WAV that plays a 1-second tone in each channel in turn, in the standard WAV order (FL, FR, C, LFE, SL, SR). A fixture for testing multichannel decoding, downmixing, and channel-order handling.

A stereo WAV with a distinct tone in each channel — 440 Hz on the left, 880 Hz on the right — so you can verify channel mapping, panning, and L/R routing unambiguously.

The clip as raw AAC in an ADTS stream — the codec behind most streaming and mobile audio. For testing AAC decoders and remux into MP4.

The clip as AC-3 (Dolby Digital) — the multichannel codec used in DVD/broadcast. Rendered here in stereo; for testing AC-3 decoding and conversion.

The clip as Apple Lossless (ALAC) in an M4A container — lossless, unlike the AAC M4A twin. For testing ALAC decoding and lossless conversion.

The clip resampled to 8 kHz mono and encoded as AMR narrowband — the telephony/voice-note codec. For testing AMR decoding and speech-codec conversion.

The clip as FLAC — free lossless audio compression. Byte-for-byte recoverable to the source PCM; for testing lossless decoders and conversion.

The clip as AAC in an MP4/M4A container (faststart) — Apple's default audio container. For testing M4A parsing and MP4 audio conversion.

The clip as an M4R iPhone ringtone — AAC in an MP4 container with the ringtone extension. For testing that a converter maps M4R↔M4A correctly.

The source clip as constant-bitrate MP3 (LAME, 192 kbps) — the most universally supported lossy audio format. For testing MP3 decoders, players, and conversion.

The same clip as variable-bitrate MP3 (LAME V2) — for testing VBR handling, seeking, and duration estimation against the CBR twin.

The clip as Ogg Vorbis — a royalty-free lossy codec. Browser-playable; for testing Vorbis decoding and Ogg conversion.

The clip as Opus — the modern low-latency codec used by WebRTC and streaming. Browser-playable; for testing Opus decoding and conversion.

The lossless PCM WAV source for the audio conversion set — the same 3-second tone every other codec in this group is encoded from. Use it as the reference when diffing encoders.

The clip as Windows Media Audio (WMA v2) in an ASF container — Microsoft's lossy codec. For testing WMA decoding and conversion to open formats.

A pure 1 kHz sine tone stored as AIFF — Apple's big-endian 16-bit PCM container, 3 seconds, 44.1 kHz mono. A clean reference for testing AIFF decoders and WAV↔AIFF conversion.

A pure 440 Hz sine tone stored as AIFF — Apple's big-endian 16-bit PCM container, 3 seconds, 44.1 kHz mono. A clean reference for testing AIFF decoders and WAV↔AIFF conversion.

A pure 1 kHz sine tone stored as a Sun/NeXT AU file — big-endian 16-bit PCM, 3 seconds, 44.1 kHz mono. A compact reference for AU decoding and format conversion.

A pure 440 Hz sine tone stored as a Sun/NeXT AU file — big-endian 16-bit PCM, 3 seconds, 44.1 kHz mono. A compact reference for AU decoding and format conversion.

A DTMF (touch-tone) dialing sequence dialing a reserved 555-01xx number — each digit is the standard dual-tone pair. A fixture for testing DTMF decoders, Goertzel detectors, and tone analysis.

Synthetic DTMF tone sequence for digits “*9#”. Pair with the transcript JSON.

Synthetic DTMF tone sequence for digits “042”. Pair with the transcript JSON.

Synthetic DTMF tone sequence for digits “1234”. Pair with the transcript JSON.

Synthetic DTMF tone sequence for digits “13579”. Pair with the transcript JSON.

Synthetic DTMF tone sequence for digits “567890”. Pair with the transcript JSON.

Ground-truth digit transcript for the DTMF sequence *9#.

Ground-truth digit transcript for the DTMF sequence 042.

Ground-truth digit transcript for the DTMF sequence 1234.

Ground-truth digit transcript for the DTMF sequence 13579.

Ground-truth digit transcript for the DTMF sequence 567890.

1 s 440 Hz tone at linear amplitude 0.55 (~-5.2 dBFS peak) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.15 (~-16.5 dBFS peak) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.95 (~-0.45 dBFS peak) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.35 (~-9.1 dBFS peak) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.05 (~-26.0 dBFS peak) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.0 (digital silence) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.75 (~-2.5 dBFS peak) — rung on the Wave E loudness ladder.

1 s 440 Hz tone at linear amplitude 0.02 (~-34.0 dBFS peak) — rung on the Wave E loudness ladder.

Real DTS Coherent Acoustics audio payload encoded from the same deterministic 0.7-second mono reference tone for legacy import coverage. Stable P8 artifact p8-convert-legacy-dts.

Real Yamaha SMAF audio payload encoded from the same deterministic 0.7-second mono reference tone for legacy import coverage. Stable P8 artifact p8-convert-legacy-mmf.

Clean 220 Hz tone (amp 0.45) — reference twin for the hard-clipped pair.

Hard-clipped twin of the 220 Hz tone (×3.2 then clip to ±1) for peak/limiter tests.

Clean 440 Hz tone (amp 0.45) — reference twin for the hard-clipped pair.

Hard-clipped twin of the 440 Hz tone (×3.2 then clip to ±1) for peak/limiter tests.

Clean 880 Hz tone (amp 0.45) — reference twin for the hard-clipped pair.

Hard-clipped twin of the 880 Hz tone (×3.2 then clip to ±1) for peak/limiter tests.

Clean 660 Hz tone (amp 0.45) — reference twin for the hard-clipped pair.

Hard-clipped twin of the 660 Hz tone (×3.2 then clip to ±1) for peak/limiter tests.

Clean 330 Hz tone (amp 0.45) — reference twin for the hard-clipped pair.

Hard-clipped twin of the 330 Hz tone (×3.2 then clip to ±1) for peak/limiter tests.

Clean 550 Hz tone (amp 0.45) — reference twin for the hard-clipped pair.

Hard-clipped twin of the 550 Hz tone (×3.2 then clip to ±1) for peak/limiter tests.

A pure 1000 Hz sine tone, 3 seconds, 16-bit 44.1 kHz mono — a clean reference signal for level metering, spectrum analysis, and waveform rendering.

A pure 220 Hz sine tone, 3 seconds, 16-bit 44.1 kHz mono — a clean reference signal for level metering, spectrum analysis, and waveform rendering.

A pure 440 Hz sine tone, 3 seconds, 16-bit 44.1 kHz mono — a clean reference signal for level metering, spectrum analysis, and waveform rendering.

A 5-second linear sweep from 20 Hz to 20 kHz across the full audible range — for testing frequency response, spectrograms, and playback fidelity.

Three seconds of pink noise with a 1/f power spectrum — the standard reference for loudness and room-calibration testing.

Three seconds of white noise with a flat power spectrum — a reference for testing noise handling, gating, and spectral tools.

A 1 kHz sine tone sampled at 16000 Hz — part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

A 1 kHz sine tone sampled at 44100 Hz — part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

A 1 kHz sine tone sampled at 48000 Hz — part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

A 1 kHz sine tone sampled at 8000 Hz — part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

A 1 kHz sine tone sampled at 96000 Hz — part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

Same 440 Hz tone rendered at 16000 Hz sample rate — twin in the 8k/16k/44.1k set.

Same 440 Hz tone rendered at 44100 Hz sample rate — twin in the 8k/16k/44.1k set.

Same 440 Hz tone rendered at 8000 Hz sample rate — twin in the 8k/16k/44.1k set.

880 Hz tone at 16000 Hz — second ladder for resampling QA.

880 Hz tone at 44100 Hz — second ladder for resampling QA.

880 Hz tone at 8000 Hz — second ladder for resampling QA.

A 440 Hz tone driven past full scale and hard-clipped so its peaks are flat-topped — an intentionally clipped signal for testing clip and true-peak detection, distortion metering, and de-clipping tools.

Two seconds of true digital silence (all-zero samples) — a fixture for testing silence detection, auto-trim thresholds, and noise-floor handling.

A two-second melody playing an ascending C-major scale (C4 to C5) with per-note envelopes — a simple musical signal for testing pitch detection, onset detection, and playback.

A short four-note notification chime (a C-major arpeggio with a decaying final note) — a pleasant, recognisable UI sound for testing playback, short-clip handling, and notification pipelines.

A very quiet modulated tone (~peak 0.04) for testing loudness normalization and gain staging.

A 440 Hz stereo tone with a loud left channel and a near-silent right channel — for balance and mono-mix testing.

A clean 440 Hz tone interrupted by a short high-amplitude click at 0.4s — for click/pop removal tools.

440 Hz tone with 0.5s of digital silence in the middle (0.2s + silence + 0.8s). Fixture for mid-gap trim / split tools.

440 Hz tone with 0.5s of digital silence in the middle (0.8s + silence + 0.2s). Fixture for mid-gap trim / split tools.

440 Hz tone with 1.5s of digital silence in the middle (0.4s + silence + 0.4s). Fixture for mid-gap trim / split tools.

440 Hz tone with 0.75s of digital silence in the middle (0.5s + silence + 0.5s). Fixture for mid-gap trim / split tools.

440 Hz tone with 0.25s of digital silence in the middle (0.4s + silence + 0.4s). Fixture for mid-gap trim / split tools.

440 Hz tone with 0.08s of digital silence in the middle (0.35s + silence + 0.35s). Fixture for mid-gap trim / split tools.

A 440 Hz tone preceded by exactly 1 second of digital silence. A direct fixture for auto-trim tools — the tone should start at 1.000s.

A 440 Hz tone preceded by exactly 3 seconds of digital silence. A direct fixture for auto-trim tools — the tone should start at 3.000s.

A 440 Hz tone preceded by exactly 5 seconds of digital silence. A direct fixture for auto-trim tools — the tone should start at 5.000s.

A 440 Hz tone followed by exactly 1 second of digital silence. The tone should end at 2.0s — a direct fixture for testing trailing-silence trimming.

A 440 Hz tone followed by exactly 3 seconds of digital silence. The tone should end at 2.0s — a direct fixture for testing trailing-silence trimming.

A 440 Hz tone followed by exactly 5 seconds of digital silence. The tone should end at 2.0s — a direct fixture for testing trailing-silence trimming.

Mono 440 Hz twin — pair with the stereo file for channel-count / downmix tests.

Stereo twin of the 440 Hz tone (L=mono, R=phase-shifted) for stereo↔mono conversion tests.

Mono 523 Hz twin — pair with the stereo file for channel-count / downmix tests.

Stereo twin of the 523 Hz tone (L=mono, R=phase-shifted) for stereo↔mono conversion tests.

Mono 659 Hz twin — pair with the stereo file for channel-count / downmix tests.

Stereo twin of the 659 Hz tone (L=mono, R=phase-shifted) for stereo↔mono conversion tests.

Mono 784 Hz twin — pair with the stereo file for channel-count / downmix tests.

Stereo twin of the 784 Hz tone (L=mono, R=phase-shifted) for stereo↔mono conversion tests.

The matching FLAC reduced to STREAMINFO and the exact same audio frames, so metadata-removal tests can compare compressed essence byte for byte.

A FLAC whose STREAMINFO and audio frames are accompanied by known Vorbis comments and a 32-pixel PNG front cover.

An MP3 with a full set of ID3v2.3 tags (title, artist, album, date, genre, comment) over a synthetic C-major melody. A fixture for testing tag readers, metadata editors, and music-library importers.

The same MPEG audio frames with leading and trailing ID3 metadata removed, for checking metadata scrubbing without an audio transcode.

An MP3 with deterministic ID3v2.3 text frames and embedded PNG cover art wrapped around unchanged MPEG audio frames.
No — Wave B ASR fixtures are synthetic tone-digit utterances with JSON transcripts. They exercise ASR harnesses and loaders without shipping copyrighted speech.
Documented leading/trailing silence (1–5 s) on pure tones so auto-trim tools can be checked against exact timestamps. Wave E also adds mid-gap silence cases for split/trim tools.
WAV references plus compressed codecs (MP3, FLAC, OGG, and more). Specs list sample rate, bit depth, and duration.
Filter Browse by purpose loudness-testing for the amplitude ladder and clean↔hard-clipped peak pairs. Stereo/mono and sample-rate twins cover channel and resampler QA.
We use Google Analytics and show ads via Adsterra. Non-essential cookies and ad scripts run only after you allow the matching categories. See our cookie policy.