Speech transcript collectionBrowse audio and transcripts

LISTEN · READ · COMPARE

Speech clips across 11 languages

Explore 220 clips with two automatic transcripts each. The models agree closely, but agreement does not establish accuracy.

220clips
11languages
220 clipsAll languages

Listen to the audio

WORD-BY-WORD COMPARISON

Where the transcripts differ

WER = (replaced + missing + added) ÷ reference words
0%Word error rate
0Replaced wordsSubstitutions
0Missing wordsDeletions
0Added wordsInsertions

Aligned words

Normalized for scoring · scroll to inspect the full segment

ReferenceQwen
Match Replaced Missing from Qwen Added by Qwen

Detected changes

WER here measures agreement between Qwen3-ASR and the reference transcript after normalization. For Hindi, the reference is an upstream transcript; for the other languages, it is Parakeet. It is not a human accuracy score.