Eight short clips — 229 words
One clip per condition, recorded speech with and without background noise. "Words dropped" means words the reference had that the transcript never produced (you would not even know they were missing). "Spots to fix by hand" means runs of wrong words collapsed into one edit — a 22-word invented sentence counts as one spot.
On these short clips the tiers land close together. The cloud reference drops more words than small (7 vs 1), but the two need a very different number of hand-fixes (16 vs 2). None of the four is clearly ahead. These are the same eight clips counted on the accuracy page, where the whisper-base tier shows 46 wrong words (the English default, Moonshine base, shows 31). That 46 is the same source counted a different way: for whisper-base, 46 = substitutions + deletions + insertions (34 + 4 + 8). This page instead splits deletions ("words dropped", 4 for base) from the runs of edits you would fix by hand ("spots", 6 for base). Two lenses on one set of 229 words, not two answers.
| Tier | Words dropped | Spots to fix by hand | Model weights to download (MiB, uncompressed) |
|---|---|---|---|
| tiny | 5 | 20 | 38.9601 |
| base (ships by default) | 4 | 6 | 73.3324 |
| small (highest accuracy, optional) | 1 | 2 | 237.5383 |
| cloud reference | 7 | 16 | 0 |
Three long files — 7,385 words
The same four tiers on three longer recordings (about 5.6, 13.1 and 27.9 minutes). Here the gaps open up, and with the retuned engine the local tiers now lead both columns — the cloud reference is shown only for comparison, not as a recommendation.
| Tier | Words dropped | Spots to fix by hand | Model weights to download (MiB, uncompressed) |
|---|---|---|---|
| tiny | 203 | 273 | 38.9601 |
| base (ships by default) | 82 | 133 | 73.3324 |
| small (highest accuracy, optional) | 45 | 79 | 237.5383 |
| cloud reference | 111 | 592 | 0 |
Read the two columns separately. With the retuned engine the local tiers lead both columns on these long files: small needs the fewest hand-fixes (79, against the cloud reference's 592) and drops the fewest words (45, against 111). The cloud reference is printed only as a reference point, not a recommendation — it is a different system that uploads your audio, whereas the local tiers keep it on your device. On the eight short clips, by contrast, the tiers still land close enough that we say "about the same".
Why the cloud row is a reference, not a ranking
We print the cloud reference so you can see the local tiers against a commercial API — but it is not a recommendation, and it is not the same kind of system:
- On the 8 clips the four tiers land close: the cloud reference drops 7 words and needs 16 hand-fixes; our small tier drops 1 and needs 2. The same order of magnitude — that is why we say "about the same".
- On the long files the retuned local engine leads both columns: small drops 45 words and needs 79 hand-fixes, against the cloud reference's 111 and 592. Better on this material — but the cloud reference is a different system that uploads your audio.
- Model download: the cloud reference is 0 MiB but you must send the audio to a server; the local tiers download 38–238 MiB but keep the file on your device.
So the local tiers are not categorically "more accurate" — they are a different trade (privacy and a download, versus handing your audio to a server). We show the cloud row so you can judge that trade yourself, not so you can crown a winner.
About the cloud reference row
We also ran the same audio through a commercial cloud speech API, as a reference point, not a recommendation. Every number on that row was measured on a named service with named settings. If a row like this cannot be written out completely, it does not go on the page.
| Service and API version | Google Cloud Speech-to-Text, v1 — methods speech:recognize (synchronous) and speech:longrunningrecognize (long-running). |
|---|---|
| Method | 8 short clips: the synchronous recognition method. Three long files: the long-running recognition method — the synchronous one rejects audio longer than 1 minute (it returns Sync input too long). |
| Model | model not specified, so the API's default model was used. We did not select Chirp / Chirp 2. |
| Date called | 2026-09-27 (8 clips) and 2026-09-28 (the three long files). |
| Parameters | languageCode: en-US, enableAutomaticPunctuation: true, encoding: MP3. |
| How it was called | From this machine, with Application Default Credentials; the project was named on each call. Long files were uploaded to Cloud Storage first and passed in as a gs:// URI. |
| Audio path | The same MP3 files, not two sets. The local tiers and the cloud API ran on the identical files. The MP3s are re-encodes of the LibriSpeech dev-clean FLAC originals: 16 kHz mono, 64 kb/s for the three long files, 24 kb/s for the 8 short clips (A-clean is 24 kHz / 48 kb/s — the one exception, and the same file for both sides). |
What the figure has to include
- The audio — 11 clips, 7,614 words of English reference text in total. Language forced (
en-US), not detected. - The reference — transcripts that come with the corpus: 8 clips 229 words, L1 823, L2 2,054, L3 4,508.
- The comparison — lower-cased, every non-letter/non-digit character replaced by a space, runs of whitespace collapsed, trimmed. Word-level edit distance, counting substitutions + deletions + insertions.
- The output — the raw model text is in the repository, so you can re-score it with different rules and say so.
On the three long files, the cloud API dropped 111 words and needed 592 hand-fixes; our local small tier dropped 45 and needed 79. On this material the local tiers needed fewer fixes — but the cloud reference is a different kind of system (it uploads your audio) and we print it only as a reference, never a recommendation. The cloud row is the same GCP measurement as before and was not re-run for this update, so treat the long-file cloud comparison as a fixed reference point, not a fresh race.
How the number was produced
One mechanical rule, applied identically to every tier and every clip. This sentence is the evidence, not the explanation — you can run it yourself:
- Normalisation: transcript and reference are lower-cased, every non-letter/non-digit character (including punctuation) becomes a space, runs of spaces collapse, ends are trimmed.
- Alignment: word-level edit distance by dynamic programming, back-tracked into substitutions, deletions (words dropped) and insertions (words added).
- "One spot": a run of consecutive edits counts as one spot; the moment one reference word matches, the run breaks.
- Overall number: weighted — total wrong words ÷ total reference words — not the average of the per-clip rates.
"Words dropped" and "spots to fix by hand" are different counts. A 22-word invented sentence is one spot but 22 dropped-or-wrong words; a file that loses a whole sentence silently can score a low "spots" number while hiding the most damaging error. Always read the two columns together.
Reading boundaries you must know
- tiny on long files loops and repeats. It can fall into a recite/hallucination loop — this row is reported as-is and is not fixed. A side effect is that tiny produces fewer real words on long files (it is looping, not transcribing); do not read that as an advantage.
- tiny on the hardest short clip is not reproducible. The same clip on the same tier, run three times, returned 84.5% / 94.4% / 111.3% error. That one cell is a range, not a point.
- The two batches are not the same material. The 8 clips contain added noise (light, heavy, telephone-band); the three long files are clean read speech. Do not compare the 8-clip column to the long-file column — the long files are easier by construction, not because they were handled better.
- How the audio is cut. Each window is 30 seconds. The engine reads the word timestamps the model already produced and ends every window at a real word boundary, then starts the next window exactly where the previous one stopped — so a word is never chopped across a seam (the usual cause of long-file deletions). A 30-second slice that is literally silent is skipped, but a window is never dropped for being quiet.
What the material is
All 11 clips come from LibriSpeech dev-clean (CC BY 4.0), clear read speech. The text is 19th-century British art criticism with dense proper nouns (Quilter, Frederick, Ithaca, Linnell, Birket, Jingo, Idylls). The cloud API's 748 substitutions land heavily on exactly those names. So the absolute error rates are high for this material and do not generalise to everyday speech — read the relative comparison, not the absolute level.
The long files are built by concatenating complete utterances: L1 is one speaker, L2 two, L3 four. The sentences are butted together with no natural pause, unlike a real long recording.
Where the raw data is
Every number on this page comes from two files in the benchmark repository — the eight-clip edit-load table and the three-long-file edit-load table — alongside the measurement scripts and the reference transcripts, so the arithmetic can be checked or re-run.
- Raw data and the measurement scripts — github.com/supersophia8888-cloud/clipquill-asr-benchmark
- Archived copy with a DOI — 10.5281/zenodo.22826968, the concept DOI: that link always resolves to the newest version of this dataset, all versions.
- The same data as a dataset — huggingface.co/datasets/sophia8888/clipquill-asr-benchmark
For the same eight clips counted as a word-error rate, see the accuracy page.
Run it on your own file
The tool is on the front page of this site. Drop in an audio or video file you already have, pick the spoken language, and read the transcript. Nothing is uploaded — the model runs in your browser tab and the file never leaves your device. English runs on Moonshine base and the other languages on whisper; a larger, more robust tier (whisper-small) can be switched to before running a file — pick it when the audio is noisy.