ClipQuill

Video transcribe the language count is not the accuracy

To video transcribe is to turn the speech in a video file into written text, and the pages ranking for the term sell it the same way: a language count, plus one accuracy figure — 99%, 99.9%, 95%+. Both of those stop meaning what you think they mean when the spoken language is not written with spaces. On one 23.088-second Mandarin clip with 97 reference characters, the same model output scored 43.3% wrong compared character-for-character against a Simplified reference, 7.2% wrong once the script was converted, and 6.2% wrong once digits were normalised as well. Same audio, same output, three numbers — because the model writes Traditional characters and the reference was Simplified. Every figure below is from our own run and ships as CSV.

What was measured

Nothing about the audio or the output changed between these three numbers - only the reference. A language count tells you none of this, which is why 99 languages and 99% accuracy can both be true at once.
Nothing about the audio or the output changed between these three numbers - only the reference. A language count tells you none of this, which is why 99 languages and 99% accuracy can both be true at once.

This is a single measurement, published as a single measurement. It is not a rate and it is not an average over anything.

  • Clip length — 23.088 s of Mandarin speech
  • Reference — 97 characters, transcribed by hand from the audio
  • Spoken language — set to zh by hand, not auto-detected
  • Model — onnx-community/whisper-base, quantised ONNX, served from this site
  • Machine — 2-core AMD EPYC 9754, 3.94 GB RAM, real (non-headless) Chrome window driven over the Chrome DevTools Protocol
  • Date — 2026-09-18, against the live public site
  • Script conversion — the zhconv library, in the scorer, not in the model

Source: chinese-cer-one-clip.csv and the cer-zh.py scorer in the published repository. The scorer aligns the output against the reference by dynamic programming and counts substitutions, deletions and insertions separately, so the three numbers below are not three different opinions about the same comparison — they are the same comparison run three times with a different amount of normalisation applied first.

Three ways to score the same output

The model returned Traditional Chinese. The reference was written in Simplified Chinese. Comparing the two directly means every character that is a legitimate variant of the reference character is counted as an error, which is why the first row is as large as it is.

  • Raw comparison, Traditional against Simplified — 42 of 97 characters wrong, 43.3%
  • After converting the output to Simplified — 7 of 97 characters wrong, 7.2%
  • After converting the script and normalising digits — 6 of 97 characters wrong, 6.2%
  • Deletions and insertions, in all three runs — 0 and 0

The third row differs from the second because the model writes a number as a numeral where the reference spells it out. Normalising digits to their written form removes one more mismatch. Nothing else was normalised: no punctuation was stripped beyond the scorer's fixed list, no fuzzy matching, no edit-distance threshold, and no manual correction of any character.

One figure in our method notes does not reconcile exactly with the CSV, and we are publishing that rather than quietly picking one. The method note records 36 of the 97 characters differing by script alone; the CSV shows 42 wrong before conversion and 7 after, a difference of 35. We do not have a mechanism for the one-character gap and we are not going to invent one. Both figures are reproduced here as recorded. Source: chinese-cer-one-clip.csv; the 36 appears in METHOD.md.

The six characters that stayed wrong

After both normalisations, six characters out of ninety-seven were still wrong. What they are is more useful than how many there are.

  • They sit at five places in the text, not six. One two-character word has both of its characters wrong, so two of the six errors are the same word counted twice.
  • Every one is a substitution. There were no deletions and no insertions anywhere in the run, which means the model never dropped a character and never inserted one. The transcript is the right length and the right shape.
  • Five of the six are substitutions between characters that sound the same, or nearly the same, in Mandarin. The sixth is not: it is a different sound with a different tone.
  • The errors cluster at the word level, not the sentence level. Nothing is missing from the transcript; a small number of characters were replaced by a plausible neighbour.

That pattern matters for anyone judging a transcript by eye. A transcript with a 43.3% character mismatch against a Simplified reference looks unusable, and it is not: it is complete, correctly ordered, correctly sized, and wrong in six places. The failure mode here is substitution, not omission — which is the difference between a transcript you proofread and a transcript you throw away.

We are publishing the six as a count of characters out of a known total, not as a percentage, and not as an accuracy figure. One clip is a count. Source: chinese-residual-errors.md, which lists all six with the reference character, the character the model produced, and a note. That file is in the repository; the characters themselves are not reproduced on this page.

What the pages ranking for this term leave out

We read the organic results for video transcribe on 2026-09-23 before writing this page, through two independent search engines — one with the market set to en-US, one through a US-English endpoint. Their first pages overlapped on 8 of 10 results, and the union covered 10 distinct sites. Our fetcher could read 7 of them; the other 3 answered with a bot-check interstitial and are recorded below as not read, not as “not stated”.

  • Language count is the headline everywhere. Across the seven pages we could read, the same sentence appears in seven different costumes: 25 languages, 63 languages, 90+, 100+. One of them is explicit about what its 25 are — English plus 24 European languages. Not one of them tells you what happens when the language you need is not written in the Latin alphabet.
  • Accuracy is a constant, not a measurement. Three of the seven attach a number: 99% and 99.9% on one page, repeated five times, and 95%+ on another. None of them says on what audio, in what language, on how many clips, or with what scorer. A fourth offers two unnamed tiers, “standard accuracy” and “highest accuracy”, with no figure attached to either.
  • The one page that qualifies its number still stops short. A fifth page is the most careful of the seven: it says a large model of this class is near-human, puts the typical English error rate in the low single digits, extends that only to other well-supported languages, and warns that noise, crosstalk, strong accents and quiet recordings all make it worse. That is as honest as the first page of results gets — and it still never mentions script, character counting, or what a “word” is in a language that does not write words apart.
  • Nobody explains the metric. The word word in “word error rate” is doing a lot of work, and it does not survive contact with Chinese, Japanese, Korean or Thai. Not one of the seven pages says how it counts errors, what its reference text was, or whether it converts scripts before scoring. Ours does, in a file anyone can re-run.
  • Nobody publishes a script-handling figure. No page states what its output looks like for a Simplified-Chinese video, whether it emits Traditional characters, or what a reader should do about it. This is the whole gap this page exists to fill.
  • No page publishes its data. Not one of the seven links to a CSV, a dataset, a script, or a DOI. Their numbers cannot be re-run, checked or contested.

To be exact about the limits of this comparison: it covers the organic first page returned for this term on the day of writing, through the two engines named above. Where a page did not state something we record it as not stated. Where our fetcher was blocked, we record the page as not read and draw no conclusion about its content. We have not tested any of these seven tools, and nothing on this page is a claim about their accuracy.

What this means if your video is in Chinese

  • Do not score the transcript character-for-character before converting the script. If you compare a Traditional output against a Simplified reference, you will measure the script difference, not the recognition. On our clip that was the difference between 43.3% and 7.2%.
  • Expect substitutions, not gaps. All six residual errors were replacements. If your transcript reads as complete but a few characters look off, that is the expected failure mode, and proofreading is the right response.
  • Set the language yourself. Our run had zh selected by hand. Most pages ranking for this term advertise automatic language detection, and we did not test detection on this clip at all — so this page has nothing to say about whether auto-detect would have produced the same output.
  • The number is a count, and it is one clip. Six of ninety-seven, one speaker, one accent, one noise-free recording. It is not a rate and it should not be turned into one.

What we did not measure

  • One clip, one speaker, n = 1. No second Mandarin clip exists in our data. The 6-of-97 count is the entire Mandarin evidence base we have.
  • One non-Latin language. Japanese, Korean, Thai, Arabic, Hindi and everything else are untested. The model carries 99 language tokens and 97 of them have never been run. We are not going to estimate what they would do.
  • No auto-detection run. The language was forced, so this page cannot tell you whether automatic detection picks the right language on a Mandarin video.
  • The model is whisper-base. The seven pages above may run entirely different models. Nothing here is a statement about them.
  • One machine, one browser, one OS. Timing and memory differ by hardware; recognition output can differ by model version. This page is about the scoring problem, which is the part that does not change with your laptop.

Where the raw data is

Two files carry everything on this page: chinese-cer-one-clip.csv (the three normalisation levels, with substitutions, deletions and insertions broken out) and chinese-residual-errors.md (the six remaining characters, listed one by one). The scorer that produced them, cer-zh.py, is published next to them, along with the method notes and an explicit list of what the numbers do not cover.

Run it on your own file

The transcriber is on the front page of this site. Drop in a video or audio file you already have, pick the spoken language — including Chinese, which you must select rather than rely on detection for — and read the transcript. Nothing is uploaded; the model runs in the browser tab and the file never leaves your device. Expect the first run to spend 78.4 MiB on the download and about a gigabyte of RAM, and every run after that to start straight away.

Transcribe a file