ClipQuill

Video to text 35 of the 42 wrong characters in our Chinese transcript were the same character in a different script

On one 23.088-second Mandarin clip with 97 reference characters, a raw character-for-character comparison counted 42 characters wrong — 43.3%. We went back through all 42 one at a time: 35 were the same word written in a different script, 1 was a number written as a digit instead of spelled out, and 6 were genuine substitutions. The audio was heard correctly 36 times out of 42. We also scanned the 50,258-entry vocabulary of the two models this site ships: 935 CJK ideographs, 360 Hangul syllables, 136 kana, and 0 full-width forms. Of the 30 vendor pages we scanned for this term, 0 mention Simplified or Traditional characters at all.

What was measured, and on what

Two separate measurements, both run today on files we produced earlier. Neither is a rate and neither is averaged over anything.

  • The clip — 23.088 s of Mandarin speech, one speaker, no added noise. Reference: 97 characters, transcribed by hand from the audio. Spoken language forced to zh, not auto-detected.
  • The run — onnx-community/whisper-base, quantised ONNX, served from this site. 2-core AMD EPYC 9754, 3.94 GB RAM, real Chrome window driven over the Chrome DevTools Protocol, 2026-09-18. Single run, never repeated, so no jitter range is reported.
  • The re-scoring — run today, 2026-09-24, on the saved raw output of that run. Deterministic: same input, same answer every time, no timing involved.
  • The vocabulary scan — run today over vocab.json for both whisper-base and whisper-tiny, the two models this site actually serves. The two vocabularies are identical: 50,258 entries each.
  • Numbers and downloads — not applicable to either measurement. No model was downloaded or run for this page; it is arithmetic over files already on disk.

The scoring script and its full output are published next to this page: scan-cjk.py and scan-out.json. The script carries the character mapping inline, every entry of it checkable by hand.

The 42 mismatches, sorted by what they actually are

Stacked bar of the 42 mismatched characters: 35 script variants, 1 numeral form, 6 real substitutions.
42 mismatches out of 97 characters. Thirty-five of them are the same word in a different script.

The output was 97 characters long against a 97-character reference, with zero deletions and zero insertions, so every position lines up and each mismatch can be classified on its own.

  • 42 — characters counted wrong in a raw comparison, 43.3% of 97
  • 35 — the same word, written in a different script. The reference character and the output character are the two standard forms of one word
  • 1 — a number the reference spells out as a character and the output writes as the ASCII digit 3
  • 6 — genuine substitutions, where the model produced a different word
  • 27 — how many distinct script pairs those 35 mismatches come from; one pair accounts for 5 of them and another for 3
  • 0 and 0 — deletions and insertions, at every level of normalisation

Applying the conversion ourselves, character by character, reproduces the three figures already published for this clip exactly: 7.2% after script conversion (7 of 97) and 6.2% after numerals are normalised as well (6 of 97). That match is the check on our mapping — it was produced independently of the library that produced the published figures, and it lands on the same two numbers.

One reconciliation, published rather than quietly settled. Our method notes say the raw comparison counts 36 characters differing by script. Counting character by character today gives 35 script pairs plus 1 numeral form — 36 under the wider definition, 35 under the narrow one. Both are right about what they counted; we are not going to pick one and forget the other.

Why it writes Traditional: what is actually in the vocabulary

The output script is not a preference the model expresses. It is a consequence of which characters it has. These counts are over the full 50,258-entry vocabulary, identical for whisper-base and whisper-tiny.

  • 935 — CJK unified ideographs that exist as a token of their own. The same 935 is the total that appear anywhere in the inventory, so every ideograph it has, it has whole
  • 360 — Hangul syllables as tokens of their own; 566 distinct syllables appear somewhere inside longer tokens
  • 136 — kana: 69 hiragana and 67 katakana
  • 12 — CJK punctuation marks
  • 0 — full-width forms, anywhere in the entire inventory

Now the 27 script pairs from our clip, checked against that inventory:

  • 12 of 27 — both the Simplified and the Traditional character have a token of their own. The model can go either way, and mixing is available to it
  • 12 of 27 — only the Traditional character has a token. For these, Traditional is not a choice, it is the only form in the inventory
  • 3 of 27 — neither form is in the inventory, and the model wrote all three anyway

Measured the other way round, on the two texts themselves: of the 74 distinct characters in our Simplified reference, 53 exist as tokens of their own. Of the 74 distinct characters the model produced, 66 do. The inventory is better stocked on the side the output lands on.

The three characters it produced with no token to draw on are the interesting ones, and they are named by codepoint in the scan output: U+6E2C, U+8A0E and U+9435. They appear nowhere in the vocabulary, in any token of any length, yet the transcript contains them. Byte-level encoding lets any character be assembled from byte pieces, so an absent character is not an impossible one — it is one the model has to spell.

Japanese: it has the old forms, not the new ones

We sampled characters that exist only in Japanese, and pairs where Japanese reformed a character and Chinese did not. Every result below is a lookup against the same inventory. We have never run a Japanese clip, so none of this is a performance claim.

  • 0 of 7 — Japanese-only characters present. U+50CD, U+99C5, U+7551, U+8FBB, U+8FBC, U+585A and U+5CE0 are all absent from the inventory entirely
  • U+7D4C absent, U+7D93 present — the post-reform form is missing while the pre-reform one is there
  • U+8EE2 absent, U+8F49 present — same pattern again
  • U+5E83 and U+5EE3 both absent, U+52B4 and U+52DE both absent — neither form of these two is in the inventory
  • U+5186 present, U+5713 absent — here the reformed form is the one it has
  • U+5909 and U+8B8A both present, U+5B9F and U+5BE6 both present — two pairs where it can go either way

The pattern is not "Japanese is unsupported". It is worse than that and more specific: the inventory was built from text that did not consistently use the current Japanese forms, so where the reformed character is missing, the pre-reform character is often sitting there in its place.

Korean, and how numbers are written

  • Korean has no script-variant problem. It is written in an alphabet. Its problem is coverage: 360 syllables as tokens of their own against 11,172 in the standard syllable set. Of thirteen common syllables we sampled, 12 have a token of their own and 1 exists only inside longer tokens
  • Chinese numerals are well covered, with three gaps. U+4E09, U+5341, U+767E, U+5343, U+4E07, U+842C, U+5104, U+4E24 and U+5169 are all present. U+3007, U+96F6 and U+4EBF are absent
  • All ten ASCII digits are present; no full-width digit is. This is why the model reaches for 3 where our reference spells the number out, and why that mismatch survives script conversion
  • We have never run a Korean clip either. 97 of the model's 99 language tokens have never been run by us

The questions

Is a Traditional character where the reference has a Simplified one an error?

Not a recognition error. On this clip, 35 of the 42 mismatches were the same word in a different script: the reference had one standard form, the output had the other. Counting them as errors measures the alphabet, not the hearing. The 6 that remain after conversion are the ones worth arguing about, and 2 of those 6 are the same word counted twice, which is why six errors sit at five places in the text.

Why does the transcript mix Simplified and Traditional characters?

A complaint posted in r/anime, in a thread about a streaming service launching in Taiwan: “Taiwanese people have complained about poor subtitle quality like bad machine translations or mixed simplified and traditional characters in the subtitles.” One person posting; treat it as one person's report, not a measurement.

Because the model decides per character, not per document. 12 of the 27 script pairs in our clip have both forms available as tokens of their own, so the output is free to choose differently at each position and there is nothing enforcing consistency. In 12 further pairs only the Traditional form exists, so those positions are not choices at all. We have not measured how often a single output actually mixes forms beyond this one clip, so we are not putting a frequency on it.

Should I ask for Traditional or Simplified output?

Asked in r/ChineseLanguage (30 comments), as a thread title: “Should I start learning with Traditional or Simplified characters?” That thread is about learning to read, not about transcripts, and the answer in it was about which materials are easier to find — but it is the same question people ask of a transcript, and the answer here is different.

You cannot ask this model. There is one Chinese output and no script switch on it. What you control is the side you score against: of the 74 distinct characters in our Simplified reference, 53 are tokens of their own; of the 74 distinct characters the model produced, 66 are. Convert one side to the other before scoring, convert it with a tool you can name, and write down which direction you went — because that direction is most of your number.

Does the script of the transcript follow the language I select?

Described in r/PLAUDAI (15 comments) as a workaround: “open the recording, go to the Language setting, and re-generate the transcript with English selected instead of Chinese.”

Yes, and on this clip it is the biggest effect we have measured. Same file, same machine, same model. Language forced to Chinese: 97 characters of Chinese script, 62.4 s reported by the page. Language forced to English: 243 characters of Latin script, 38.2 s, 97 substitutions and 146 insertions against the same 97-character reference, a character error rate of 250.5%. The rate is over 100% because insertions count and the reference is short. Nothing was dropped — the model wrote a different script in more than twice as many characters.

Does the model write numbers as digits or as words?

Digits. Our reference spelled a number out as U+4E09; the model wrote 3. That is 1 of the 42 mismatches and the only one that script conversion does not touch, which is why the published figures move from 7.2% to 6.2% when numerals are normalised on both sides. Every ASCII digit is in the inventory; no full-width digit is. If your reference spells numbers out, normalise both sides or you will count the notation as an error.

Can the model write a character that is not in its vocabulary?

Yes. 3 of the characters it produced on our clip appear nowhere in the 50,258-entry inventory — not as a token of their own, not inside any longer token. It wrote them anyway. Byte-level encoding means the alphabet is not closed: any character can be assembled from byte pieces. The practical consequence is not impossibility but cost, and we have not measured what that cost is.

Does the same problem apply to Japanese?

On the evidence of the inventory, the Japanese case is worse than the Chinese one. All 7 Japanese-only characters we sampled are absent, and in 2 pairs the post-reform form is missing while the pre-reform form is present. Chinese at least has both forms for 12 of its 27 pairs. This is a statement about what the model can write. We have never run a Japanese clip, and we are not going to estimate what one would score.

What about Korean?

No script variants to argue about, but a coverage gap of a different shape: 360 syllables with a token of their own and 566 distinct syllables somewhere in the inventory, against 11,172 in the standard set. Of thirteen common syllables we sampled, 12 have their own token and 1 only appears inside longer tokens. We have never run a Korean clip.

Can the model write full-width digits or full-width punctuation?

Not from its inventory. Across all 50,258 entries in both models, 0 full-width forms appear anywhere: no full-width digits, no full-width Latin letters, no full-width comma or full stop. 12 CJK punctuation marks do appear. Byte-level encoding still makes the characters reachable, but nothing in the model is built to produce them, so do not plan a pipeline that expects CJK-width punctuation.

Can I fix the script afterwards instead?

Yes, and on our clip that is most of the fix: 43.3% compared character for character, 7.2% after conversion, 6.2% after numerals as well. Same audio, same output, three numbers. Converting afterwards changes the measurement and not the transcript, so it is the right move for scoring and the wrong move if you are trying to find out what the model did.

Is the gap between marketed language support and real accuracy real?

Posted in r/SaaS by someone who disclosed they build a competing product — their position, not ours: “The accuracy gap between marketed language support and real-world performance is massive… the claimed 99 language support is technically true but practically misleading.”

Our own numbers point the same way, on one clip and eight English ones. 97 of 99 language tokens have never been run by us. The one non-Latin language we did run scored 6 wrong characters out of 97 after normalisation, against a 13.5% word error rate across all eight English clips (7.1% on typical files). That is two data points on two different metrics, and it is not enough to call a trend — which is exactly the discipline the quote above is asking vendors to drop.

What the pages ranking for this term say about it

We scanned 30 pages — ten each from three sites, taken from their own sitemaps — on 2026-09-23, before writing any of this.

  • 0 of 30 mention Simplified or Traditional characters
  • 29 of 30 use the word accuracy; 28 of 30 print a percentage
  • 0 of 30 state what their figure was measured on, and 0 of 30 publish raw data
  • 14 of 30 mention Chinese, and every one of those mentions is a single word inside a language list

That scan was of the pages ranking for this term on the day it was run. It is not a claim about how any of those tools perform — we have not run them.

What this page does not cover

  • One clip, one speaker, n = 1. Every character figure here comes from 97 characters of one recording. It is a count, not a rate
  • No Japanese or Korean audio has ever been run. Those sections are about the output inventory, not about accuracy
  • The mapping is ours. The script conversion is a character-by-character table written into the scan script. It reproduces the two published figures exactly, which is the check we have; it is not a standard library
  • Vocabulary is not capability. A character absent from the inventory can still be produced, as three of them were. Absence says what the model has to reach for, not what it cannot do
  • One model family. Both models scanned are Whisper quantised ONNX and share one vocabulary. Nothing here describes any other model
  • No second run. The clip was run once on 2026-09-18 and never repeated, so there is no reproducibility figure for the transcript itself

Where the raw data is

Measure it on your own file

The transcriber is on the front page of this site. Drop in a video or audio file you already have, pick the spoken language — including Chinese, which you have to select rather than rely on detection for — and read the transcript. Nothing is uploaded; the model runs in the browser tab and the file never leaves your device. Expect the first run to spend 78.4 MiB on the download and about a gigabyte of RAM, and every run after that to start straight away. If you are scoring the result, convert the script on one side first and write down which way you went.

Transcribe a file