ClipQuill

Video to text five questions people actually asked, answered with measured numbers

These five questions were asked by real people in public forums, quoted here as they were written. They are answered below with what we have measured on our own runs — on the audio most people actually have, clean or light-noise real speech, the English default is 7.1% wrong and the optional small tier 8.2% (85 reference words); on deliberately harsh audio it is higher, up to 21.3% on heavy background noise, a 30-minute per-file cap, and 78.4 MiB crossing the network before the first run starts. Where we have no measurement, this page says so instead of filling the gap with a number. That distinction is the point: 0 of the 30 vendor pages we scanned state what their accuracy figure was measured on.

Where the questions came from

Each question below is quoted from a public discussion, with the subreddit, the thread it came from and that thread's comment count, so you can go and read the whole exchange. A comment count is not evidence of a trend: the two threads above 50 comments are the only ones where more than a handful of people took part, and everything else on this page is one person speaking, not a chorus.

  • Collected — 2026-09-23, from an archive of public discussion threads
  • Screened — 138 comments matched the first pass, 44 were actually about turning speech into text, and 94 were false matches on the word transcript meaning a written record
  • Edited — only for length. The wording inside each quote is the original

One of the five answers quotes a comment whose author disclosed that they make a competing product. That is flagged where it appears. People promoting their own tool are still real people, but they arrive with a position, and it should be visible.

The questions

Word error rate on realistic files: no added noise 7.1%, light background noise 7.0%, typical files combined 7.1% on the English default. Deliberately harsh audio is worse: telephone band with echo and hiss 18.2%, heavy background noise 21.3%.
The same model on the same clips: 7.1% wrong on typical clean or light-noise real speech, against 21.3% on deliberately harsh heavy-noise audio. The spread across conditions is wider than the spread between two tools measured on one condition.

Is there a real accuracy difference between local models and cloud-based transcription?

Asked in r/macapps, August 2026, in a thread about a local Mac transcription app (53 comments): “Did you notice a big difference in transcription accuracy and speed between local models and cloud-based solutions?” The same person added that with their current cloud tool, “sometimes the first few words can be completely wrong.”

We have not run a local-versus-cloud comparison, so there is no number here and we are not going to invent one. What we have measured is the cost of the local side, on a 2-core, 3.94 GB machine: 78.4 MiB downloaded before the first run, a 13-second clip finished in 45.2 s cold and 23.5 s warm, and a browser process tree peaking at 821–1128 MB of RAM. On accuracy, the figure we can give you is our own: on the audio most people actually record — clean or light-noise real speech, 85 reference words — the English default is 7.1% wrong and the optional small tier 8.2%; on deliberately harsh clips (telephone band, heavy noise) it reaches 18.2% to 21.3%. On the first-few-words problem specifically: unmeasured here. It is a real and commonly reported behaviour, and we have not put a number on it.

Can you transcribe a 3-hour audio file for free?

Asked in r/audiovisual, June 2026 (21 comments): “Can you transcribe a 3-hour audio for free?”

Not here. The hard cap is 30 minutes of audio and 512 MiB per file, and the file is refused before any work starts rather than failing halfway. The longest file we have actually timed is 26 min 50 s, which took 550.4 s cold and 485.7 s warm on the same 2-core machine. “Free” in the sense that does apply: nothing is charged per minute, because the model runs in your browser tab and the file never leaves your device. A three-hour recording would have to be cut into segments under the cap — and we have not measured whether accuracy changes when you do that, so treat the result of splitting as unknown rather than as equivalent to one continuous run.

If most tools run on Whisper variants and accuracy is close, what should I compare instead?

Answered in r/automation, September 2026, in a thread asking which transcription tool to use (27 comments). The comment argued that at 15–25 hours a month “the plan structure matters more than the model, since most of these tools run on Whisper variants and raw accuracy is close,” and that the comparison should be on speaker labels, on specialist vocabulary, and on export.

The premise is half right, and the half that is wrong is the useful part. Raw accuracy is not close — it is close between two tools on the same clip, and wildly different between two conditions. Across our eight fixed clips the same model went from 6.1% wrong on clean synthetic speech to 21.3% wrong on real speech with heavy background noise, with light noise at 7.0% and a telephone band with echo and hiss at 18.2%. Two tools that look identical on a clean sample can be several times apart on a noisy one. Compare on the condition your audio is actually in, then on the two things that comment got right: how fast you can correct a word the model got wrong, and what it exports.

Why does the transcript come back grammatically wrong, with missing punctuation?

Posted as a complaint in r/PetPeeves, June 2026 (14 comments), under the title “When voice to text is grammatically wrong.”

Punctuation and capitalisation are decisions the model makes during recognition, not formatting applied afterwards, which is why they are wrong in the same ways the words are. They are also the least documented part of the entire process: of the 30 vendor pages we scanned, 2 mention punctuation or capitalisation at all. We have not measured how often our own output gets punctuation wrong, so there is no figure here. What we can point to is how much a scoring convention moves a result: on one unchanged 23.088-second Mandarin clip, the same model output scored 43.3% wrong compared character for character, 7.2% after the script was converted, and 6.2% once digits were normalised as well. Before blaming the grammar, establish which comparison you are making.

Is the real problem accuracy, or that people stop reviewing once a tool is usually right?

Argued in r/artificial, August 2026, in a thread about AI scribes in healthcare (40 comments): “The review step thing is the real issue imo, not the transcription accuracy itself… once a tool is right 95% of the time, people stop actually reading closely and start rubber-stamping.”

Both, and the first one is worse than the marketing implies. Our measured figure on typical files — clean or light-noise real speech, the audio most people actually have — is 7.1% wrong across 85 reference words, about one word in fourteen. On deliberately harsh audio it is worse: 21.3% on heavy background noise, about one word in five. A transcript at that rate does not deserve rubber-stamping; it deserves reading. Which is why every number on this site ships with the raw output and the scoring script: so you can re-score it, disagree, and see the errors listed rather than summarized into a percentage.

What this page is not

  • Not a survey. Five quotes from five threads. The two threads above 50 comments are the only ones with real participation behind them; the rest are individuals.
  • Not a comparison against named competitors. We have not run their tools, so this page makes no claim about how they perform.
  • Not a full answer to any of the five. Two of them we answer only partially, and we say which parts are unmeasured.
  • One model, one machine. whisper-base quantised, on a 2-core 3.94 GB machine, run on 2026-10-03. Different models and faster machines will differ.

Where the raw data is

Run it on your own file

The transcriber is on the front page of this site. Drop in a video or audio file you already have, pick the spoken language, and read the transcript. Nothing is uploaded; the model runs in the browser tab and the file never leaves your device. Expect the first run to spend 78.4 MiB on the download and about a gigabyte of RAM, and every run after that to start straight away.

Transcribe a file