Where the questions came from
Each question below is quoted from a public discussion, with the subreddit, the thread it came from and that thread's comment count, so you can go and read the whole exchange. A comment count is not evidence of a trend: the two threads above 50 comments are the only ones where more than a handful of people took part, and everything else on this page is one person speaking, not a chorus.
- Collected — 2026-09-23, from an archive of public discussion threads
- Screened — 138 comments matched the first pass, 44 were actually about turning speech into text, and 94 were false matches on the word transcript meaning a written record
- Edited — only for length. The wording inside each quote is the original
One of the five answers quotes a comment whose author disclosed that they make a competing product. That is flagged where it appears. People promoting their own tool are still real people, but they arrive with a position, and it should be visible.
The questions
Is there a real accuracy difference between local models and cloud-based transcription?
Asked in r/macapps, August 2026, in a thread about a local Mac transcription app (53 comments): “Did you notice a big difference in transcription accuracy and speed between local models and cloud-based solutions?” The same person added that with their current cloud tool, “sometimes the first few words can be completely wrong.”
We have not run a local-versus-cloud comparison, so there is no number here and we are not going to invent one. What we have measured is the cost of the local side, on a 2-core, 3.94 GB machine: 78.4 MiB downloaded before the first run, a 13-second clip finished in 45.2 s cold and 23.5 s warm, and a browser process tree peaking at 821–1128 MB of RAM. On accuracy, the figure we can give you is our own: on the audio most people actually record — clean or light-noise real speech, 85 reference words — the English default is 7.1% wrong and the optional small tier 8.2%; on deliberately harsh clips (telephone band, heavy noise) it reaches 18.2% to 21.3%. On the first-few-words problem specifically: unmeasured here. It is a real and commonly reported behaviour, and we have not put a number on it.
Can you transcribe a 3-hour audio file for free?
Asked in r/audiovisual, June 2026 (21 comments): “Can you transcribe a 3-hour audio for free?”
Not here. The hard cap is 30 minutes of audio and 512 MiB per file, and the file is refused before any work starts rather than failing halfway. The longest file we have actually timed is 26 min 50 s, which took 550.4 s cold and 485.7 s warm on the same 2-core machine. “Free” in the sense that does apply: nothing is charged per minute, because the model runs in your browser tab and the file never leaves your device. A three-hour recording would have to be cut into segments under the cap — and we have not measured whether accuracy changes when you do that, so treat the result of splitting as unknown rather than as equivalent to one continuous run.
If most tools run on Whisper variants and accuracy is close, what should I compare instead?
Answered in r/automation, September 2026, in a thread asking which transcription tool to use (27 comments). The comment argued that at 15–25 hours a month “the plan structure matters more than the model, since most of these tools run on Whisper variants and raw accuracy is close,” and that the comparison should be on speaker labels, on specialist vocabulary, and on export.
The premise is half right, and the half that is wrong is the useful part. Raw accuracy is not close — it is close between two tools on the same clip, and wildly different between two conditions. Across our eight fixed clips the same model went from 6.1% wrong on clean synthetic speech to 21.3% wrong on real speech with heavy background noise, with light noise at 7.0% and a telephone band with echo and hiss at 18.2%. Two tools that look identical on a clean sample can be several times apart on a noisy one. Compare on the condition your audio is actually in, then on the two things that comment got right: how fast you can correct a word the model got wrong, and what it exports.
Why does the transcript come back grammatically wrong, with missing punctuation?
Posted as a complaint in r/PetPeeves, June 2026 (14 comments), under the title “When voice to text is grammatically wrong.”
Punctuation and capitalisation are decisions the model makes during recognition, not formatting applied afterwards, which is why they are wrong in the same ways the words are. They are also the least documented part of the entire process: of the 30 vendor pages we scanned, 2 mention punctuation or capitalisation at all. We have not measured how often our own output gets punctuation wrong, so there is no figure here. What we can point to is how much a scoring convention moves a result: on one unchanged 23.088-second Mandarin clip, the same model output scored 43.3% wrong compared character for character, 7.2% after the script was converted, and 6.2% once digits were normalised as well. Before blaming the grammar, establish which comparison you are making.
Is the real problem accuracy, or that people stop reviewing once a tool is usually right?
Argued in r/artificial, August 2026, in a thread about AI scribes in healthcare (40 comments): “The review step thing is the real issue imo, not the transcription accuracy itself… once a tool is right 95% of the time, people stop actually reading closely and start rubber-stamping.”
Both, and the first one is worse than the marketing implies. Our measured figure on typical files — clean or light-noise real speech, the audio most people actually have — is 7.1% wrong across 85 reference words, about one word in fourteen. On deliberately harsh audio it is worse: 21.3% on heavy background noise, about one word in five. A transcript at that rate does not deserve rubber-stamping; it deserves reading. Which is why every number on this site ships with the raw output and the scoring script: so you can re-score it, disagree, and see the errors listed rather than summarized into a percentage.
What this page is not
- Not a survey. Five quotes from five threads. The two threads above 50 comments are the only ones with real participation behind them; the rest are individuals.
- Not a comparison against named competitors. We have not run their tools, so this page makes no claim about how they perform.
- Not a full answer to any of the five. Two of them we answer only partially, and we say which parts are unmeasured.
- One model, one machine. whisper-base quantised, on a 2-core 3.94 GB machine, run on 2026-10-03. Different models and faster machines will differ.
Where the raw data is
- Raw data and the measurement scripts — github.com/supersophia8888-cloud/clipquill-asr-benchmark
- Archived copy with a DOI — 10.5281/zenodo.22826968, the concept DOI: that link always resolves to the newest version of this dataset, all versions.
- The eight clips, bucket by bucket — the measured word error rate, including the clips that went badly
- Why one language's script changes the score — the three numbers from one Mandarin clip
- What the tool costs the machine it runs on — measured download size, timings and memory
- What vendor pages leave out — 30 pages scanned, 0 state their test conditions
Run it on your own file
The transcriber is on the front page of this site. Drop in a video or audio file you already have, pick the spoken language, and read the transcript. Nothing is uploaded; the model runs in the browser tab and the file never leaves your device. Expect the first run to spend 78.4 MiB on the download and about a gigabyte of RAM, and every run after that to start straight away.