ClipQuill

Transcribe audio five pages make an accuracy claim, and not one of them says what it was measured on

Transcribing audio means turning the speech in a recording into written words, and the ten tools ranking for the phrase all do that. What they do not do is show the work. 5 of 10 make a claim about accuracy and 2 of 10 put a number on it — the highest is 99%. 5 of 10 name the conditions that move the result: background noise, crosstalk, accents, very quiet recordings. 0 of 10 attaches a number to any of those conditions. 0 of 10 names a test set, a clip count or a word count. 1 of 10 cites any source at all, and the source is the model's training corpus, not a test of the tool. Every count was read from our saved copies of the results page on 2026-10-03.

Ten pages, and not one number with a denominator

Five different headline accuracy numbers from the same eight clips, same tool, same day, whisper-base: the worst single clip reads 40.8 percent, the two heavy-noise clips read 34.8 percent, all eight clips read 20.1 percent, the four clean and light real-speech clips read 9.4 percent, and the best single clip reads 6.1 percent. The small tier is lower on every one of the five: 18.3, 16.9, 9.6, 8.2 and 0.0 percent. The denominator behind all of them is 8 clips, 92.7 seconds of audio and 229 reference words, with two speakers, one noise type (pink noise), six of eight clips read speech, English only and every file mono. The smallest clip is 4.8 seconds and 11 reference words, so a single wrong word moves that clip by 9.1 points; on the 33-word clip one word is 3.0 points.
Same tool, same day, same eight files: the headline number is 40.8% or 9.4% depending only on which clips you decide to count.

We read the first page of results for transcribe audio on 2026-10-03. Ten organic results, 8 distinct domains — two domains hold two results each. We then fetched seventeen pages across the first page of two independent engines: 10 returned readable text, six answered HTTP 403, and one served a JavaScript shell with no content. Every count below is out of those ten.

  • The claims are everywhere; the evidence is nowhere. 5 of 10 readable pages say something about accuracy. Two of them put a number on it. 0 of 10 says how many clips, how many words, or which recordings its figure came from.
  • Not one page even has the vocabulary. Across all seventeen saved copies, the words test set, corpus, benchmark, dataset, validation set, ground truth and sample size appear zero times. A page can call its number industry-leading precision without ever having a noun for what it was measured on.
  • The conditions are named and never quantified. 5 of 10 name the things that move the result — background noise, crosstalk or overlapping speakers, strong accents, very quiet recordings. 0 of 10 attaches a figure to any of them. The closest anyone gets is one page’s phrase dropping a few points on noisy or multi-accent recordings, which is a direction, not a size.
  • Two pages point the opposite way on the same variable. One advertises 99% Accuracy Rate — precision that handles accents, background noise, and multiple speakers with ease. Another warns that heavy accents, background noise, and overlapping speakers may reduce accuracy. Same variable, opposite sign, no measurement on either side.
  • One page cites a source, and it is the wrong kind of number. A page running the model in the browser cites the research paper: trained on 680,000 hours of audio, achieving near-human accuracy. That is the size of the training corpus. It is not a test of that page, of that engine, or of your recording.

Counted on 2026-10-03 from our saved copies of the results page for transcribe audio. Every count is out of the ten pages that returned readable text; the six HTTP 403 responses and the one JavaScript shell are excluded from all of them. Each count was assigned by printing the full context of every match and reading it, not by tallying a regular expression — the method that caught false positives on this site on 2026-09-26, 2026-09-27, 2026-09-28 and 2026-10-01. A mention that appears only in a navigation bar, a footer, or a list of other tools was not counted.

The five claims, quoted, and what is missing from each

Here is every accuracy claim in the ten, in the pages’ own words. We are not naming the sites, because the point is what the ten do and do not say, not which one says it.

  • “Up to 99% accuracy — near-human level precision for reliable, professional results.” — a badge on the first screen. The same page’s FAQ answers How accurate is the AI transcription? with a different and lower number: Our engine averages 95%+ accuracy on clear, single-speaker audio in supported languages, dropping a few points on noisy or multi-accent recordings. That FAQ sentence is the most specific accuracy statement anywhere in the ten: it names the audio it holds for (clear, single-speaker) and the direction the number moves in. It still does not say how many recordings it averages over, or how many words a few points is worth.
  • “99% Accuracy Rate — Industry-leading precision that handles accents, background noise, and multiple speakers with ease.” — a first-screen feature block. The same page later promises that its engine automatically identifies and labels different speakers. One number, three variables it claims to absorb, no measurement of any of them, and no clip count.
  • “Whisper’s research paper describes training on 680,000 hours, achieving near-human accuracy across dozens of languages.” — the only citation of an external source in the ten. It is a fact about how the model was built. Nothing on the page reports what that model did on any audio the page itself chose.
  • “On clear speech with a decent microphone, a Whisper-class large model is near-human — typically a low single-digit word error rate in English and other well-supported languages. Accuracy drops with heavy background noise, crosstalk, strong accents, or very quiet recordings.” — the only page in the ten that uses the term word error rate at all, and the most honest sentence in the set. It names four conditions and gives none of them a size. Low single-digit is as close as any of the ten comes to a denominator, and it is a band, not a number.
  • “AI transcription can make mistakes, especially with overlapping speech, background noise, names, or technical terms. Check important details and quotations against the original recording.” — a warning, which is more than most of the ten offer. It is also a list of the same four variables, again with nothing measured, and it is the page’s only statement about accuracy anywhere.

Read them together and the shape is consistent. Every page knows which variables matter — noise, accents, how many people are talking, how good the microphone is. Not one of them has measured any of those variables, and not one of them will tell you how much audio its own figure came from. That is what makes the numbers incomparable: a 99% and a 95% and a low single-digit word error rate may be three descriptions of the same engine on three different piles of recordings, and nothing on any of the ten pages would let you tell.

All five quotations are copied from our own saved copies of the pages, retrieved on 2026-10-03. We have not tested any of them and make no claim about whether their numbers are right — only that the numbers arrive without the audio they came from.

Our own number, with the denominator attached

We publish an accuracy figure on this site, so the same criticism applies to us until we show the work. Here is the whole of it, and it is not flattering.

  • The test set is 8 clips, 92.7 seconds of audio, 229 reference words. Under a minute and a half of speech is the entire basis of every accuracy number we publish. One clip is synthetic text-to-speech; the other seven are real recordings. Two speakers. One noise type, pink noise, at two levels. 6 of 8 clips are read speech. English only, and every file is mono.
  • The same eight clips produce five different headline numbers. Counting the worst single clip gives 40.8%. Counting the two heavy-noise clips gives 34.8%. Counting all eight gives 20.1%. Counting only the four clean and light real-speech clips gives 9.4%. Counting the single cleanest clip gives 6.1%. Same tool, same day, same files — a factor of 6.7 between the highest and the lowest, decided entirely by which clips you put in the denominator.
  • The percentage is quantised, and the smallest clip is tiny. The smallest clip in the set is 4.8 seconds and 11 reference words. One wrong word moves that clip by 9.1 points. On the 33-word clip, one word is 3.0 points. At this sample size a tenth of a percent is not a measurement, it is an artefact of which clip you happened to include.
  • We publish the bad end too. The smallest tier, on the heaviest-noise clip, produced 113 error words against a 71-word reference — a word error rate of 159.2%, which is possible because insertions and deletions both count. We could have left that clip out. It is in the table, named, with its 29.4 seconds and its 71 words.
  • The figure that turns a percentage into something you can feel. The reference text runs at 148.4 words per minute. At our overall 20.1% that is 46 wrong words against 229 — roughly one error every five words. On the heaviest clip it is 29 against 71, closer to one error every two and a half words. Nobody in the ten gives you this, because nobody in the ten gives you the word count.

Our headline is 9.4%, and it is honest as long as you can see that it is four clips out of eight, chosen to be clean and lightly noisy real speech, and that the full eight-clip figure is 20.1%. A reader who is told only 9.4% has been told a true thing in a way that will mislead them about their own recording. That is the same move the ten pages make, which is why we would rather show you the spread.

Every figure above is read from wer-by-sample.csv and wer-by-condition.csv in the public benchmark repository, measured 2026-09-17 and 2026-09-18. The 92.7 seconds is the sum of the eight audio_seconds values; the 229 reference words and the word-count convention are described in that directory’s README, which also states the corpus limits quoted below. Nothing on this page was re-measured for this article.

What our denominator does not cover

The point of publishing a denominator is that it tells you where the number stops. These are the limits, in the same file the numbers come from.

  • Two speakers, one noise type. Every real recording in the set comes from the same two voices and the same noise, pink noise, at two levels. A four-person meeting with a phone on the table, a car, a cafĂ©, a wind-buffeted field recording — none of that is in the eight clips, and all of it is in the use-case lists the ten pages print.
  • Six of the eight clips are read speech. Read speech is the easy case: the speaker knows what they are going to say, there are no false starts and no one talks over anyone. Spontaneous speech is what a meeting or an interview actually sounds like, and it is two clips out of eight.
  • One language out of ninety-nine. The model carries 99 language tokens. 97 of them have never been run on this site. Everything above is English, and the one non-English measurement we do have is a single 23.088-second Mandarin clip, where the model returns Traditional characters against a Simplified reference and 36 of the 97 characters differ by script alone.
  • The 9.4% headline is a grouping, not a condition. It is four clips taken from two different acoustic conditions, clean and light noise, averaged together because that is the realistic case. The README says this explicitly. A grouping chosen for realism is still a choice, and it is ours.
  • One row is a single run. The longest timing measurement in the repository is a single 1,610-second run with no repeat, so its number has no error bar at all. Where a clip turns out not to be reproducible — the heaviest-noise clip is the one that moves — we say so on the page rather than printing one value and calling it settled.

None of the ten pages has anything equivalent to this section, and that is the real asymmetry. It is not that their numbers are worse than ours. It is that ours can be argued with, and theirs cannot, because there is nothing there to argue with.

The corpus limits, the 97-of-99 figure, the script difference in the Mandarin clip and the single-run caveat are all stated in data/README.md and METHOD.md in the public benchmark repository. The grouping note is in the same README.

The other sense of the word, sitting in the same ten results

There is a second thing the ten pages do not tell you, and it is about the phrase itself. Transcribe audio does not say what the output is. Two of the ten results read the phrase the other way.

  • Two of the ten results are the same desktop program, and its job is music. Its own description: the world’s leading software for helping musicians to work out music from recordings. The detail page that ranks is headed software to help transcribe recorded music, and what it offers is pitch change, named loops, slowing down and speeding up, and support for a foot pedal. That is a tool for learning a part by ear.
  • Both pages do add the other use, further down. The homepage line is and also for speech transcription; the detail page repeats it and points at a section of the program’s own help file for advice. So the text-transcription use is real. It is also the second paragraph of a page whose headline, feature list and screenshots are all about music.
  • None of the ten mentions the split. A reader who types transcribe audio and lands on a music tool gets a page that will not produce a transcript, and the other eight results say nothing about the fact that a fifth of their own results page is a different job.

This is worth stating because it is the kind of thing a results page never explains about itself. The word transcribe carries two jobs — writing down speech, and working out what notes were played — and on this phrase they share the first page.

Both pages were read on 2026-10-03 and the phrases above are copied from our saved copies. We are not naming the program, because the point is the ambiguity of the phrase, not the product.

Questions this page answers

How accurate is transcribing audio, and how would I know?

You would know by reading the denominator, and none of the ten readable pages gives you one. Five make a claim about accuracy; two put a number on it, the highest being 99%; 0 of 10 says how many clips or how many words the number came from, and the words test set, corpus, benchmark, dataset and sample size appear nowhere in any of them. A percentage with no denominator is not a measurement. Ours is 8 clips, 92.7 seconds, 229 reference words, and the same eight clips will give you 40.8% or 9.4% depending on which ones you count.

Why do different transcription pages give different accuracy figures for the same job?

Because the figure moves with the audio, and nobody publishes the audio. Our eight clips span 6.1% to 40.8% — a factor of 6.7 on one tool, one day, one set of files. A page reporting 99% and a page reporting 95% may be describing the same engine over two different piles of recordings. Without the clip list, the two numbers are not comparable, and neither is checkable.

Does background noise change the accuracy of transcribing audio?

Yes. 5 of 10 pages name it, along with accents, overlapping speakers and very quiet recordings, and 0 of 10 attaches a number to any of those conditions. Two of the five point in opposite directions on the same variable — one advertises precision that handles accents, background noise, and multiple speakers with ease, another warns those same things may reduce accuracy. In our published data the variable is the whole story: 6.1% on the clean synthetic clip, 34.8% across the two heavy-noise clips.

What audio was your accuracy number measured on?

8 clips, 92.7 seconds of audio, 229 reference words. One synthetic text-to-speech clip and seven real recordings; two speakers; one noise type (pink noise, two levels); 6 of 8 read speech; English only; every file mono. The limits are stated in the same file as the numbers: 97 of 99 language tokens have never been run, the 9.4% headline is a cross-condition grouping of four clips rather than an acoustic condition, and the longest timing row is a single run. Under two minutes of audio is a small denominator, which is exactly why we show it.

Is every result for transcribe audio about turning speech into text?

No. Two of the ten results are the same desktop program, described on its own page as software for helping musicians to work out music from recordings, with a feature list of pitch change, named loops and speed control. Both pages add, further down, that it is also used for speech transcription. So a fifth of the first page for this phrase is the other job, and none of the ten pages mentions the split.

Where the raw data is

The eight clips, the reference words and every percentage on this page come from the same published measurement run as the rest of this site. The reading of the ten results pages is the only new thing here.

Run it on your own file

The transcriber is on the front page of this site. It reads the audio in the tab and writes the words from it, so the file does not leave your machine. Expect the first run to spend 78.4 MiB on the download and about a gigabyte of RAM, then 18 to 20.5 seconds of wall clock for every minute of audio after that. Accuracy is the 20.1% figure with the clips named — and if your recording is clean, read speech, the honest comparison is the 9.4% grouping, not the headline. The caps are 30 minutes of audio and 512 MiB per file, checked before any work starts.

Transcribe a file