ClipQuill

YouTube video transcript nine pages sell the text and none of them shows it

A YouTube video transcript is the spoken text of a video, written out so it can be read instead of watched, and every page ranking for the phrase will make you one. Not one of the nine we could read shows you a single line of what that text looks like. 0 of 9 print a sample. 0 of 9 mention punctuation or capitalisation anywhere on the page. 0 of 9 mention paragraph or line breaks. 2 of 9 mention speaker labels or speaker recognition at all, both inside a control or a feature list. 4 of 9 describe their output with the words clean or readable and add no other property to it. Every count below was read from our saved copies of the results page on 2026-10-01.

Nine pages, one word, and no sample of the text

What the nine readable pages ranking for youtube video transcript show: 0 of 9 print a sample of the transcript, 0 of 9 mention punctuation or capitalisation, 0 of 9 mention paragraph or line breaks, 2 of 9 mention speaker labels, and 4 of 9 describe the output only as clean or readable. Below, three long recordings and 7,385 reference words counted two ways: words dropped 356 for tiny, 558 for base, 379 for small and 111 for the cloud reference; spots to fix by hand 601, 404, 268 and 592.
Nine pages describe a document none of them shows, and the property they reach for instead is the one that cannot be checked.

We read all ten results for youtube video transcript on 2026-10-01. Ten results, 9 distinct domains — one domain holds two of them — and 9 of 10 returned readable HTML. The tenth answered with HTTP 429, so it is excluded from every content count below and every count is out of nine.

  • The pages are about producing text, and none of them shows text. 8 of 10 page titles contain the word Generator or Extractor, and 8 of 9 readable pages put a box you paste a link into above everything else. 0 of 9 prints a sample of the transcript it produces. The nearest thing to an example is a section headed Examples on one page, which lists four recordings with their durations — 29:04, 01:41:39, 04:15, 01:30 — and no text from any of them.
  • Punctuation is never mentioned. 0 of 9 pages contain the words punctuation or capitalisation in any form, on any part of the page, including the FAQ blocks that 9 of 10 of them ship as structured data. A transcript that arrives as one unpunctuated run and one that arrives properly punctuated are sold with identical copy.
  • Neither is the paragraph. 0 of 9 pages mention paragraph breaks, line breaks, or where one speaker’s turn ends and the next begins. One page comes close by saying its text has no timestamps clutter, no auto-generated mess — just clean text, which tells you what it removes and not what is left.
  • Speaker labels: two pages, and neither says how a turn is marked. 2 of 9 mention speaker labels or speaker recognition anywhere, and both mentions sit inside a list of tool features or a control — one beside 99.9% accuracy, 200+ languages and unlimited minutes, the other as an option in a panel. Neither of them says what the text does when two people talk over each other, or what separates one speaker’s turn from the next. For a document you read, that is basic structure.
  • Four pages use the same two words and nothing else. 4 of 9 describe their own output as clean, readable, or clean, readable text, and add no other property to it: no punctuation, no paragraph, no speaker, no word count, no sample. Those two words are the entire specification of the document you are being sold.

Counted on 2026-10-01 from our saved copy of the organic first page for youtube video transcript, fetched through two independent engines that returned the same ten results. Every count is out of nine readable pages. Each count was assigned by printing the full context of every match and reading it, not by tallying a regular expression — the method that caught a false positive on this site on 2026-09-26, when a competitor’s own product name was counted as a claim about captions. A word appearing only in a navigation sidebar, a list of other tools, or a blog-post title in a footer was not counted.

The one page that does define the three words, and what it still does not show

There is one genuinely useful paragraph on the results page, and it is on one page out of nine. Asked what the difference is between a transcript, subtitles and closed captions, that page answers:

  • “A transcript is the full spoken text of the video, as flowing text or timestamped lines. Subtitles are that text split into short, timed on-screen cues (SRT or VTT) meant to be read while watching. Closed captions (CC) are subtitles that also note non-speech audio like [music] or [applause].” — a page answering its own FAQ, 2026-10-01

That is the most concrete structural statement anywhere in the ten. It tells you that a transcript is not a fixed shape — flowing text or timestamped lines — and it is the only place on the results page where the non-speech events show up at all: a caption track can carry [music] and [applause], and whether the text you receive carries them is a decision no other page raises.

What it still does not do is show you the text. The same page offers SRT, VTT, TXT and PDF exports and describes its lines as carrying start and end times in milliseconds, and nowhere on it is a single sentence of output. So the most informative page in the ten still asks you to take the shape of the document on trust.

Why this matters for the word transcript in particular: a reader searching the noun is not shopping for a feature list, they are trying to find out what they will be reading. Eight of the nine pages answer a different question — how to get the text — and the ninth answers what the text is, in prose, without an example.

The quotation is copied from our own saved copy of that page, retrieved on 2026-10-01. We are not naming the site, because the point is what the ten do and do not say, not which one says it.

No accuracy figure in this market can tell you whether the text is readable

Here is the part that costs a reader something. Three of the nine readable pages print an accuracy figure: 99.9%, 98.7% and 95%. Not one of the nine mentions punctuation anywhere, and not one of the nine publishes a word error rate or says what audio the figure was measured on.

The reason the two facts are connected is the way these numbers are computed. This site publishes its own rule, and it is the ordinary one. Before anything is compared, the transcript and the reference are lower-cased, and every non-letter and non-digit character — punctuation included — is replaced by a space; runs of spaces are collapsed and the ends trimmed. Only then does the word-level edit distance run.

Under that rule a transcript with no full stops, no commas and no capitals and a transcript with perfect punctuation score identically. The punctuation is deleted before the counting starts. So an accuracy figure, whichever page prints it, is a statement about words that survived the deletion, and it is silent on the one property that decides whether a reader can get through the text.

This is not a criticism we exempt ourselves from. Our own 20.1% word error rate is computed the same way, over the same deletion rule, and it says nothing about punctuation either. What we can do about it is show the text: the raw model output behind every number on this site is in the public repository, so the scoring can be run again under different rules — including a rule that keeps punctuation and charges for losing it. None of the nine pages offers that, and none of them prints an example.

The normalisation rule quoted above is the one published on this site’s benchmark page and used for every error rate on this site. The three competitor figures are copied from our saved copies of the results page, 2026-10-01. We are not asserting that those figures are wrong; we are pointing out that the standard method makes them unable to speak to punctuation, ours included.

Readability is measurable, just not as one number

Four pages call their output clean or readable and stop there. We can do better than that, not because we have a better adjective but because we counted the same output two ways and published both columns.

The two columns measure two different things a reader experiences:

  • Words dropped — words the reference had that the transcript never produced. A reader cannot see these. The word is simply absent, and nothing on the page marks the hole.
  • Spots to fix by hand — runs of wrong words collapsed into one edit. A reader sees these: you stop, back up, and decide what was meant. A twenty-two-word invented sentence counts as one spot.

On the eight short clips — 229 reference words — the four tiers land close together and neither column separates them: tiny 15 dropped and 31 spots, base 13 and 23, small 7 and 18, cloud reference 7 and 16.

On three longer recordings — about 5.6, 13.1 and 27.9 minutes, 7,385 reference words — the two columns stop agreeing and rank the tiers in opposite directions. tiny 356 dropped and 601 spots, base 558 and 404, small 379 and 268, cloud reference 111 and 592.

Read the columns separately and the result is uncomfortable: small needs the fewest hand-fixes, 268 against the cloud reference at 592, and the cloud reference drops the fewest words, 111 against small at 379. Neither wins both. That is why this site says the tiers are about the same and does not say one is more accurate.

The same two columns are also why clean is not a specification. A page that drops whole sentences silently can score a low spot count while hiding the most damaging kind of error, and a page that never drops a word can still make you stop every few seconds. Both are readable text by the copy on four of the nine pages.

All eight figures are read from the tables on the benchmark page, which is generated from the same measurement run as the rest of the data on this site. The four tiers are tiny and base (both shipped by this page), small, and a commercial cloud speech API run as a reference point, not a recommendation. The 7,385 reference words are the sum of the three long files’ reference transcripts.

The same text has two word counts, and we publish both

There is one more reason a page can get away with saying clean, readable and nothing else: the properties it is hiding do not have single values either.

The clearest case is the simplest question you can ask about a document — how many words is it. In our own measurement files the scorer splits possessive apostrophes into separate tokens, so a plain whitespace count of the same reference text comes out lower on three of the eight clips:

  • C2-real-clean — 10 words counted by whitespace, 11 by the scorer. One apostrophe.
  • D2-real-noise-light — 24 against 25. One apostrophe.
  • E1-real-noise-heavy — 68 against 71. Three apostrophes.

Both counts are defensible; they disagree about what a word is. The 229 reference words published across the eight clips are the scorer counts, and the file that says so is in the repository with the rest of the data.

None of the nine ranking pages gives any word count for its output at all, and 0 of 9 say how the length of a transcript relates to the length of the recording. So the reader is left with an adjective. An adjective cannot be wrong, which is exactly the problem with it.

The two counts and the three clips come from the conventions section of data/README.md in the public benchmark repository, which also states the total of 229 reference words. Nothing was recounted for this page.

The other half: what this page does not tell you

  • We did not run any of the nine. Every count above is about what the pages say, not about what their tools return. A page that prints no sample may still produce excellent text. We are reporting that it does not show you, which is a different and smaller claim.
  • The punctuation claim is about copy, not output. We counted mentions of punctuation on the pages. We did not obtain the output of any of the nine, so we cannot say their text lacks punctuation; only one of them would have had to write the sentence for the count to change.
  • We cannot do speaker labels either. Two of nine pages name them and we do not do them at all. That is a real hole in our own output, not a feature, and it is the same hole the other seven pages leave.
  • Our own numbers are narrow. 8 clips, two speakers, one noise type (pink noise), 6 of 8 read speech, English only; the long-file batch is three recordings. The two-column split is the more useful of the two batches for this question and it is also the smaller one.
  • One thing the top ten does better than us. A link-based tool can take a video you cannot download, including one you did not upload. This site cannot: the file has to be on your device. That is a limitation of the route we chose, and it is why none of these pages’ paste-a-link convenience exists here.
  • The pages were read once, on one day. The ten results were saved on 2026-10-01. These tools rewrite their copy often, and a page that prints no sample today may print one next month.

Questions this page answers

What is a YouTube video transcript, and does any page ranking for it show you one?

A YouTube video transcript is the spoken text of a video, written out so it can be read instead of watched. Every page ranking for the phrase will produce one for you, and not one of the nine we could read shows you a sample of the text it produces. 0 of 9 print an example, 0 of 9 mention punctuation or capitalisation, 0 of 9 mention paragraphs, and 2 of 9 mention speaker labels or speaker recognition at all. The closest thing to an example is a section headed Examples that lists four recordings with their durations, not text.

Why do YouTube transcripts come out without punctuation or paragraphs?

The results page never answers this, so what follows is the mechanism rather than a measurement. A caption track is written for on-screen display, in short timed cues, and the one page that defines the words says so directly: subtitles are the transcript text split into short, timed on-screen cues. Text built from that track inherits the cue breaks. Text built by running recognition on the audio inherits whatever the recogniser decides, and how a recogniser decides where a full stop goes is documented on none of the nine pages. We are not putting a number on this, because our own scoring deletes punctuation before it counts anything.

Can I trust a 99% accuracy claim on a YouTube video transcript?

Not as a statement about readability, and not as a statement about punctuation. Three of the nine readable pages print a figure and the highest is 99.9%. The standard way to compute one lower-cases the text and replaces every non-letter and non-digit character, punctuation included, with a space before any comparison happens — the rule this site publishes for its own numbers. Under that rule an unpunctuated transcript and a perfectly punctuated one score identically. So no accuracy percentage in this market, ours included, can tell you whether the text reads, and 0 of 9 pages mention punctuation anywhere.

How many words is a YouTube video transcript?

The same transcript has more than one defensible word count, and that is documented rather than estimated. In our own files the scorer splits possessive apostrophes into separate tokens, so a whitespace count comes out lower on three of eight clips: 10 against 11 on C2-real-clean, 24 against 25 on D2-real-noise-light, and 68 against 71 on E1-real-noise-heavy. The 229 reference words published across the eight clips are the scorer counts. None of the nine ranking pages gives any word count at all.

What is the difference between a YouTube transcript and YouTube’s captions?

One page out of nine answers this and answers it well: a transcript is the full spoken text of the video, as flowing text or timestamped lines; subtitles are that text split into short, timed on-screen cues; closed captions are subtitles that also note non-speech audio, with [music] and [applause] given as the examples. That is the most concrete structural statement anywhere on the results page. The other eight do not draw the distinction, so a reader who wanted flowing text and got cue-by-cue lines — or the reverse — finds out afterwards.

Where the raw data is

The error rates, the two columns and the word-count conventions above come from the same measurement run as everything else published here. The reading of the ten results pages is the only new thing on this page.

Run it on your own file

The transcriber is on the front page of this site. It writes the words from the audio, so the text it produces is its own rather than a caption track’s, and it can be re-scored because the output is published. Expect the first run to spend 78.4 MiB on the download and about a gigabyte of RAM, then 18 to 20.5 seconds of wall clock for every minute of audio after that. Accuracy is the 20.1% figure with the clips named, not a single headline number, and the two columns above are the honest description of what that means for a reader. The caps are 30 minutes of audio and 512 MiB per file, checked before any work starts.

Transcribe a file