ClipQuill

Transcribe Video to Text — 30 Languages, Free, Private, No Signup

The file is processed in your browser. Zero bytes are uploaded — your video never leaves your device, and nothing is kept. Free, no signup, no account, no daily cap. It transcribes 30 languages — English, 中文, Español, 日本語, 한국어 and 25 more; pick the one that is spoken and the text comes back in that language. Drop in any video or audio file you already have — a phone recording, a Zoom or Teams capture, a Reel, or a file saved from YouTube, TikTok or Facebook — and read the text, copy it, or export it as TXT, SRT or VTT. It runs in the tab you already have open, and it works on a phone as well as a desktop. The measured figures, the method and the raw data are further down this page (see accuracy and where the numbers come from).

Drop a video or audio file here

Nothing to install — no download, no plug-in, no app, no signup. It runs in the tab you already have open.

English runs on Moonshine base — the more accurate of the two on everyday speech, and the smaller download. The other 29 languages run on whisper. Choose the larger tier for heavy background noise, or when you need frame-accurate SRT/VTT subtitle timing. The maximum-accuracy tier (distil-whisper for English, whisper-large-v3-turbo for other languages) roughly halves the error rate again — but it downloads 550–730 MB and is several times slower, so it suits short recordings and one-off jobs, not a long file on a phone.

The model shipped here cannot work out the language on its own — it has to be told. Pick the language that is spoken in the file; if English is picked for a Chinese clip, the transcript comes back in English. The choice is remembered on this device.

This page works on a file that is already on your device. It does not fetch a video from a link.

MP4 · MOV · WEBM · M4A · MP3 · WAV · AAC · OGG · FLAC · OPUS · 3GP · M4B. That list is what we have tested; a file outside it may still work, but we do not promise it. You can also paste a file straight in with Ctrl+V.

Three steps, and nothing else

  1. Drop the file in — a video or an audio file you already have. Nothing is uploaded; the file stays in this tab.
  2. Tell it which language is spoken — the model has to be told, it does not guess for you. Pick the wrong one and the text comes out wrong.
  3. Read it, copy it, or export it — TXT, SRT or VTT. There is nothing to unlock afterwards.

The first file you run also downloads the model, so it takes longer than the ones after it. Files up to 30 minutes are accepted; past that this page stops rather than hang.

What this page supports

This is a video to text converter that runs entirely on your own device. The recognition model is downloaded once, then all work happens locally. Two different things happen here, and it is worth keeping them apart: your file is never uploaded — not one byte of it leaves your device, it is never stored, and nothing is sent to this site; the model is downloaded to your browser once, about 76 MiB of model files, and your browser caches it after that.

  • Input formats — MP4, MOV, WEBM, M4A, MP3, WAV, AAC, OGG, FLAC, OPUS, 3GP, M4B. That list is what we have tested; a file outside it may still work, but we do not promise it.
  • Output formats — TXT, SRT, VTT, plus one-click copy.
  • Limits — 30 minutes of audio and 512 MiB (536,870,912 bytes) per file. These are the figures for whisper-base, the multilingual tier; English runs on Moonshine base, and the tool also offers a larger, more robust tier (whisper-small) you can switch to before running a file. A file past either one is turned away before it starts rather than run until the browser runs out of memory and the tab dies. That is the trade for doing the work on your own device instead of on a server, and we would rather state it than imply “any length” — this page does not.
  • Privacy — the file is processed in your browser. Zero bytes are uploaded.
  • Accounts — none. No signup, no login, no daily quota, no credits, no paywall.

Timed on this machine — 2026-09-18, Windows, Chrome 153, whisper-base (the multilingual tier) quantized, end-to-end in a real browser tab, on the same build served here (cross-origin-isolated, same code path), on a machine with 2 logical CPUs — an AMD EPYC 9754 host with two CPUs allocated to it — and 3.9 GiB of RAM. A first visit transfers about 78.4 MiB (82.2 MB) with brotli in transit, and about 97.5 MiB (102.3 MB) once decompressed: the 76.0 MiB model files plus the 21.5 MiB runtime that runs them. Both numbers were counted file by file on clipquill.com on 2026-09-19 — every file a first visit pulls fetched twice, once with brotli and once uncompressed, and the bytes summed. The weights are already quantised and barely compress, so 73.3324 MiB of that 78.4 MiB is model weights crossing the wire at full size. Your browser caches all of it, so from the second file on nothing is downloaded again and later runs only do the transcription itself, and the raw data and the scripts behind every number on this page are published — see where the numbers come from.
Every figure below was timed on clipquill.com itself, on 2026-09-18, over the public internet, in a real window on this machine, with whisper-base, the multilingual tier. Cold means the browser cache was emptied first, so that run carried the first download of the model and the runtime and then the transcription itself; it does not include reading the file or decoding the audio; cached means the model was already in the cache, so only the transcription happened.
  · 13 s of audio → 44.1 s cold, 23.1 s cached
  · 60 s of audio → 37.2 s cold, 22.8 s cached
  · 277 s of audio (4 min 37 s) → 96.3 s cold, 102.2 s cached
  · 1609 s of audio (26 min 49 s) → 543.4 s (9 min 3 s) cold, 479.3 s (7 min 59 s) cached
Timed on a 92-second loop of mixed material (clean synthetic speech, real human speech, and real speech under background noise) run through whisper-base, so these are harder than a clean studio recording. These are the seconds the page itself reports, from handing the file over to the finished transcript; wall-clock runs a few seconds longer because it also counts reading and decoding the file.
On a phone-sized window (375 px wide, same machine) the layout holds — no horizontal overflow, no broken layout. That is a narrow window on a desktop, not a phone. On a real phone — a Samsung Galaxy Z Fold5 — we ran two audio files: an English clip of 31 s (0.5 MB) and a Chinese clip of 28 s (0.38 MB). Both finished and gave text we could copy. We did not put a stopwatch on them, so no phone timings are listed here and none are claimed. On a phone this will occupy the whole device, longer files take longer, and the phone gets warm — on those two runs it was warm to the touch, though the phone stayed responsive and did not stutter. We have not tested it while charging.
We do not say “any length” — at 30 minutes this page stops. A file longer than that is refused up front rather than run until the browser runs out of memory and the tab dies. Clips of one to three minutes are where it feels smoothest. Long files still work, they just take real time, and on a phone a long file means a long wait.

Accuracy and languages

What we measured. On 2026-10-03 we ran a set of fixed test clips through the English model (Moonshine base) and compared the result word by word against a reference transcript, counting every word that came back wrong, missing, or added. The clips are real human speech — some recorded clean, some with light background noise — which is the audio most people actually transcribe. Across those typical files: 6 wrong words out of 85 reference words — a word error rate of 7.1% on the default English tier, and 8.2% on the optional higher-accuracy tier (whisper-small). By how hard the audio is:

  • real speech, no added noise — 2 clips, 28 words — 7.1% wrong
  • real speech with light background noise — 2 clips, 57 words — 7.0% wrong

Four clips of English, single speaker, on one machine — a sample, not a benchmark, and not a promise about your audio. The two rows above are the typical case; we also test deliberately harsh audio — a clip pushed through a telephone band with echo and hiss, and two clips recorded in heavy background noise — and accuracy drops there: 18.2% on the telephone clip and 21.3% on heavy noise. That is the honest edge of what this model does on a phone-sized machine, and it is why we say read the transcript rather than trust it. The full set, including the clips that went badly, is on the measured word error rate page. Read the result before you use it; numbers, names and technical terms are what you should check yourself. Because everything happens on your own device, you can re-run a file as many times as you like at no cost. Music under speech, several people talking at once, and strong accents are still unmeasured, so nothing is claimed about them.

Languages. The model bundled with this page is multilingual and carries 99 language tokens, so files in those languages can be transcribed. You have to pick the language that is spoken — this build of the model cannot work it out on its own. Pick the wrong one and the transcript comes back in the language you picked: choose English for a Chinese clip and you get English, not Chinese — we tried it, and a 23-second Chinese clip came back as English prose about buying a latte, which is not what was said. Whatever you pick is remembered on this device. Of those 99 we have run English audio and one short Chinese clip ourselves; that is the whole of what we have tested. The two are not equally documented: for English we publish the word error rate measured above, and for Chinese we have that one clip — 23 seconds, 97 characters of reference. It came back in Traditional characters, not Simplified; converted, 6 of those 97 characters were wrong (上线 as 上限, 设计 as 涉及, 那家 as 大家, 来得及 as 来的急). One clip is not enough to call a percentage, so for Chinese we print the count and not a rate.

How to save an Instagram Reel to your phone

Same for YouTube, TikTok and Facebook — anything you can get onto your device works here.

This tool reads a file you already have, so the first step is getting the Reel onto your device. Pick whichever matches your situation:

Your own Reel

  1. Open Instagram and go to your profile.
  2. Open the Reel you want.
  3. Tap the three dots, then Save (Android) or Save to Camera Roll (iPhone).
  4. The Reel is now in your gallery or Photos app — upload it above.

Someone else's Reel

  1. Open the Reel and tap the three dots.
  2. Choose Copy link.
  3. Paste that link into Instagram's own Save feature if it is offered, or ask the creator for the file.
  4. If Instagram does not offer saving for that Reel, you can instead record your screen while it plays, then upload the recording.

On a computer

  1. Right-click the Reel and choose Copy link if you need the reference.
  2. To get an actual file, use Instagram's built-in save on mobile and transfer it, or record the screen with your system's screen recorder.
  3. Upload the resulting file above.

Deliberately simple: this page does not fetch videos from Instagram for you. A field that promises to pull a Reel from a link would have to send that link to a server, and that is exactly what this page refuses to do.

Who made this, and where the numbers come from

There is no About page here and no author byline, so there is no name to trust instead of the data. What there is: every figure on this page was measured by whoever runs this page, and the raw data, the measurement scripts and the reference transcripts are published so that anyone can check them or re-run them.

If a figure on this page disagrees with the published data, the published data is right and this page is wrong. Say so and it gets fixed here.

FAQ

Is this Instagram transcript generator really free?

Yes. There is no signup, no daily cap and no credit system. The model is downloaded to your browser once and every transcript after that costs nothing on our side, because your own device is doing the work.

Do I need to upload my Reel?

No. The file stays in your browser tab. An Instagram reels transcript is produced locally and disappears when you close the tab. Nothing is kept.

Can I just paste a link?

No. This page reads a file that is already on your device, and it never sends anything anywhere — which also means it cannot fetch a video from a link for you. Save the video first, then upload it. The three ways to do that for a Reel are written out above, under How to save an Instagram Reel to your phone. The same goes for YouTube, TikTok and Facebook: whatever you can save to your own device, this page can read. Where the file came from does not matter, only that you have it.

How long does it take?

It depends on your device, not on us. The measured figures above are real runs on real hardware, not estimates. Two things dominate: the first file you run also downloads about 78.4 MiB, and after that the time scales with the length of the audio — a one-minute clip is quick, a half-hour file is not.

How accurate is it?

Most accurate on quiet audio with one person speaking clearly. It drops when there is background music, several people talking over each other, or a strong accent — that is true of any automatic Instagram video to text tool, not just this one. Whatever the audio, read the result before you use it: numbers, names and technical terms are what you should check yourself. On typical files (clean or light-noise real speech, the audio most people actually have) the English default is 7.1% wrong and the optional higher-accuracy tier 8.2%; on deliberately harsh audio it climbs to 21.3%. Because everything runs locally, you can re-run as many times as you like at no cost.

Can I get subtitles for my Reel?

Yes. Export as SRT or VTT and you have a subtitle file ready to drop into your video editor. That is the usual way people turn an Instagram reels to text result back into captions.

Can I download the transcript?

Yes — TXT, SRT and VTT download buttons appear once the transcript is ready. There is no Instagram transcript downloader account to create and nothing to unlock.

Does it work on a phone?

The layout is built for a phone screen and the file picker works the same way. We have now run it on a real phone — a Samsung Galaxy Z Fold5 — on two audio files, an English clip of 31 s (0.5 MB) and a Chinese clip of 28 s (0.38 MB). Both finished and gave text we could copy. We did not put a stopwatch on them, so no phone timings are listed here and none are claimed. On a phone this will occupy the whole device, longer files take longer, and the phone gets warm — on those two runs it was warm to the touch, though the phone stayed responsive and did not stutter. We have not tested it while charging. Practical advice: start with a short clip and be on Wi-Fi, because the first run downloads about 78.4 MiB.

Can I transcribe video to text for free, without an account?

Yes. There is no account, no daily cap and no credit system. The model is downloaded to your browser once, and after that every run costs nothing on our side because your own device is doing the work. Free here is not a trial with a meter on it, and it is not a quota that resets tomorrow.

Is a video to transcript the same thing as captions?

They are the same text, taken two ways. A transcript is the plain words in order — export it as TXT and that is what you get. Export the same words as SRT or VTT and you have a subtitle file with timings, ready to drop into a video editor. One run of video to text transcription here gives you both, so you do not have to decide before you start.

What happens to my file after the transcript is done?

Nothing, because nothing was sent. The file stays in the browser tab and is gone when you close it. There is no copy on a server to ask us to delete and no retention period to read — which is exactly what most services that convert video to text do not put in writing.