How it works
From file to captions, step by step.
Everything happens on your device. There is no server in the middle, no queue, and no account to create.
1. Your file stays where it is
When you choose a video or audio file, it is read straight from your device into the page. The audio is decoded locally and resampled to a format the transcription model understands. At no point is your file uploaded to us or anyone else.
2. The model downloads once
The first time you transcribe, your browser downloads the transcription model — about 40–150 MB. It comes from public CDNs, and once it arrives it is cached on your device, so later visits start almost instantly. You can watch the real download progress while it happens.
3. Transcription runs in your browser
An open speech-recognition model (Whisper, base size) listens to your audio in 30-second pieces and writes out what it hears, with start and end times for every line. You can let it detect the language automatically, or choose English, Hindi, or Urdu. Longer files take longer — a three-minute clip can take a few minutes on a laptop, and mobile phones are slower still.
4. You review and export
Every caption lands in an editable timeline. Click a timestamp to hear that exact moment, fix a word, split a long line in two, merge two short ones, or delete a line. One-step undo has your back. When you’re happy, copy the text or download it as SRT (for YouTube), VTT (for the web), or plain TXT.
Good to know
- Files are limited to 3 minutes and 25 MB on the free tool.
- Clear, single-speaker audio gives the best results. Heavy background noise or crosstalk will show in the captions.
- Always skim the captions before publishing — names and unusual words are worth a second look.
- Works best in a recent Chrome, Edge, Firefox, or Safari, on desktop or mobile.