How it works
How your video becomes subtitles, without leaving your device
Subtitle Generator runs a speech recognition model inside your browser tab. Here is each step from the file to the cues, what your browser downloads and keeps, and what we have and have not measured.
Five steps, all inside your browser tab
1. Your browser opens the file
You pick a video or audio file and the browser hands the page that one file. The page passes it to the engine inside the tab. Nothing is uploaded.
2. The sound is read in slices
Where your browser can, a media reader in a Web Worker reads the sound 30 seconds at a time with your browser's own decoders, and turns it into 16 kHz mono, the form the model takes.
3. Windows are cut at quiet points
The sound is grouped into windows of about 3 minutes. Each cut falls on the quietest half second near the 3-minute mark, so it rarely splits a word.
4. The model writes each window
A voice detector marks the stretches with speech, and the speech recognition model writes them, with a start and an end for every word.
5. Words become cues
The words are cut into cues short enough to read, the cues appear on the page, and you edit them. The SRT or VTT file is written when you press download.
The model: Whisper, in two sizes
The words come from Whisper, a speech recognition model that OpenAI published in 2022 with its weights under the MIT License. It runs through whisper.cpp, an open-source C++ version of it, which we compile to WebAssembly so that it can run in a web page. You choose between two sizes:
- The smaller model (the default) is Whisper base, stored with its weights rounded to 5 bits: a 59.7 MB download.
- The larger model is Whisper small, stored with 8-bit weights: a 264.5 MB download that needs more memory.
In OpenAI's paper, the larger Whisper models did better in most of its tests, with the smallest gains for English, but we have not yet measured how these two compare on this engine, so the page calls them "smaller" and "larger" rather than promising that one is more accurate. Before the model, a small voice detector, Silero VAD, marks where people speak, so the model spends its time on speech. When you leave the language on automatic, the model identifies it in the first window that has speech, and the page keeps that language for the rest of the file.
Each window is written with the last words of the one before it as a prompt (about 200 characters), so a sentence that runs across a cut keeps its wording and punctuation. As soon as a window is done, its cues appear on the page, and you can edit them while the next window is written. Make subtitles from your video to see it.
How words become cues
The model gives segments of text with a start and an end for every word. A segment can be a long sentence, so the page cuts it into cues a viewer can read in time:
- At most two lines of 42 columns per cue; a wide character, as in Chinese or Japanese, counts as two columns. Scripts written without spaces are split into words with your browser's word segmenter.
- At most 7 seconds on screen, and at least 1 second when the gap to the next cue allows.
- Cuts fall between words: at the end of a sentence first, then at a comma or a pause of half a second or more, then where the two parts come out even. A closing mark stays with the word before it.
- The cut is timed from the word times. Where the model gave none, the time is shared in proportion to the text.
- Cues never overlap, and none ends after the file does.
The times are the model's own estimate of where each word starts and ends, so a cue can appear a little early or late. We have not measured by how much on this engine yet. If a cue is off, you can fix a cue by hand: move its start or end by 0.1 second, split it, or merge it with the next.
Which engine build runs
The engine exists in three builds, and the page picks one when you add your first file:
- WebGPU, tried first when your browser offers a graphics adapter that passes the page's check. If it fails while loading or while writing a window, the page switches to a processor build and writes that window again.
- WebAssembly on several threads, one thread per logical processor your browser reports. Threads need a cross-origin isolated page, which is why the site sends the COOP and COEP headers.
- WebAssembly on one thread, the last resort for a browser that cannot run threads, and the build used when your browser reports two logical processors or fewer.
The line under the preview names the build that ran, for example "WebAssembly on several threads". We have not measured the WebGPU build on a real graphics card yet, so the site makes no claim about its speed. Whatever the build, your device does the work: the status line shows an estimate of the time left, measured on your device as it goes, instead of a promised speed.
What your browser downloads, and what it keeps
Everything comes from this site. The models are split into parts of at most 20 MiB, because the host serves static files of up to 25 MiB, and each part is checked against its SHA-256 fingerprint before use, so a damaged download is refused instead of producing wrong subtitles. Sizes are those of the files as published, in megabytes of a million bytes:
| File | Size | Kept in |
|---|---|---|
| The engine, multi-threaded WebAssembly build (when WebGPU is not used and your browser reports more than two logical processors) | 1.6 MB | Your browser's ordinary cache |
| The engine, single-thread build (when threads are not available, or when your browser reports two logical processors or fewer) | 1.5 MB | Your browser's ordinary cache |
| The engine, WebGPU build (only when the page finds a usable GPU) | 3.4 MB | Your browser's ordinary cache |
| The smaller model, Whisper base (the default) | 59.7 MB, in 3 parts | Cache Storage, after each part's SHA-256 check |
| The larger model, Whisper small (only if you choose it) | 264.5 MB, in 13 parts | Cache Storage, after each part's SHA-256 check |
| The voice detector, Silero VAD | 0.9 MB, in 1 part | Cache Storage, after its SHA-256 check |
| The media reader (part of the page's scripts) | 0.3 MB | Your browser's ordinary cache |
The engine, the model, and the detector download only once you add a file, not when you open the page. Your file, its sound, and your subtitles are never stored by the page: they live in the open tab and go when you close it. To remove the kept model, clear this site's data in your browser. The network test step by step shows how to watch all of this in your browser's developer tools.
How we test it
An automated browser test runs on every change to the site. It loads a built copy of the site in Chromium, the engine of Chrome and Edge, with the site's own security headers, and checks, among other things, that:
- a 3.5-second English sentence from the LibriSpeech corpus, "Concorde returned to its place amidst the tents.", comes out as cues with at most two words wrong, with the smaller model, from a WAV file, an MP4 video, and a WebM file;
- after an edit, a 0.1-second move, and a split, the SRT and VTT files read back to exactly the cues on the page, and Chromium's own subtitle reader loads the VTT;
- on a 60-second file with speech every 10 seconds, cues appear window by window and each sentence starts at its own mark;
- Cancel stops a run, and the next file reuses the model the browser kept, with no new download;
- a silent clip gives no subtitles, a damaged part of a file keeps the cues before it, and a file that is not media is refused;
- every request goes to this site, as a GET or HEAD with no body.
That test uses one short sentence from one speaker, in one language. It shows that the pieces work together; it is not a measure of how accurate the subtitles are.
Limits worth knowing
- Accuracy and timing are not measured yet on this engine, so no page gives a figure. Music, noise, and people talking over each other make more mistakes.
- No speaker names, no sound descriptions, no translation. The subtitles are the words spoken, in the language spoken.
- Long files take long. Where your browser can, the page reads a slice at a time and never holds the whole file's sound at once, but every minute of audio is work for your processor. The longest file we have run so far is 12 minutes.
- Some files are read whole. Where your browser cannot decode a file's sound in slices, the page decodes it at once and stops at 500 MB or 1 hour.
- Phones are untested. The larger model may need more memory than a phone has free; the page says so when the browser reports running out.
The open-source components and their licenses are listed on the open-source credits page.
Making subtitles
Continue exploring this topic
Related topics: Privacy and trust
Try it on your own video
Runs in your browser. Your file never leaves your device.