How it works

How your video becomes subtitles, without leaving your device

Subtitle Generator runs a speech recognition model inside your browser tab. Here is each step from the file to the cues, what your browser downloads and keeps, and what we have and have not measured.

Five steps, all inside your browser tab

1. Your browser opens the file

You pick a video or audio file and the browser hands the page that one file. The page passes it to the engine inside the tab. Nothing is uploaded.

2. The sound is read in slices

Where your browser can, a media reader in a Web Worker reads the sound 30 seconds at a time with your browser's own decoders, and turns it into 16 kHz mono, the form the model takes.

3. Windows are cut at quiet points

The sound is grouped into windows of about 3 minutes. Each cut falls on the quietest half second near the 3-minute mark, so it rarely splits a word.

4. The model writes each window

A voice detector marks the stretches with speech, and the speech recognition model writes them, with a start and an end for every word.

5. Words become cues

The words are cut into cues short enough to read, the cues appear on the page, and you edit them. The SRT or VTT file is written when you press download.

Diagram of where your file goes: your video or audio is opened with the browser's File API, a media reader in a Web Worker reads its sound 30 seconds at a time, and the speech model writes each window of about 3 minutes on WebGPU or WebAssembly; your tab shows the cues, you edit them and save an SRT or VTT file. The file is never sent. The only traffic is the page, the engine, the model (about 60 or 264 MB), and a small voice detector, which come from the site's server on the first run and are then kept.
Your file stays in your browser tab: the page, its engine, and the model come in, and nothing goes out.

The model: Whisper, in two sizes

The words come from Whisper, a speech recognition model that OpenAI published in 2022 with its weights under the MIT License. It runs through whisper.cpp, an open-source C++ version of it, which we compile to WebAssembly so that it can run in a web page. You choose between two sizes:

  • The smaller model (the default) is Whisper base, stored with its weights rounded to 5 bits: a 59.7 MB download.
  • The larger model is Whisper small, stored with 8-bit weights: a 264.5 MB download that needs more memory.

In OpenAI's paper, the larger Whisper models did better in most of its tests, with the smallest gains for English, but we have not yet measured how these two compare on this engine, so the page calls them "smaller" and "larger" rather than promising that one is more accurate. Before the model, a small voice detector, Silero VAD, marks where people speak, so the model spends its time on speech. When you leave the language on automatic, the model identifies it in the first window that has speech, and the page keeps that language for the rest of the file.

Each window is written with the last words of the one before it as a prompt (about 200 characters), so a sentence that runs across a cut keeps its wording and punctuation. As soon as a window is done, its cues appear on the page, and you can edit them while the next window is written. Make subtitles from your video to see it.

How words become cues

The model gives segments of text with a start and an end for every word. A segment can be a long sentence, so the page cuts it into cues a viewer can read in time:

  • At most two lines of 42 columns per cue; a wide character, as in Chinese or Japanese, counts as two columns. Scripts written without spaces are split into words with your browser's word segmenter.
  • At most 7 seconds on screen, and at least 1 second when the gap to the next cue allows.
  • Cuts fall between words: at the end of a sentence first, then at a comma or a pause of half a second or more, then where the two parts come out even. A closing mark stays with the word before it.
  • The cut is timed from the word times. Where the model gave none, the time is shared in proportion to the text.
  • Cues never overlap, and none ends after the file does.
Four steps from speech to subtitle cues: the sound is cut into windows of about 3 minutes at the quietest half second; a voice detector marks the speech and the model writes it with a time for each word; the last words of each window go to the next as context; the words are shaped into cues of at most two lines of 42 columns and 7 seconds, cut at a sentence end, a comma, or a pause.
How the page turns the sound of your file into cues. The limits are the ones in the cue code.

The times are the model's own estimate of where each word starts and ends, so a cue can appear a little early or late. We have not measured by how much on this engine yet. If a cue is off, you can fix a cue by hand: move its start or end by 0.1 second, split it, or merge it with the next.

Which engine build runs

The engine exists in three builds, and the page picks one when you add your first file:

  • WebGPU, tried first when your browser offers a graphics adapter that passes the page's check. If it fails while loading or while writing a window, the page switches to a processor build and writes that window again.
  • WebAssembly on several threads, one thread per logical processor your browser reports. Threads need a cross-origin isolated page, which is why the site sends the COOP and COEP headers.
  • WebAssembly on one thread, the last resort for a browser that cannot run threads, and the build used when your browser reports two logical processors or fewer.

The line under the preview names the build that ran, for example "WebAssembly on several threads". We have not measured the WebGPU build on a real graphics card yet, so the site makes no claim about its speed. Whatever the build, your device does the work: the status line shows an estimate of the time left, measured on your device as it goes, instead of a promised speed.

What your browser downloads, and what it keeps

Everything comes from this site. The models are split into parts of at most 20 MiB, because the host serves static files of up to 25 MiB, and each part is checked against its SHA-256 fingerprint before use, so a damaged download is refused instead of producing wrong subtitles. Sizes are those of the files as published, in megabytes of a million bytes:

Files the page downloads from this site
FileSizeKept in
The engine, multi-threaded WebAssembly build (when WebGPU is not used and your browser reports more than two logical processors)1.6 MBYour browser's ordinary cache
The engine, single-thread build (when threads are not available, or when your browser reports two logical processors or fewer)1.5 MBYour browser's ordinary cache
The engine, WebGPU build (only when the page finds a usable GPU)3.4 MBYour browser's ordinary cache
The smaller model, Whisper base (the default)59.7 MB, in 3 partsCache Storage, after each part's SHA-256 check
The larger model, Whisper small (only if you choose it)264.5 MB, in 13 partsCache Storage, after each part's SHA-256 check
The voice detector, Silero VAD0.9 MB, in 1 partCache Storage, after its SHA-256 check
The media reader (part of the page's scripts)0.3 MBYour browser's ordinary cache

The engine, the model, and the detector download only once you add a file, not when you open the page. Your file, its sound, and your subtitles are never stored by the page: they live in the open tab and go when you close it. To remove the kept model, clear this site's data in your browser. The network test step by step shows how to watch all of this in your browser's developer tools.

How we test it

An automated browser test runs on every change to the site. It loads a built copy of the site in Chromium, the engine of Chrome and Edge, with the site's own security headers, and checks, among other things, that:

  • a 3.5-second English sentence from the LibriSpeech corpus, "Concorde returned to its place amidst the tents.", comes out as cues with at most two words wrong, with the smaller model, from a WAV file, an MP4 video, and a WebM file;
  • after an edit, a 0.1-second move, and a split, the SRT and VTT files read back to exactly the cues on the page, and Chromium's own subtitle reader loads the VTT;
  • on a 60-second file with speech every 10 seconds, cues appear window by window and each sentence starts at its own mark;
  • Cancel stops a run, and the next file reuses the model the browser kept, with no new download;
  • a silent clip gives no subtitles, a damaged part of a file keeps the cues before it, and a file that is not media is refused;
  • every request goes to this site, as a GET or HEAD with no body.

That test uses one short sentence from one speaker, in one language. It shows that the pieces work together; it is not a measure of how accurate the subtitles are.

Limits worth knowing

  • Accuracy and timing are not measured yet on this engine, so no page gives a figure. Music, noise, and people talking over each other make more mistakes.
  • No speaker names, no sound descriptions, no translation. The subtitles are the words spoken, in the language spoken.
  • Long files take long. Where your browser can, the page reads a slice at a time and never holds the whole file's sound at once, but every minute of audio is work for your processor. The longest file we have run so far is 12 minutes.
  • Some files are read whole. Where your browser cannot decode a file's sound in slices, the page decodes it at once and stops at 500 MB or 1 hour.
  • Phones are untested. The larger model may need more memory than a phone has free; the page says so when the browser reports running out.

The open-source components and their licenses are listed on the open-source credits page.

Try it on your own video

Runs in your browser. Your file never leaves your device.

Open the subtitle generator