Explicit model consent
The measured pinned q8 download starts only after approval and remains cached in the browser.
Advertisement
Transcribe speech locally, review every cue and decide whether to download a caption file, apply the track in the Video Editor, or burn approved text into a copy of the clip.
Result proof
Input
Local video audio
Local pass
Pinned Whisper q8 + cue grouping
Result
Editable SRT · VTT · TXT
How to verify it: Play the source with the cue list, verify names and numbers, inspect every start/end time, and download the reviewed format—not the first model draft.
Caption workflow
Pinned local speech recognition creates a typed cue track with explicit timestamps, while split, merge, delete and apply-all controls keep the human review in charge.
SRT · VTT · TXT
Editable caption outputs
The measured pinned q8 download starts only after approval and remains cached in the browser.
Edit text and timecodes, split long cues, merge neighbours, or delete mistakes before applying.
Optional burn-in muxes the source audio back into the captioned result.
Know before you process: Speech recognition can mishear names, accents and noisy recordings. Always review captions; clips without detected speech remain unchanged.
Create a first caption pass, correct product names and export WebVTT for the player.
Style short, readable cues and burn an approved track into a delivery copy.
Export plain text for review while retaining SRT timecodes for editing.
Replace auto-generated platform captions with a reviewed, portable track.
No. The browser decodes the audio locally and runs the pinned Whisper model in a worker. The file and generated cues remain in this tab.
The pinned q8 revision uses 76.9 MB of model weights and about 82 MB including tokenizer and configuration files. It downloads only after explicit approval and is normally browser-cached.
Yes. Every cue has editable text, start and end times. You can also split, merge and delete cues, then apply one style and position across the track.
SRT, WebVTT and plain TXT are generated from the reviewed track. Timestamp rounding carries correctly across second boundaries.
Yes. The quality-preserving browser render keeps source resolution, frame rate and audio, with a hard limit of 60 seconds for opaque output or 30 seconds when preserving WebM transparency. Longer clips can still export SRT/VTT/TXT.
A transcription run is limited to 15 minutes and also has a device-aware PCM memory preflight. Oversized, no-audio, and no-speech inputs report the condition and leave the source and existing caption track unchanged.
Cancel terminates the model worker, discards that generation, and recreates a fresh worker on retry. A late result cannot overwrite a newer request.
Pinned local model, explicit consent, no media upload.
Generate editable captions →