Advertisement

Private video AI

Auto subtitles with an editable caption track

Transcribe speech locally, review every cue and decide whether to download a caption file, apply the track in the Video Editor, or burn approved text into a copy of the clip.

Result proof

A caption track you can audit before publishing

01

Input

Local video audio

02

Local pass

Pinned Whisper q8 + cue grouping

03

Result

Editable SRT · VTT · TXT

How to verify it: Play the source with the cue list, verify names and numbers, inspect every start/end time, and download the reviewed format—not the first model draft.

Editable local captions: Video + speech, Transcribe + review, SRT · VTT · videoEditable local captionsINPUTVideo + speechTRANSCRIBE + REVIEWSpeech becomeseditable cues00:07.42OUTPUTSRT · VTT · videoCOMPLETETimed cues editable

Caption workflow

Transcribe, correct and style before anything reaches the timeline

Pinned local speech recognition creates a typed cue track with explicit timestamps, while split, merge, delete and apply-all controls keep the human review in charge.

SRT · VTT · TXT

Editable caption outputs

01

Explicit model consent

The measured pinned q8 download starts only after approval and remains cached in the browser.

02

Cue-level correction

Edit text and timecodes, split long cues, merge neighbours, or delete mistakes before applying.

03

Audio-safe result

Optional burn-in muxes the source audio back into the captioned result.

Know before you process: Speech recognition can mishear names, accents and noisy recordings. Always review captions; clips without detected speech remain unchanged.

What people use this for

Tutorials and demos

Create a first caption pass, correct product names and export WebVTT for the player.

Social clips

Style short, readable cues and burn an approved track into a delivery copy.

Interview transcripts

Export plain text for review while retaining SRT timecodes for editing.

Accessibility cleanup

Replace auto-generated platform captions with a reviewed, portable track.

Limits worth knowing

  • Names, accents, overlapping speakers and noisy audio require human correction.
  • Transcription is limited to 15 minutes per run and can refuse sooner when decoded PCM would exceed the device-aware memory ceiling.
  • Browser codec support determines whether the audio stream can be decoded; the first run needs explicit approval for an approximately 82 MB pinned model bundle.
  • Quality-preserving burn-in is limited to 60 seconds for opaque output or 30 seconds while preserving transparency; caption-file export remains available for supported longer transcripts.

How the local workflow runs

  1. Choose a clip and approve the exact pinned model revision.
  2. Decode the audio locally and generate timed caption cues.
  3. Edit text and timing; split, merge or delete cues as needed.
  4. Set the apply-all style and position.
  5. Download SRT/VTT/TXT, or create a supported-length captioned copy and verify playback.

Frequently asked questions

Is my video uploaded for transcription?

No. The browser decodes the audio locally and runs the pinned Whisper model in a worker. The file and generated cues remain in this tab.

How large is the model download?

The pinned q8 revision uses 76.9 MB of model weights and about 82 MB including tokenizer and configuration files. It downloads only after explicit approval and is normally browser-cached.

Can I correct the generated captions?

Yes. Every cue has editable text, start and end times. You can also split, merge and delete cues, then apply one style and position across the track.

Which caption formats can I download?

SRT, WebVTT and plain TXT are generated from the reviewed track. Timestamp rounding carries correctly across second boundaries.

Can I burn captions into the video?

Yes. The quality-preserving browser render keeps source resolution, frame rate and audio, with a hard limit of 60 seconds for opaque output or 30 seconds when preserving WebM transparency. Longer clips can still export SRT/VTT/TXT.

What happens with long clips, no audio, or no speech?

A transcription run is limited to 15 minutes and also has a device-aware PCM memory preflight. Oversized, no-audio, and no-speech inputs report the condition and leave the source and existing caption track unchanged.

What does Cancel do during transcription?

Cancel terminates the model worker, discards that generation, and recreates a fresh worker on retry. A late result cannot overwrite a newer request.

Ready to try it?

Pinned local model, explicit consent, no media upload.

Generate editable captions