← All AI tools

Speech to Text

Transcribe audio

Drop an audio or video file, run Whisper locally, then copy the transcript or download SRT / VTT.

Runs in your browser. Model weights may download once from Hugging Face and cache locally — your prompts and files are not uploaded to BrowserSpaces.

Whisper tiny

Compact q8 model (~40MB) cached after first Hugging Face download.

Timestamps

Sentence chunks for seeking and subtitles.

Exports

TXT, SRT, and VTT downloads.

Private

Audio is decoded and transcribed in this tab.

Use the tool

First load downloads Whisper (~40MB, q8) from Hugging Face and caches it. Your audio is decoded and transcribed in this tab — not uploaded to BrowserSpaces.

Model

Load Whisper, then drop an audio file.

Decode locally, transcribe locally

Your browser decodes the media to 16 kHz mono PCM. Whisper (via Transformers.js) turns speech into text with optional chunk timestamps.

Before you start

  • Shorter clips finish faster on CPU-only devices.
  • Pick the spoken language when you know it.

Three steps

  1. 1

    Add a file

    MP3, WAV, M4A, OGG, or a video with audio.

  2. 2

    Transcribe

    Load Whisper once, then run on the clip.

  3. 3

    Export

    Copy text or download subtitles.