Turn speech into text.Privately, for free.

Drop in an audio or video file and get a clean transcript, powered by an on-device AI speech recognition model. Nothing is ever uploaded anywhere. Export as plain text, or as SRT / VTT subtitles with timestamps.

ScriptFlow runs a small Whisper speech-recognition model (Xenova/whisper-tiny.en) directly in your browser using WebAssembly, entirely on your own hardware. The first time you use it, your browser downloads the model (roughly 40–75MB) and caches it locally — every transcription after that works fully offline, with no server, no API key, and no cost. Because transcription runs on your device rather than a data center GPU, longer files will take longer to process, especially on older or lower-powered machines.

This can take a moment — play a round of Snake while you wait
Score: 0

Arrow keys / WASD, or swipe on mobile.

Why ScriptFlow

100% private

Your file never leaves your device. There is no server to send it to.

Free, no account

No sign-up, no API key, no usage limits. Ever.

TXT, SRT & VTT

Export plain text, or timestamped subtitles for video editors and players.