Turn speech into text.Privately, for free.
Drop in an audio or video file and get a clean transcript, powered by an on-device AI speech recognition model. Nothing is ever uploaded anywhere. Export as plain text, or as SRT / VTT subtitles with timestamps.
ScriptFlow runs a small Whisper speech-recognition model (Xenova/whisper-tiny.en) directly in your browser using WebAssembly, entirely on your own hardware. The first time you use it, your browser downloads the model (roughly 40–75MB) and caches it locally — every transcription after that works fully offline, with no server, no API key, and no cost. Because transcription runs on your device rather than a data center GPU, longer files will take longer to process, especially on older or lower-powered machines.
Idle
This can take a moment — play a round of Snake while you wait
Arrow keys / WASD, or swipe on mobile.
Why ScriptFlow
100% private
Your file never leaves your device. There is no server to send it to.
Free, no account
No sign-up, no API key, no usage limits. Ever.
TXT, SRT & VTT
Export plain text, or timestamped subtitles for video editors and players.