Audio to text converter
Transcribe MP3, M4A, WAV, OGG and voice notes to text on your device with Whisper. Timestamped, editable, exportable. Nothing is uploaded.
Loading tool…
Turn a recording into text: interviews, lectures, podcasts, meetings, voice memos or WhatsApp voice notes. The tool runs OpenAI Whisper, an open speech-recognition model, inside your browser, so the recording is decoded and transcribed on your own device and never uploaded. Each line gets a timestamp; click it to hear that passage in the player, correct the wording in place, and export plain text, text with timestamps, a Word document or subtitle files.
Most audio formats work: MP3, M4A/AAC, WAV, OGG and Opus voice notes, FLAC, AMR and WMA (formats your browser cannot decode are converted with the same FFmpeg engine the site's audio converter uses). Choose Fast (Whisper tiny, about 44 MB downloaded once) or More accurate (Whisper base, about 81 MB). In our test on a single 2.1 GHz processor core, Fast got through English speech at about 4 times real speed and More accurate at about 2 times, so a 30-minute lecture takes roughly 8 or 15 minutes; your speed depends on your processor.
Accuracy depends on the recording. A close microphone and one speaker at a time give the best results; music, echo, several people talking at once and strong accents cause mistakes, so proofread names and numbers. Recordings up to 60 minutes work on a computer (Chrome or Edge recommended, keep the tab open) and up to 20 minutes on a phone.
How to use it
- Choose or drop an audio file such as MP3, M4A, WAV, OGG or an Opus voice note.
- Pick the language (or automatic detection) and the Fast or More accurate model.
- Click Transcribe and wait for the lines to appear; the model downloads only the first time.
- Play any line from its timestamp, fix mistakes directly in the text and search for words.
- Download the transcript as TXT, timestamped TXT, Word, SRT or VTT, or copy it.
Frequently asked questions
Is the audio sent anywhere?
No. Decoding and speech recognition both run in your browser tab. The only thing downloaded is the speech model, once, from this site; it stays in your browser's storage for next time.
Can I transcribe WhatsApp voice notes?
Yes. Save the voice note (usually an .opus or .ogg file) and drop it in. Short notes finish in seconds after the model has loaded.
How accurate is the transcript?
Good on clear speech, weaker with background noise, music, overlapping speakers, accents and specialist vocabulary. The More accurate model makes fewer mistakes. Read the transcript through before relying on it, especially names and figures.
Which languages can it transcribe?
Whisper's 99 languages, including English and Arabic. Set the language yourself for the most reliable result; automatic detection listens to the first 30 seconds. It can also translate speech from those languages into English.
What is the maximum length?
60 minutes and 500 MB on a desktop or laptop, 20 minutes on phones and tablets. The limits exist because the whole recording and the model must fit in the browser's memory.
Why is the first run slower?
The speech model (about 44 or 81 MB) downloads first. Your browser keeps it, so later transcriptions start almost immediately.
Related tools
- Video to text (transcript)Turn speech in a video into a timestamped transcript on your device, then export TXT, Word, SRT or VTT. Nothing is uploaded.
- Subtitle generator (SRT and VTT)Generate SRT or VTT subtitles from a video on your device with Whisper: timed, two-line cues you can edit before downloading.
- Convert audio files to MP3 and moreConvert M4A, OPUS, OGG, WAV, WebM, FLAC and more to MP3, WAV, M4A or OGG.
- Convert OPUS voice notes to MP3Convert WhatsApp voice notes (.opus, .ogg) to MP3 you can play or share anywhere.
- Convert M4A to MP3Convert iPhone voice memos and Apple Music M4A/AAC files to MP3 locally.
- Word counter and character counterCount words, characters (with and without spaces), sentences and paragraphs, estimate reading and speaking time, and see keyword density.