1. Building whisper.cpp on a Mac in 33 Seconds, and the Two Failures That Exit Zero
whisper.cpp cloned, configured and built in 33 seconds on an M4 Pro, then transcribed a 16.6 minute speech in 8.7 seconds. It also reported success on a file it could not read.
The WJS Desk
Sep 30, 2026 · 8 min read

whisper.cpp turns speech into text on your own computer. No account, no API key, no audio leaving the machine. It is a C and C++ port of OpenAI's Whisper model by Georgi Gerganov, the same person behind llama.cpp, and it sits at 53,982 stars on GitHub with an MIT license, more than 100 contributors and a last push four days before we wrote this.
This is part 1 of a four part course. We built it from source, downloaded six models, transcribed three real public domain recordings, and deliberately fed it the files a normal person would actually have. This part gets you from nothing to a working transcript, and covers the two failures that report success.
What you will end up with
A command that takes an audio file and writes a plain text transcript and a subtitle file next to it. On our machine the first real test, a 16 minute 34.9 second radio address by Franklin Roosevelt, came back as 1,761 words of text in 8.7 seconds. That is 114 times faster than listening to it.
| Step | Time on our machine |
|---|---|
| Clone v1.9.4 (shallow, 50 MB) | 12.3 s |
| Configure with CMake | 4.6 s |
| Build everything (39 binaries) | 16.3 s |
| Download the base.en model (148 MB) | 8.1 s |
| First ever transcription (11 s clip) | 16.9 s |
| Same transcription, second time | 0.24 s |
The conditions matter, so here they are. An Apple M4 Pro with 12 cores and 48 GB of memory, macOS 26.5.1, Apple clang 17.0.0, CMake 4.4.3, plugged into power. Download times depend on your connection and will differ. Build times on an 8 GB M1 or an Intel Mac will be longer, and we did not test either.
Why a local transcriber at all
Transcription is the unglamorous job behind a lot of content work: show notes for a podcast, subtitles for a video, a searchable record of an interview, quotes you can check instead of paraphrase. The paid services charge by the minute and need your audio uploaded. whisper.cpp does it on the laptop you already own, and on Apple Silicon it uses the GPU through Metal without you configuring anything.
We did not take the speed claims on trust. The rest of this course measures them on recordings that are not the project's own demo clip.
Prerequisites
- A Mac with the Xcode command line tools (
xcode-select --installifclang --versionfails) - CMake and git. With Homebrew:
brew install cmake git - About 150 MB of disk for the source and build, plus 78 MB to 1.6 GB per model
- Realistic time: five minutes, most of it downloads
Step 1: clone and build
We pinned the release tag so the numbers in this course stay reproducible:
git clone --depth 1 --branch v1.9.4 https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
cmake -B build
cmake --build build -j --config Release
The configure step printed one warning, ARM -march/-mcpu not found, -mcpu=native will be used, which is harmless. The build produced 39 files in build/bin and took 32 MB. The one you want is build/bin/whisper-cli, an 845 KB binary. Metal support is on by default, so there is no flag to find.
Pro tip: if you only want the command line tool, cmake --build build -j --config Release --target whisper-cli skips the servers, benchmarks and tests. From a clean directory, configure plus that build took us 18.4 seconds, so it saves disk (23 MB against 32 MB) more than time.
Step 2: download a model
The code is useless without a model file. The repo ships a script that pulls converted models from Hugging Face:
sh ./models/download-ggml-model.sh base.en
That fetched a 148 MB file in 8.1 seconds. The .en models are English only. There are 30 model names in the script. These are the six we downloaded, with sizes as they landed on disk:
| Model | Size on disk | Download time |
|---|---|---|
| tiny.en | 78 MB | 3.5 s |
| base.en | 148 MB | 8.1 s |
| small.en | 488 MB | 19.5 s |
| large-v3-turbo-q5_0 | 574 MB | 51.0 s |
| medium.en | 1,534 MB | 109.0 s |
| large-v3-turbo | 1,625 MB | 109.3 s |
Start with base.en. Part 2 measures which model is actually worth its size.
Step 3: the first transcription, and the 17 second surprise
./build/bin/whisper-cli -m models/ggml-base.en.bin -f samples/jfk.wav
The bundled 11 second Kennedy clip came back word perfect. The command took 16.9 seconds, even though whisper's own timing report said the whole job took about 150 milliseconds. We ran it three more times: 0.24, 0.24 and 0.25 seconds.
To pin it down we took a second build that had never run and timed it in pieces. Its first --help took 4.0 seconds and its second took 0.03. Its first real transcription then took 14.6 seconds. So a freshly built whisper.cpp costs roughly 18 seconds, once, before it goes quiet. We believe this is macOS checking the new binaries and libraries plus the GPU driver preparing its compute kernels, but we did not instrument it further, so treat that as our reading rather than a diagnosis. The practical point is simpler: do not benchmark your first run, and do not assume it is hung.
Step 4: a real file
The demo clip proves nothing. We downloaded Roosevelt's fireside chat of 5 June 1944 on the fall of Rome from the Internet Archive, where it is marked public domain. It is a 16 MB MP3, 994.9 seconds long, recorded off a 1944 broadcast.
mkdir -p out
./build/bin/whisper-cli -m models/ggml-base.en.bin -f fdr-1944-rome.mp3 -otxt -osrt -of out/fdr
MP3 works directly, no conversion. -otxt writes out/fdr.txt, -osrt writes out/fdr.srt with timestamps, and -of sets the file name without the extension. Two runs took 8.9 and 8.7 seconds. The output began:
Ladies and gentlemen, the President of the United States.
My friends, yesterday on June 4, 1944, Rome fell to American and Allied troops.
The first of the access capitals is now in our hands.
Roosevelt said Axis capitals. That is the error pattern worth knowing about: the words that sound alike and matter most are exactly the ones a small model guesses. Here is the odd part. When we cut out only the first 60 seconds and transcribed that, base.en wrote Axis correctly, and changed see there monuments to see their monuments. Same audio, same model, different words, depending on where the 30 second windows fell. If a quote matters, check it against the recording.
What broke
We tried the wrong turns a reader would take. Two of them fail and report success.
An m4a file fails, and exits 0
Voice Memos, most phones and a lot of Zoom exports produce .m4a. whisper.cpp decodes audio with miniaudio, which handles WAV, MP3 and FLAC but not AAC. It prints this and stops:
read_audio_data: failed to read audio data
error: failed to read audio file 'fdr-60s.m4a'
Then it exits with status 0. The Homebrew build does the same. A shell script that checks $? will happily move on to the next file. The fix is to convert first, and macOS ships a converter so you do not need ffmpeg:
afconvert -f WAVE -d LEI16@16000 -c 1 memo.m4a memo.wav
./build/bin/whisper-cli -m models/ggml-base.en.bin -f memo.wav -otxt -of out/memo
That writes 16 kHz mono 16 bit WAV, which is what the model consumes internally anyway.
A missing output folder fails, and exits 0
Our very first run on the FDR file pointed -of at a folder that did not exist yet. whisper.cpp transcribed all 16 minutes, printed failed to open 'out/fdr.txt' for writing, threw the transcript away, and exited 0. Create the folder first.
Gotcha: in whisper-cli 1.9.4 an unreadable input and an unwritable output both exit 0. A missing model exits 3. If you script it, check that the output file exists and is not empty rather than trusting the exit status.
Our Mac went to sleep in the middle of the benchmarks
This one was our setup, not the tool. We ran these tests unattended, and our first build reported 1,075.7 seconds of wall time against 47.9 seconds of CPU. The power log showed the machine dropping into maintenance sleep and only running our commands in short wake windows. caffeinate -i did not stop it. caffeinate -s, which blocks system sleep while on mains power, did, and every number in this course was taken under it, with the power log checked for sleep events during each run. If you queue a long batch job on a laptop and walk away, wrap it:
caffeinate -s ./build/bin/whisper-cli -m models/ggml-base.en.bin -f fdr-1944-rome.mp3 -otxt -of out/fdr
Common mistakes
- Running it without
-m. The default model path ismodels/ggml-base.en.binrelative to the folder you are in, not the install. From anywhere else it fails with exit 3. - Feeding it m4a or video files and trusting the exit code. Convert with
afconvertfirst. - Using a
.enmodel on non-English audio. It is English only by design. Part 3 shows what happens when you try. - Timing the first run and concluding it is slow.
- Looking for
whisper-cppin Homebrew. The formula was renamed towhisper.cpp. The old name still resolves with a warning.
What we did not test
Honest limits for this part. The machine running these tests already had Homebrew's whisper.cpp 1.9.4 installed, and we did not uninstall it to time a fresh brew install, so we have no install time for that route. It pulls in ggml, llama.cpp and sdl2-compat as dependencies, and it uses its own ggml build (0.25.3) rather than the one vendored in the source tree. We did not test an Intel Mac, a base 8 GB machine, the optional Core ML encoder, or Linux and Windows. Every timing above is one machine.
The escape hatch
Nothing here touches your system. The whole install is one folder:
cd ..
rm -rf whisper.cpp
That removes the source, the build and every model you downloaded.
Next
Part 2 runs six models over a 65 minute audiobook chapter with a known text, and scores every word. That is where the model sizes above stop being abstract.


