Tutorial2 hours ago

1. Building whisper.cpp on a Mac in 33 Seconds, and the Two Failures That Exit Zero

whisper.cpp cloned, configured and built in 33 seconds on an M4 Pro, then transcribed a 16.6 minute speech in 8.7 seconds. It also reported success on a file it could not read.

The WJS Desk

Sep 30, 2026 · 8 min read

Photo by Andreu Marquès on Pexels

whisper.cpp turns speech into text on your own computer. No account, no API key, no audio leaving the machine. It is a C and C++ port of OpenAI's Whisper model by Georgi Gerganov, the same person behind llama.cpp, and it sits at 53,982 stars on GitHub with an MIT license, more than 100 contributors and a last push four days before we wrote this.

This is part 1 of a four part course. We built it from source, downloaded six models, transcribed three real public domain recordings, and deliberately fed it the files a normal person would actually have. This part gets you from nothing to a working transcript, and covers the two failures that report success.

What you will end up with

A command that takes an audio file and writes a plain text transcript and a subtitle file next to it. On our machine the first real test, a 16 minute 34.9 second radio address by Franklin Roosevelt, came back as 1,761 words of text in 8.7 seconds. That is 114 times faster than listening to it.

StepTime on our machine
Clone v1.9.4 (shallow, 50 MB)12.3 s
Configure with CMake4.6 s
Build everything (39 binaries)16.3 s
Download the base.en model (148 MB)8.1 s
First ever transcription (11 s clip)16.9 s
Same transcription, second time0.24 s

The conditions matter, so here they are. An Apple M4 Pro with 12 cores and 48 GB of memory, macOS 26.5.1, Apple clang 17.0.0, CMake 4.4.3, plugged into power. Download times depend on your connection and will differ. Build times on an 8 GB M1 or an Intel Mac will be longer, and we did not test either.

Why a local transcriber at all

Transcription is the unglamorous job behind a lot of content work: show notes for a podcast, subtitles for a video, a searchable record of an interview, quotes you can check instead of paraphrase. The paid services charge by the minute and need your audio uploaded. whisper.cpp does it on the laptop you already own, and on Apple Silicon it uses the GPU through Metal without you configuring anything.

We did not take the speed claims on trust. The rest of this course measures them on recordings that are not the project's own demo clip.

Prerequisites

  • A Mac with the Xcode command line tools (xcode-select --install if clang --version fails)
  • CMake and git. With Homebrew: brew install cmake git
  • About 150 MB of disk for the source and build, plus 78 MB to 1.6 GB per model
  • Realistic time: five minutes, most of it downloads

Step 1: clone and build

We pinned the release tag so the numbers in this course stay reproducible:

git clone --depth 1 --branch v1.9.4 https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
cmake -B build
cmake --build build -j --config Release

The configure step printed one warning, ARM -march/-mcpu not found, -mcpu=native will be used, which is harmless. The build produced 39 files in build/bin and took 32 MB. The one you want is build/bin/whisper-cli, an 845 KB binary. Metal support is on by default, so there is no flag to find.

Pro tip: if you only want the command line tool, cmake --build build -j --config Release --target whisper-cli skips the servers, benchmarks and tests. From a clean directory, configure plus that build took us 18.4 seconds, so it saves disk (23 MB against 32 MB) more than time.

Step 2: download a model

The code is useless without a model file. The repo ships a script that pulls converted models from Hugging Face:

sh ./models/download-ggml-model.sh base.en

That fetched a 148 MB file in 8.1 seconds. The .en models are English only. There are 30 model names in the script. These are the six we downloaded, with sizes as they landed on disk:

ModelSize on diskDownload time
tiny.en78 MB3.5 s
base.en148 MB8.1 s
small.en488 MB19.5 s
large-v3-turbo-q5_0574 MB51.0 s
medium.en1,534 MB109.0 s
large-v3-turbo1,625 MB109.3 s

Start with base.en. Part 2 measures which model is actually worth its size.

Step 3: the first transcription, and the 17 second surprise

./build/bin/whisper-cli -m models/ggml-base.en.bin -f samples/jfk.wav

The bundled 11 second Kennedy clip came back word perfect. The command took 16.9 seconds, even though whisper's own timing report said the whole job took about 150 milliseconds. We ran it three more times: 0.24, 0.24 and 0.25 seconds.

To pin it down we took a second build that had never run and timed it in pieces. Its first --help took 4.0 seconds and its second took 0.03. Its first real transcription then took 14.6 seconds. So a freshly built whisper.cpp costs roughly 18 seconds, once, before it goes quiet. We believe this is macOS checking the new binaries and libraries plus the GPU driver preparing its compute kernels, but we did not instrument it further, so treat that as our reading rather than a diagnosis. The practical point is simpler: do not benchmark your first run, and do not assume it is hung.

Step 4: a real file

The demo clip proves nothing. We downloaded Roosevelt's fireside chat of 5 June 1944 on the fall of Rome from the Internet Archive, where it is marked public domain. It is a 16 MB MP3, 994.9 seconds long, recorded off a 1944 broadcast.

mkdir -p out
./build/bin/whisper-cli -m models/ggml-base.en.bin -f fdr-1944-rome.mp3 -otxt -osrt -of out/fdr

MP3 works directly, no conversion. -otxt writes out/fdr.txt, -osrt writes out/fdr.srt with timestamps, and -of sets the file name without the extension. Two runs took 8.9 and 8.7 seconds. The output began:

Ladies and gentlemen, the President of the United States.
My friends, yesterday on June 4, 1944, Rome fell to American and Allied troops.
The first of the access capitals is now in our hands.

Roosevelt said Axis capitals. That is the error pattern worth knowing about: the words that sound alike and matter most are exactly the ones a small model guesses. Here is the odd part. When we cut out only the first 60 seconds and transcribed that, base.en wrote Axis correctly, and changed see there monuments to see their monuments. Same audio, same model, different words, depending on where the 30 second windows fell. If a quote matters, check it against the recording.

What broke

We tried the wrong turns a reader would take. Two of them fail and report success.

An m4a file fails, and exits 0

Voice Memos, most phones and a lot of Zoom exports produce .m4a. whisper.cpp decodes audio with miniaudio, which handles WAV, MP3 and FLAC but not AAC. It prints this and stops:

read_audio_data: failed to read audio data
error: failed to read audio file 'fdr-60s.m4a'

Then it exits with status 0. The Homebrew build does the same. A shell script that checks $? will happily move on to the next file. The fix is to convert first, and macOS ships a converter so you do not need ffmpeg:

afconvert -f WAVE -d LEI16@16000 -c 1 memo.m4a memo.wav
./build/bin/whisper-cli -m models/ggml-base.en.bin -f memo.wav -otxt -of out/memo

That writes 16 kHz mono 16 bit WAV, which is what the model consumes internally anyway.

A missing output folder fails, and exits 0

Our very first run on the FDR file pointed -of at a folder that did not exist yet. whisper.cpp transcribed all 16 minutes, printed failed to open 'out/fdr.txt' for writing, threw the transcript away, and exited 0. Create the folder first.

Gotcha: in whisper-cli 1.9.4 an unreadable input and an unwritable output both exit 0. A missing model exits 3. If you script it, check that the output file exists and is not empty rather than trusting the exit status.

Our Mac went to sleep in the middle of the benchmarks

This one was our setup, not the tool. We ran these tests unattended, and our first build reported 1,075.7 seconds of wall time against 47.9 seconds of CPU. The power log showed the machine dropping into maintenance sleep and only running our commands in short wake windows. caffeinate -i did not stop it. caffeinate -s, which blocks system sleep while on mains power, did, and every number in this course was taken under it, with the power log checked for sleep events during each run. If you queue a long batch job on a laptop and walk away, wrap it:

caffeinate -s ./build/bin/whisper-cli -m models/ggml-base.en.bin -f fdr-1944-rome.mp3 -otxt -of out/fdr

Common mistakes

  • Running it without -m. The default model path is models/ggml-base.en.bin relative to the folder you are in, not the install. From anywhere else it fails with exit 3.
  • Feeding it m4a or video files and trusting the exit code. Convert with afconvert first.
  • Using a .en model on non-English audio. It is English only by design. Part 3 shows what happens when you try.
  • Timing the first run and concluding it is slow.
  • Looking for whisper-cpp in Homebrew. The formula was renamed to whisper.cpp. The old name still resolves with a warning.

What we did not test

Honest limits for this part. The machine running these tests already had Homebrew's whisper.cpp 1.9.4 installed, and we did not uninstall it to time a fresh brew install, so we have no install time for that route. It pulls in ggml, llama.cpp and sdl2-compat as dependencies, and it uses its own ggml build (0.25.3) rather than the one vendored in the source tree. We did not test an Intel Mac, a base 8 GB machine, the optional Core ML encoder, or Linux and Windows. Every timing above is one machine.

The escape hatch

Nothing here touches your system. The whole install is one folder:

cd ..
rm -rf whisper.cpp

That removes the source, the build and every model you downloaded.

Next

Part 2 runs six models over a 65 minute audiobook chapter with a known text, and scores every word. That is where the model sizes above stop being abstract.

Share

whisper.cpp turned a 16.6 minute speech into text in 8.7 seconds on a Mac, no API key. It also exits 0 when it cannot read your file. #OpenSource #Whisper #SpeechToText #macOS

Never miss a ship

The best stuff that shipped this week, delivered every Thursday. Free, no spam. We read all the boring stuff so you get the fun parts.

Keep reading