4. Talking to It Instead of Typing
Two voice features that look the same. One is metered against a rolling five hour allowance and one is free. Most people run out using the wrong one.
The WJS Desk
Sep 16, 2026 · updated 8 days ago · 5 min read

There are two ways to use your voice with the desktop app. They look similar, they are not the same feature, and one of them is metered while the other is not. Nearly everyone who runs out of voice time is using the wrong one for the job.
The difference in one table
| Dictation | Voice Mode | |
|---|---|---|
| What happens | Your speech becomes text you can edit | A live spoken conversation |
| It replies by | Text, as normal | Speaking back |
| You can edit before sending | Yes | No |
| Plan | Available broadly | Paid plans |
| Counts against a voice allowance | No | Yes, rolling five hour window |
That last row is the practical one. Voice Mode draws on a plan-dependent allowance measured over rolling five hour windows, with one active voice conversation at a time. Dictation does not touch it. You can draft by voice all day and spend nothing from the allowance.
Dictation is the one you will use daily
It is the unglamorous feature and it is the one that changes your week. You hold a key, talk, and your words appear in the input box as editable text. Then you fix the two words it misheard and send.
Why that is better than typing for a specific kind of message: most people speak at roughly three times their typing speed, and the messages worth using AI for are the ones with a lot of context in them. The context is exactly what people leave out when typing, because typing it is work.
So dictation is not really a speed feature. It is a context feature. Compare what you would type:
write an email declining the project
with what you would say out loud in the same number of seconds:
I need to decline this project. It is a website
rebuild for about eight thousand, they want it in
three weeks, and I am already booked until November.
I have worked with them before and it went fine, so
I want to keep the relationship. Say no clearly,
offer to look at it in the new year, and do not
apologise more than once.
The second one produces something you can send. It is also the exact prompt shape from our Beginner lesson on roles and reasons, and dictation is the cheapest way to actually supply that much context.
The habit worth forming. When a request feels like it needs explaining, stop typing and dictate it. Ramble. Include the things you would leave out in writing because they seemed like too much detail. Those details are the whole reason the output will be usable.
When Voice Mode earns its allowance
Voice Mode is a conversation. It listens, it speaks, and it handles interruption. The cases where that beats typing:
- Your hands are busy. Cooking, driving, walking, carrying something.
- Thinking out loud about a decision. Back and forth, five or six exchanges, none of which you want written down.
- Practising something spoken. An interview, a difficult conversation, a language. This is where it is genuinely without substitute, because typing a rehearsal is not a rehearsal.
- Explaining something complicated. Faster to say than to write, and it can interrupt to ask.
Where it is worse, and this catches people out: anything you need to keep. A spoken answer is hard to scan, hard to copy, and you cannot skim it. If the output is something you will paste somewhere, type or dictate and read the reply.
Getting either to work
Both need the microphone permission from lesson 2. If you declined it, the feature will not appear rather than explaining itself, which is the confusing failure mode:
System Settings, Privacy and Security, Microphone
then enable ChatGPT
Three practical notes from using them:
Speak normally. Over-enunciating makes transcription worse, not better. Read a sentence aloud the way you would to a colleague.
Say the punctuation only when it matters. Dictation infers most of it. Saying "new paragraph" is worth it; saying "comma" every time is not.
Always read before sending. Names, technical terms and numbers are where transcription fails, and those are exactly the parts that change the meaning. A misheard "fifteen" for "fifty" survives a quick glance.
Dictating in a second language
Worth its own section because it is where dictation helps most and is least written about.
If English is not your first language, the gap between what you can say and what you can type quickly is usually much wider than it is for a native speaker. Dictation closes that gap, and then the model handles the grammar. You are not asking it to guess what you meant from four typed words; you are giving it a full spoken explanation and letting it tidy the English.
I said that in a rush and the grammar is wrong.
Keep exactly what I meant, fix the English, and
do not make it more formal than I was.
That last clause matters. Left alone it tends to raise the register, so a relaxed message to a colleague comes back sounding like a letter from a bank.
Three things that improve transcription
- Get closer to the microphone than feels necessary. Laptop microphones are several feet from your mouth and pick up the room in between. This does more than any setting.
- Pause between thoughts rather than filling the gap. A short silence is a clean break. "Um" is a word it will try to transcribe.
- Spell unusual names once. Product names, surnames and acronyms are where it fails, and they are usually the load-bearing words in the sentence.
A test worth running once
Take a message you already sent this week, one where you put effort in. Dictate the same request from scratch without looking at it, rambling for thirty seconds.
Compare the two. Most people find the dictated version contains three or four facts the typed one left out, not because they forgot, but because typing them felt like too much work at the time. Those facts are the difference between a generic answer and a usable one.
The honest limits
- Accents and background noise. Accuracy varies, and no setting fixes a noisy room.
- Voice Mode needs a paid plan. Dictation is the one to lean on otherwise.
- One voice conversation at a time. There is no running two.
- It is still the same model. Speaking to it does not make it more accurate, and the checking habits from the Beginner track apply unchanged.
Before the next lesson
Take the next message you were going to type with real context in it, and dictate it instead. Ramble deliberately. Compare what comes back to what your typed version would have produced.
Next: handing it your actual files, and the size limits nobody tells you about until you hit one.


