Voice input Guide

Where dictation saves time, and where it quietly costs you

Speaking is faster than typing. That was never in dispute. The question is what the text costs you after you stop talking, and the answer depends almost entirely on what kind of writing it is.

Published
Reading
8 minutes
Kind
Guide
Basis
Editorial judgement
A two-axis field in which sample points cluster into the two opposite quadrants, one filled blue and one filled orange.
The two axes that decide it. How finished the sentence already is when you start, against how much precision the output needs. Dictation pays in one corner and costs you in the opposite one.

Every dictation product sells a speed multiple. None of them are lying about it, and almost none of them are measuring the thing you care about.

Two numbers, and only one is about writing

The comparison you are usually shown is a speaking rate against a typing rate. Something in the region of 200 words a minute spoken against 40 to 50 typed. Both figures are roughly right for an average person under easy conditions, and the ratio between them is where the familiar “several times faster” claims come from.

The problem is not the arithmetic. It is that a speaking rate measures how fast words leave your mouth, and writing is not the business of getting words out of your mouth. It is the business of producing text that is finished enough to send. Between those two things sits a variable cost that no words-per-minute figure captures: the editing pass.

The useful number is not how fast you can say it. It is how much work is left when you stop.

That cost is not constant. For some writing it is close to zero, and dictation is a straightforward win. For other writing it is larger than the time you saved, and dictation is a net loss that feels like a win, because the speaking part felt fast. This is the whole difficulty: the saving is visible and immediate, and the cost arrives later and attaches itself to a different task.

The test: how finished is the sentence?

There is a single question that predicts this better than anything else we know of.

The question

  • Before you start, is the sentence already finished in your head, so that the only remaining work is getting it down?
  • If yes, you are doing transcription, and speaking is faster than typing at transcription. Dictate it.
  • If no, and you are working out what you think as you go, you are doing composition, and speaking is not obviously faster at composition. It is often slower, because revising spoken text is harder than revising typed text.

The reason composition behaves differently is mechanical rather than mysterious. When you type, revising is continuous and nearly free: you back up a few words, change them, and carry on, and the version you abandoned leaves no trace. When you speak, revision is expensive. You either say the correction out loud and hope the cleanup pass resolves it, or you stop, read what has landed, and fix it with the keyboard, at which point you have switched input methods mid-sentence and lost the flow you were dictating to preserve.

Worth noticingPeople who dictate well often report that they plan more before they start. That is a real skill and it is worth developing. It is also not free, and it is not what the speed comparison is measuring.

This is also why dictation tends to feel better for people who talk through their work anyway, and worse for people who think by writing. Neither is the correct way to work. They are different, and a tool that suits one will not suit the other, which is the sort of thing feature comparisons are structurally unable to tell you.

Four kinds of writing, sorted

Sorting real tasks by the two things that actually vary: how finished the sentence is when you start, and how much precision the output needs.

Where the time goes, by kind of writing
Kind of writingMostlyPrecision neededDictation
The long reply you have been putting offTranscriptionModerateClear win
First drafts, notes to yourself, thinking out loudTranscriptionLowClear win
Short precise messages, names, numbers, timesCompositionHighUsually a loss
Code, commands, structured syntaxCompositionVery highUsually a loss
Anything under a confidentiality obligationEitherEitherCheck first

The first two rows are where dictation genuinely earns its place, and they are not small categories. A great deal of professional writing is a message you already know how to write and have been avoiding for two days. Removing the physical cost of producing it is a real intervention on a real problem, and it works partly for a reason that has nothing to do with speed: speaking is a lower activation barrier than typing, so the thing you were avoiding gets started.

The third and fourth rows are where the arithmetic reverses. A message with three proper nouns, a date and a figure in it is almost all precision and almost no volume. You will spend longer checking the four things that had to be exactly right than you would have spent typing forty words.

Misheard is not the same as misspoken

This is the failure mode to understand before you rely on dictation for anything that matters, and it is a direct consequence of the cleanup pass that makes modern dictation good.

Repairable

Misspoken

You said “Tuesday, sorry, Monday”. The intent is recoverable from the sentence itself, so a cleanup model can resolve it, and generally does. This is the class of error the technology has genuinely solved.

Not repairable

Misheard

You said “Priya” and it wrote “pre-K”. Nothing in the sentence indicates a problem, so the cleanup pass has nothing to work with, and it will hand you a fluent, confident, wrong sentence.

The consequence is that dictated text needs a different kind of proofread from typed text. When you mistype, the result usually looks wrong. When a dictation system mishears, the result usually looks right. You are no longer scanning for typos; you are checking whether specific facts survived, which is slower and much easier to skip.

Our practical rule: after dictating, read back only the proper nouns, the numbers, the dates and the negations. Not the whole message. Those four categories are where a misheard word does damage, and checking them takes a few seconds while rereading everything takes long enough that you will stop doing it by Thursday.

Negations especially“can” and “can't”, “is” and “isn't”. A dropped contraction inverts the meaning of a sentence while leaving it perfectly grammatical, which is the worst combination available.

The constraint nobody plans for

People evaluate dictation at a desk, alone, in a quiet room, and then try to use it in the places they actually work. Three constraints are worth being honest with yourself about before you buy anything.

  • Other people can hear you. In an open office, on a train, in a shared house, dictating a candid message is not available, and this rules out a meaningful share of your working hours regardless of how good the software is.
  • It requires a voice. A cold, a long day of calls, or simply not wanting to talk any more are all normal, and on those days your input method is unavailable in a way a keyboard never is.
  • Most of these tools need a connection. Processing generally happens on a server, so on a bad connection dictation degrades or stops. Check whether the tool you are considering has an offline mode. Most do not.

None of these are objections to the category. They are the reason to treat voice as one input method among several rather than as a replacement, and to be suspicious of any plan that requires it to work every time.

A two week way to find out

The honest answer to “will this save me time” is that it depends on your mix of writing, which you probably have not measured. Here is a way to find out that costs almost nothing and does not require you to trust anybody's multiplier.

  1. Week one: dictate only the obvious wins

    Long replies, first drafts, notes to yourself. Nothing precise, nothing short, nothing sensitive. You are establishing whether the good case is actually good for you before you test the hard case.

  2. Keep a two column note

    Every time you dictate something, write one line: what it was, and whether you would do it that way again. That is the entire method. It takes seconds and it beats recollection, which will otherwise be dominated by whichever attempt was most annoying.

  3. Week two: push at the boundary

    Now try the cases you expect to fail. Short messages. Something with names in it. A reply while other people are in the room. You are looking for where the line is, not for a verdict.

  4. Count the corrections, not the minutes

    Time saved is hard to feel accurately and easy to fool yourself about. Corrections are countable. If a category of writing consistently needs a fix, it belongs on the typed side of your line, whatever the speed multiple says.

  5. Decide per category, not overall

    The output of this is not “dictation works” or “dictation does not work”. It is a short list of the specific kinds of writing you will dictate, which is a much more useful thing to own.

Most people who run something like this end up with a narrower list than they expected and use dictation more than they expected, because the list is reliable. That is the outcome to aim for. A tool you use confidently for four things beats a tool you use hopefully for everything.

On the numbers in this piece

The typing and speaking rates mentioned above are given as rough, widely cited ranges to explain how the standard comparison is constructed. They are not measurements taken by Workflow Vitals, and we have not run a controlled timing study of any dictation product. Where a specific vendor's figures are discussed, we say so and attribute them, as on our Wispr Flow page.

Read next