Journal

Voice is not chat with a microphone

We thought putting our chat agent on the phone was a two-week job. It took eleven, and we threw away the first four. Here is what we got wrong, in the order we got it wrong.

6 min read

We built the chat agent first, and it worked. It answered questions from a company’s own documents, it attached the passage it used, and when it could not find a source it said so instead of inventing one. People liked it.

Then a dental practice asked for the same thing on the phone. We scoped it at two weeks. Take the thing that already answers questions, put speech recognition on the front and speech synthesis on the back, done.

It took eleven weeks. We threw away most of the first four.

The mistake is in the framing, and it is worth naming plainly, because we have since watched two other teams make it. A voice agent is not a chat agent with a microphone. It is a different product that happens to share a model.

People do not talk the way they type

Typed questions are edited. Somebody thinks, writes, glances at it and presses send. What arrives is a sentence.

Spoken questions are not edited, because there is nowhere to edit them. What arrives is this:

Hi, yeah, I was in, um, two weeks ago? For the, the thing with the crown, and they said come back but I don’t know if I need to, do I need to book that or…

Our retrieval was built for questions. That is not a question. It has no subject, it trails off, and the actual request is implied rather than stated. Passing the raw transcript to a system tuned on clean queries returned nothing useful about half the time.

What fixed it was a step we had not budgeted for: a pass that turns a spoken turn into a stated intent before anything searches for anything. It costs about 200ms and it was the single highest-value thing we added.

You have about a second and a half

In chat, three seconds of a typing indicator is fine. It even helps. It reads as consideration.

On a phone call, a second and a half of silence and the caller says “hello?” At three seconds they assume the line dropped. Nobody waits politely for a pause on the telephone, because for a hundred and fifty years a pause on the telephone has meant something went wrong.

Our first working version had a budget that looked like this:

  • Speech to text, final transcript: 300ms
  • Intent pass: 200ms
  • Retrieval: 200ms
  • Model, full response: 900ms
  • Speech synthesis, first audio: 350ms

That is just under two seconds before the caller hears anything, and it was worse in practice, because none of those are averages. They are the good case. We were losing people before the agent opened its mouth.

Three things bought it back. Stream the model output and start synthesising on the first clause rather than the last. Start retrieval on the partial transcript instead of waiting for the final one. And the least sophisticated fix was the most effective. Let it say “let me check that for you” immediately, which is both true and worth about 800ms of cover.

They will interrupt, and you have to let them

Chat has no equivalent of this. Nobody interrupts a paragraph.

On a call, the moment the agent starts down the wrong path the caller talks over the top of it. If the agent keeps going, the call is finished. People will not repeat themselves to something that will not stop.

So it has to stop mid-word, discard the response it had already generated, and listen again, while telling a real interruption apart from somebody saying “mhm”. We got that wrong in both directions for a fortnight: first an agent that ploughed on regardless, then one that stopped dead every time the caller breathed.

There is no scrollback

This is the one that changed how we write the answers.

In chat, a wrong answer sits on the screen. The person re-reads it, spots the error and corrects it. The transcript is a safety net.

Spoken, the answer is gone the instant it is said. Nobody can check it. Which means a voice agent has to be shorter than a chat agent, more certain than a chat agent, and it has to read back anything with a consequence attached:

Thursday the fourteenth, ten past three, with Dr Okafor. Shall I book that?

Our chat agent answers in about sixty words. The voice agent answers in under twenty-five, and when it cannot, that is the signal it should be offering a person instead.

Names, numbers and postcodes

Speech recognition is very good at sentences and unreliable at exactly the things a booking needs. Surnames, house numbers, postcodes, dates of birth, anything spelled aloud.

There is no clever fix. You confirm digit by digit. You check against what you already hold rather than capturing from scratch. “Is this still the mobile ending 4 4 1?” is a far better question than “what is your number?” For anything you genuinely cannot verify, you hand over. We tried to be cleverer than that for three weeks and got nowhere.

What we would tell you before you start

  • Build the chat version first, even if voice is what you want. It is the same knowledge layer, and you can see what it gets wrong because it is written down.
  • Write the handover path before the happy path. The question is not whether it fails. It is whether it fails to a person or to a dial tone.
  • Cap the length. A voice agent that will not stop talking is worse than a phone menu, because at least a menu ends.
  • Log the audio, not just the transcript. Half our real problems were audible and invisible in text: someone trailing off, a pause in the wrong place, the agent sounding certain while being wrong.

Where it landed

The practice takes around forty calls a day. The agent handles the ones that are genuinely routine, such as opening hours, whether a treatment is covered, or moving an appointment. It hands over the rest, which is roughly two in three.

One in three is not the number anyone puts in a pitch deck. We think it is the honest one. The other two thirds are people who are worried about something, or whose situation is unusual, or who are simply owed a human being. A system that pretends otherwise is not saving anyone time. It is moving the cost onto the caller and calling it efficiency.

The rule we ended up with, and the one we would start from next time: if it cannot say the thing in twelve seconds, it should be offering to pass you to someone who can.

Answering the same questions all day?

Concierge is the version of this we have already built. Voice and chat on the same knowledge, with a person a sentence away.

See Concierge