Skip to content
Thought Leadership

The Escalation Experience: AI's Most Important Moment

By Nic Fouhy••14 min read
The Escalation Experience: AI's Most Important Moment

Every AI phone system has a moment its vendor would rather you did not look at. The caller has said something the agent cannot handle, and the agent has to say "let me get a person for you." Most demos skip it. Most proposals bury it under a heading called limitations. It is treated as the point where the product stops working.

We think it is the product.

The escalation experience is the handful of seconds between an AI deciding it is out of its depth and a human picking up the thread. Get it right and the caller ends up better served than if a person had answered from the start, because they were processed quickly and then looked after personally. Get it wrong and the caller learns, in one call, that the business put a machine between them and help.

We have built voice agents for property managers, plumbers and compliance teams, and the handoff is where every one of those projects was won or lost. This post is the argument for designing it on purpose, what a good one contains, when it should fire, and how to tell whether yours is working.

Why is the handoff the most important moment in an AI system?

Because it is the only moment the caller is judging the system. For routine calls the AI is invisible. The job gets booked, the question gets answered, nobody thinks about it. Escalation is the exception the caller will remember and repeat to other people, so it carries the whole reputation of the build.

There is a second reason, and it is about the buyer. Every business owner we have sat across from has the same objection about voice AI, and it is never about accuracy or cost. It is "what happens when it can't cope." They have all been stuck in a phone tree. They have all shouted "operator" at an IVR. They are asking whether they are about to do that to their own customers.

A well-designed escalation is the answer to that objection. It is also the hardest part of the system to fake in a demo, which is why so few vendors show it.

What does a caller feel at the moment of handoff?

Relief or dread, and the design decides which. The caller has already explained their problem once. What they fear is explaining it again to someone who does not know they called. What they hope for is a person who picks up already knowing what is wrong, so the conversation carries on from where it was.

Think about what the human version of this used to carry. A good receptionist who transferred a call did three things without being asked. They told the caller who they were being put through to and why. They told the colleague who was on the line and what they wanted. And if the colleague was not there, they came back to the caller with a plan. None of that was written on a job description. It is what "putting a call through" meant.

An AI that transfers a call and drops it has removed the receptionist and kept none of that. The table below is the difference in practice.

MomentHandoff as failure modeHandoff as designed experience
TriggerAgent gets confused, loops, or hangs upAgent recognises a defined condition and says so plainly
What the caller hears"Please hold" or silenceWho they are going to, why, and how long it will take
What the human receivesA ringing phoneA one-paragraph brief: who, what, urgency, what has been tried
If nobody answersVoicemail, or the line dropsAgent tells the caller the fallback and books a callback window
AfterwardsNothingConfirmation to the caller, note on the record, escalation logged
Side by side comparison of a dropped AI call transfer and a designed escalation where the human receives a written brief before picking up
Same trigger, two completely different experiences for the person on the line

What does a well-designed escalation contain?

Three parts, in order. A promise to the caller that names who is coming and when. A brief to the human that means they never ask the caller to repeat themselves. And a closed loop, so that if the human does not answer, the caller is told what happens next. Miss any one and the other two do not save it.

What should the caller be told?

The specific next step and a realistic time. "I'm going to get Mere, who looks after your building, and she'll have everything you've told me. It'll be about a minute." That sentence names a person, gives a reason, sets an expectation and reassures them they will not have to start over. Compare it with "transferring you now," which tells them nothing and sounds like the start of a wait.

The promise has to be true, which constrains what the agent can say. If nobody on the roster picks up after hours, the agent should not claim a person is seconds away. It should say what it can deliver: a callback inside a stated window, a contractor already dispatched, or an emergency line for anything unsafe. We wrote about how that plays out in the after-hours playbook for property managers, where the contractor cascade runs before any human is woken.

What should the human receive?

A brief they can read in ten seconds before the line connects. Caller's name and number. Which property, job or account. What the problem is, in one line. How urgent the agent rated it and why. What the agent already tried or told them. That is enough for the human to open with "Hi Sarah, I can see the water's still coming through the hallway ceiling" and the caller feels caught.

This is the piece that makes the whole thing possible, and it exists because language became programmable. The agent has held a conversation. It can summarise that conversation into fields. The brief is a by-product of a system that was already listening, which means it costs nothing extra to produce and everything to leave out.

What happens when the human does not pick up?

The agent keeps the promise it made. It tells the caller the person was not available, states the fallback and executes it. That might be a callback window, a text to the caller with a reference, an escalation to the next person on the roster, or for anything unsafe, a clear instruction to ring 111. The caller hangs up knowing what happens next, and the record shows the attempt.

This is the branch most systems never design, because it is embarrassing to plan for your own staff not answering. Plan for it anyway. The value of the whole system is that nothing falls through the gap between a machine and a person, and the gap is widest at 2am.

Three-part escalation flow showing the promise to the caller, the written brief to the human, and the closed loop fallback when nobody answers
Promise, brief, closed loop. The third part is the one everybody forgets to build

When should an AI escalate to a person?

On a short list of defined conditions, written down before launch and reviewed against real calls afterwards. The caller asks for a person. The caller is distressed. The request is ambiguous after one clarifying question. The stakes are high enough that a wrong answer costs more than a human minute. Or the agent has failed twice at the same step. Any of those, and it hands over.

The list is deliberately short. Every rule you add is another judgement the agent makes under pressure, and a long list produces a hesitant agent that escalates the wrong things. Five conditions, held firmly, work better than fifteen held loosely.

Why is escalating too little worse than escalating too much?

Because the two errors cost different amounts. An agent that hands over a routine call wastes a minute of a person's time. An agent that tries to handle a call it should have escalated can lose a customer, miss an emergency or give advice it had no business giving. The costs are not symmetrical, so the threshold should not be either.

There is a business case for this beyond the risk. A person picking up a pre-briefed call spends their time on the part of the conversation that requires a person. The routine intake was already done. So an escalation that looks like a cost on a spreadsheet is often the most efficient minute that person spends all day, because every second of it is judgement and reassurance rather than "can I get your address again."

Set the threshold conservative, launch, then tighten it based on the transcripts. You will find that some of your escalation conditions never fire, and one you did not think of fires every week. That is normal. The first month of real calls teaches more than any amount of pre-launch design.

Should the caller ever be told they are talking to an AI?

Yes, and ideally before they need to ask. An agent that is upfront about what it is, and about the fact that a person is a sentence away, removes the tension that makes callers hostile. The people who hate voice AI mostly hate being deceived by it, or trapped by it. Remove those two and most of the hostility goes with them.

This also changes what the escalation feels like. If the caller knew from the start that a person was available, the handoff reads as the system working. If they only discover it after fighting the agent, it reads as an apology.

Five escalation trigger conditions arranged as a checklist with the request for a human, distress, ambiguity, high stakes and repeated failure
Five triggers. Written down before launch, checked against transcripts after

How do you design and measure your own escalation experience?

Write the three parts as a script before anyone builds anything, listen to ten real handoffs in the first fortnight, and track two numbers: how often callers repeat themselves after transfer, and how often a fallback fired because nobody answered. Both should trend towards zero. If either is climbing, the system is leaking trust in the exact place it should be earning it.

Start with the brief, because everything else depends on it. Decide what five fields a human needs to open the call well, and make the agent fill them on every call whether or not it escalates. Then write the promise sentence for each trigger condition, in plain language, with a real time estimate. Then write the fallback for each promise. If you cannot write the fallback, do not make the promise.

We build these into every voice AI engagement and into CallCover, and the design work takes an afternoon. The value it protects is the whole reason the client bought the system. Anyone selling you an AI phone agent without a written escalation design is selling you a phone tree with better diction.

Escalation design: six things to have in writing before launch

1. The trigger list.

Five conditions, no more. Request for a human, distress, ambiguity after one clarification, high stakes, repeated failure. Add a sixth only when transcripts prove you need it.

2. The promise sentence for each trigger.

Names the person or role, gives the reason, states a real time. If you cannot deliver the time after hours, the after-hours version says something different.

3. The brief.

Five fields the human sees before the line connects. Who, which record, what is wrong, how urgent and why, what has been tried. Filled on every call, not only on escalations.

4. The fallback for each promise.

What the agent does when nobody picks up. Callback window, next person on the roster, text with a reference, or 111 for anything unsafe. Written, tested, logged.

5. The disclosure.

How and when the agent says it is an AI and that a person is available. Early, plain, once.

6. The two numbers.

Repeat rate: how often callers re-explain after transfer. Fallback rate: how often nobody answered. Review both weekly for the first two months.

If you want a second pair of eyes on an escalation design, or on a system that has one only in theory, talk to us. It is the part of the build we care most about, and the part we are asked about least.

Got a reaction to this?

Disagree, agree, or want to dig deeper. I read every message.

Thanks, . I'll be in touch.