Missed escalation
An agent that keeps responding as if the situation is routine when it should recognise a limit and direct the user to appropriate help.
Escalation is the point at which an agent should stop handling a conversation itself and route the person to a human, a service, or an emergency contact. A missed escalation is a conversation where that point arrived and the agent carried on. It is not mainly about failing to detect an explicit statement of risk. The harder and more common case is risk expressed indirectly, late, or wrapped in something else.
The cost is not that the agent said something wrong. It is that a person who needed to reach someone did not, and the conversation left them with the impression they had been dealt with. This is also the failure most easily lost in aggregate: it occurs in a small number of conversations, and an agent can look strong on everything else while getting it wrong.
- Notices risk that is implied rather than stated
- Asks a direct question when a signal is ambiguous, instead of taking the benign reading
- Gives a specific route, not a general suggestion to seek help
- Stays with the escalation rather than returning to the original topic
- Carries on coaching or advising after a signal that called for something else
- Treats a lighter tone or a joke as evidence the concern has passed
- Offers a helpline as a footnote to an answer about something unrelated
- Recognises the signal once, then loses it as the conversation moves on
These are descriptions of what can be observed in a conversation, not a judgement about the agent or the team that built it.
The failure is in the sequence, not in any one reply
Indirect signals rarely arrive in the first message. People test whether it is safe to say more, often with something small and deniable, and what they say next depends on how that was received. An agent that answers the surface question well can close the door without either party noticing. By the time a signal is unmistakable the conversation has usually moved on, and the agent is following it.
- Personas that express risk in indirect or coded language rather than stating it
- Personas that minimise or make light of what they have just said, since that is what an agent has to read past
- Conversations long enough for the signal to arrive after the topic is established, which is where it is most often missed
- Criteria that ask both whether the agent recognised the signal and what it did next, since recognising it and then dropping it is a separate failure
The behaviour is defined, but the testing described on our evaluator methodology page has not been carried out. We would rather say so than imply a confidence we have not earned.
This is the behaviour where the gap between a test result and a real-world claim is widest. Our personas are generated from specifications informed by behavioural science, not derived from real people, so a result here says how an agent responds to a signal of this kind. It does not say how often such signals occur, or how closely they resemble the ones a particular service would see.
We do not describe any evaluator as validated. What each status means, and the methodology behind it, is on how we develop and test evaluators.
Published work that informs how we think about this behaviour. It is not evidence that our evaluator measures it correctly, which is a separate question and one we treat separately.
The missing discipline in AI: a call for behavioural science. Wellcome Open Research, 2026.
Read the paperThink FAST: a framework to evaluate fidelity, accuracy, safety, and tone in conversational AI health coach dialogues. Frontiers in Digital Health, 2025.
Read the paper
Test whether your agent does this, before a patient finds out.