Skip to main content
Behaviour library

The behaviours we test for

Observable behaviours that matter in conversational AI, how they fail, and how we test them.

Most evaluation asks whether an answer was correct. The harder question is how an agent behaves with a person: whether it gathers what it needs before it advises, whether it holds a boundary when someone pushes, whether it stays useful when a conversation gets difficult.

Safety matters, but it is not the only dimension of behaviour worth testing.

Each entry sets out what the behaviour is, what good and bad look like, how the failure tends to emerge over a conversation rather than in a single reply, and how far the evidence behind our evaluator actually goes.

The other half of the problem

Naming a behaviour is not the same as being able to measure it

A list of behaviours tells you what to look for. It does not tell you whether the evaluator can actually recognise them. Testing an agent with an untested evaluator moves the uncertainty rather than resolving it, which is why every entry here carries the evidence status of the evaluator behind it, and why we test evaluators as instruments in their own right.

You can see how we develop and test evaluators, including where the evidence runs out, in how we develop and test evaluators. And why we test for regressions explains why a behaviour that appears fixed still needs testing again.

Find out how your agent behaves under pressure before a patient does.