Scientific foundations
An ongoing research programme in conversational AI evaluation.
PromptSafe is built on an evolving programme of behavioural science, evaluation methodology, and empirical research designed to improve how conversational AI is evaluated.
As conversational AI is increasingly used in high-stakes settings, organisations need evaluation methods that are transparent and grounded in evidence rather than subjective judgement.
This research underpins how PromptSafe develops evaluators. It makes the platform more trustworthy without changing the way you use it.
Why behavioural science
AI that talks to people does more than return information. It influences behaviour, sometimes in ways that are harmful or misleading even when the system is technically accurate. These effects are rarely measured by standard testing.
Our founder, Dr Paul Sacher, set out the case for this in a 2026 open letter, co-authored with leading behavioural scientists:
“Even when technically accurate, AI systems can influence behaviour in ways that are harmful, misleading, or misaligned with people’s interests.”
The letter was published on behalf of the Behavioral AI Institute, where Dr Sacher is a co-founder and research director.
PromptSafe was created to translate behavioural science into practical, scalable methods for evaluating conversational AI.
The missing discipline in AI: a call for behavioural science. Wellcome Open Research, 2026.
Read the paperMeasuring conversation quality
FAST, a framework co-authored by our founder and published in Frontiers in Digital Health, evaluates conversations across four dimensions:
- FidelityDoes the agent follow evidence-based behaviour-change practice, not just hand over information?
- AccuracyIs what it says correct, current, and within scope?
- SafetyDoes it recognise risk and signpost to a professional when it should?
- ToneIs it empathic, non-judgemental, and pitched at the right level?
PromptSafe draws on published evaluation research, including the FAST Framework, alongside behavioural science, measurement science, and software engineering principles.
Think FAST: a framework to evaluate fidelity, accuracy, safety, and tone in conversational AI health coach dialogues. Frontiers in Digital Health, 2025.
Read the frameworkWhether AI supports how people think and decide
Most AI evaluation measures technical performance: accuracy, robustness, reasoning, policy compliance. A 2026 preprint co-authored by our founder argues that this misses something a human-facing system does every time it speaks. It shapes how a person understands their situation and what they decide to do next.
The paper names that missing dimension psychological competence:
“The capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioural decision-making in ways that are appropriate to the user, context, and purpose.”
It sets out the properties this depends on, including framing, tone, perceived authority, and how a system communicates uncertainty, and proposes a conceptual framework and assessment methods rather than a single benchmark. It comes from the same programme of work behind PromptSafe: deciding what good conversational AI behaviour looks like before trying to measure it.
This work is at preprint stage and has not been peer reviewed. We publish it openly so the approach can be examined and challenged while it is still taking shape.
Psychological competence as a missing dimension in AI evaluation. arXiv, July 2026.
Read the preprintHow much confidence should a score carry?
A quality score only means something if there is sufficient evidence behind it. We are developing the Evaluation Evidence Framework (EEF), a transparent methodology for quantifying confidence in AI evaluation results. Rather than asking only “How well did the AI perform?”, EEF also asks “How much evidence supports that conclusion?” This work is currently in development and will be published as it matures.
We measure evaluators. We do not assume they are correct.
An evaluator is a measurement instrument. Small changes in its wording can produce very different results, so we treat evaluator design as an engineering problem that can be measured rather than guessed.
PromptSafe applies the same evidence-based methodology when developing evaluator templates and when working with enterprise customers to develop evaluators for their own AI systems.
In practice, that means treating evaluator development as a disciplined process:
- Define the behaviourWe start by defining the behaviour we want an evaluator to measure.
- Build multiple evaluatorsWe create alternative evaluator designs rather than assuming the first version is correct.
- Compare against reference examplesWhere appropriate, we compare candidate evaluators against expert-rated examples to understand which performs most reliably.
- Improve using evidenceEvaluators are version controlled and refined over time as more evidence becomes available.
When evidence is insufficient, we record it as insufficient rather than assuming an evaluator is correct. Evidence accumulates over time as evaluators are used, tested and improved.
The same methodology underpins our Enterprise Evaluator Engineering.
Every evaluation improves the next
Over time, evaluation data helps improve PromptSafe’s evaluation methodologies and understanding of conversational AI behaviour. This supports better evaluators and new research into conversational AI evaluation.
PromptSafe learns from how AI systems behave under evaluation. We do not learn from customers’ proprietary agent prompts or system designs. The evidence it draws on, including personas, evaluators, simulation outcomes and aggregated metrics, is held in the PromptSafe Research Knowledge Base under the PromptSafe Research & Data Governance Framework, processed with appropriate deidentification and governance safeguards, and never published in a way that identifies individual customers or users.
Our aim is not only to apply published research, but to contribute new methodologies and evidence that advance the evaluation of conversational AI.
Grounded in academic collaboration
Our founder is an honorary senior lecturer in the Faculty of Medicine at Imperial College London and a collaborator at the Health Impact Lab there.
The Health Impact Lab works to close the gap between health research and real-world adoption, applying implementation science to move evidence-based innovations beyond academia and into patient care. Turning research into impact is the same goal PromptSafe is built to serve: turning behavioural science into evaluation that teams actually use.
Where this is heading
- Published foundationsFAST framework
- TodayPromptSafe platform
- Current researchEvaluation Evidence Framework
- NextValidation studies
- FutureIndustry benchmarking
Planned parts of the programme
These components will expand the PromptSafe Scientific Foundations programme over time.
- Coming soonValidation studiesIndependent and internal studies testing the reliability of our evaluation methodologies.
- Coming soonBenchmark reportsAggregated, anonymised findings on how conversational agents perform across behavioural dimensions.
PromptSafe is helping advance the science of conversational AI evaluation.
Our research programme is ongoing. The published research described on this page has been peer reviewed. Other methodologies, including the Evaluation Evidence Framework (EEF), are under active development. We are explicit about the status of each.