Scientific foundations
An ongoing research programme in conversational AI evaluation.
PromptSafe is built on an evolving programme of behavioural science, evaluation methodology, and empirical research designed to improve how conversational AI is evaluated.
We do not assume that an AI evaluator is correct simply because it produces a score. We test evaluators and document the evidence supporting their use.
As conversational AI is increasingly used in high-stakes settings, organisations need evaluation methods that make their assumptions and evidence explicit, rather than relying on untested judgements.
This research underpins how PromptSafe develops evaluators, without changing the way you use it.
From the Clinical Product Thinking newsletter? The trial is free and needs no card. Your discount code applies at your first payment.
Why behavioural science
AI that talks to people does more than return information. It influences behaviour, sometimes in ways that are harmful or misleading even when the system is technically accurate. These effects are often not captured by standard model evaluation.
Our founder, Dr Paul Sacher, set out the case for this in a 2026 open letter, co-authored with leading behavioural scientists:
“Even when technically accurate, AI systems can influence behaviour in ways that are harmful, misleading, or misaligned with people’s interests.”
The letter was published on behalf of the Behavioral AI Institute, where Dr Sacher is a co-founder and research director.
PromptSafe was created to translate behavioural science into practical, scalable methods for evaluating conversational AI.
Part of what shaped it was doing that work by hand. Reviewing five versions of one agent, conversation by conversation, showed us that a behavioural fix can be confirmed and then quietly come back: why we test for regressions.
The missing discipline in AI: a call for behavioural science. Wellcome Open Research, 2026.
Read the paperMeasuring conversation quality
FAST, a framework co-authored by our founder and published in Frontiers in Digital Health, evaluates conversations across four dimensions:
- FidelityDoes the agent follow evidence-based behaviour-change practice, not just hand over information?
- AccuracyIs what it says correct, current, and within scope?
- SafetyDoes it recognise risk and signpost to a professional when it should?
- ToneIs it empathic, non-judgemental, and pitched at the right level?
PromptSafe draws on published evaluation research, including the FAST Framework, alongside behavioural science, measurement science, and software engineering principles.
Think FAST: a framework to evaluate fidelity, accuracy, safety, and tone in conversational AI health coach dialogues. Frontiers in Digital Health, 2025.
Read the frameworkWhether AI supports how people think and decide
Most AI evaluation measures technical performance: accuracy, robustness, reasoning, policy compliance. A 2026 preprint co-authored by our founder argues that this misses something a human-facing system does every time it speaks. It shapes how a person understands their situation and what they decide to do next.
The paper names that missing dimension psychological competence:
“The capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioural decision-making in ways that are appropriate to the user, context, and purpose.”
It sets out the properties this depends on, including framing, tone, perceived authority, and how a system communicates uncertainty, and defines constructs and directions for assessment rather than presenting them as validated measures. Through the PromptSafe research programme, we are investigating how constructs such as these can be operationalised as practical AI evaluations, and what evidence is required before those evaluations should be relied on.
This work is at preprint stage and has not been peer reviewed. We publish it openly so the approach can be examined and challenged while it is still taking shape.
Psychological competence as a missing dimension in AI evaluation. arXiv, July 2026.
Read the preprintWe measure evaluators. We do not assume they are correct.
An evaluator is a measurement instrument. A plausible evaluator can be repeatable, produce convincing scores, and still fail on the behaviour it is intended to measure. We therefore test evaluators rather than assuming that a well written evaluation prompt produces a trustworthy measurement.
PromptSafe uses this methodology to guide the development and testing of evaluator templates, and when working with enterprise customers to develop evaluators for their own AI systems.
In practice, that means treating evaluator development as a disciplined process:
- Define the behaviourSpecify what the evaluator is intended to assess, and the conditions under which that behaviour can meaningfully be observed.
- Operationalise itTranslate the behaviour into observable criteria, and distinguish dimensions that may need to be assessed separately.
- Build controlled challengesCreate cases designed to test whether the evaluator can distinguish the intended behaviour, while avoiding obvious confounds such as response length or unrelated differences between cases.
- Check the test casesA test case is not treated as correct simply because it was designed to represent a particular behaviour. Where appropriate, challenge cases are independently reviewed to assess whether they manipulate the intended behaviour without introducing unrelated differences.
- Test the evaluatorAssess properties such as repeatability, applicability, controlled discrimination and known failure modes. Where useful, candidate evaluators may also be stress tested across different judge models or configurations.
- Seek independent evidenceWhere appropriate, independently authored or expert-assessed cases test whether development findings generalise beyond the examples used to build the evaluator.
- Test in dynamic conversationsControlled examples establish whether an evaluator can recognise deliberately constructed differences. Dynamically generated conversations test its behaviour under more varied, multi-turn conditions.
When evidence is insufficient, we record it as insufficient rather than assuming an evaluator is correct. Evidence develops through structured testing, independent assessment where appropriate, and comparison with external benchmarks where these exist.
Some behavioural and clinical risks are specific to an AI system, population or use case, and are not adequately captured by standard evaluators. In other cases, the behaviour may be well defined in clinical or behavioural science but no established AI evaluation exists for measuring it.
Sacher AI works with clinical, behavioural and product teams to define these behaviours, develop candidate evaluators, test their behaviour under defined conditions, document the evidence and known limitations, and implement them within PromptSafe.
Projects can include
- Evaluator specifications
- Controlled challenge case development
- Test case quality and confound checks
- Repeatability and applicability testing
- Evaluator stress testing across judge models and configurations
- Documented evidence and known limitations
- Implementation within PromptSafe
Turning a behavioural construct into an evaluator
Published behavioural science can tell us which properties of human AI interaction may matter. It does not automatically tell us how to measure them.
PromptSafe is developing a methodology for translating behavioural constructs into observable conversational behaviours, then testing whether AI evaluators behave as intended under defined test conditions.
The first step is to translate the construct into an operational definition and a small number of measurable dimensions. Whether an evaluator can then distinguish those dimensions under appropriate test conditions is a separate question, and the one the evaluator methodology above is built to investigate.
Importantly, failure at one stage does not necessarily mean the underlying behavioural construct is wrong. The evaluator, the operational definition, or the conditions used to test it may be responsible.
This work is ongoing. We report the evidence supporting an evaluator separately from the theoretical basis of the construct it is intended to measure.
We investigated whether a construct relating to preservation of user agency could be translated into measurable properties of an AI conversation. Early testing suggested that two proposed dimensions could not be distinguished. Controlled testing later showed that the test design itself was contributing to that result. Once the behavioural property was varied while other features were held more constant, the dimensions separated.
This is development evidence from deliberately constructed cases, not evidence that the full construct has been successfully operationalised. It illustrates why PromptSafe treats evaluator development as a measurement problem rather than a prompt-writing task.
Testing the same evaluator across different judge models also showed why model comparisons require appropriate test coverage. An apparent difference between models on a single challenge pair reversed when tested across ten controlled pairs representing different forms of the same behaviour. This remains development evidence rather than an independent model benchmark.
Understanding the evidence behind a result
A quality score only means something in the context of the evaluator and testing conditions that produced it. We are developing the Evaluation Evidence Framework (EEF) to make that evidence more explicit.
Rather than asking only “How did the AI perform?”, EEF also asks what evidence supports relying on that judgement, what has been tested, and what remains unknown.
Different forms of evidence are reported separately rather than collapsed into a single validation or confidence score. This work is currently in development and will be published as it matures.
Research and continuous improvement
Evaluation data can support ongoing research into PromptSafe’s evaluation methodologies and conversational AI behaviour.
PromptSafe research uses evidence about how AI systems behave under evaluation. We do not use customers’ proprietary agent prompts or system designs for research. Where we use evaluation evidence, including personas, evaluators, simulation outcomes and aggregated metrics, it will be held in the PromptSafe Research Knowledge Base under the PromptSafe Research & Data Governance Framework, with a documented deidentification process before workspace content enters it. Published research will contain only aggregated or anonymised findings, and will not identify individual customers, organisations or users without their explicit written permission.
Our aim is not only to apply published research, but to contribute new methodologies and evidence that advance the evaluation of conversational AI.
Grounded in academic collaboration
Our founder is an honorary senior lecturer in the Faculty of Medicine at Imperial College London and a collaborator at the Health Impact Lab there.
The Health Impact Lab works to close the gap between health research and real-world adoption, applying implementation science to move evidence-based innovations beyond academia and into patient care. That focus on translating research into practice aligns with PromptSafe’s aim of turning behavioural science into evaluation methods that teams can use.
This is an affiliation of our founder. It does not imply institutional endorsement or independent validation of PromptSafe’s evaluation methods.
Where this is heading
- Published foundationsFAST and behavioural science research
- TodayPromptSafe simulation and evaluation
- Current researchEvaluator methodology and Evaluation Evidence Framework
- PlannedEvaluator studies
- Longer termEvaluation coverage and industry benchmarking
Planned parts of the programme
These components will expand the PromptSafe Scientific Foundations programme over time.
- Coming soonEvaluator studiesInternal and independent studies examining evaluator performance, repeatability and generalisation under defined conditions.
- Coming soonBenchmark reportsAggregated, anonymised findings on how conversational agents perform across behavioural dimensions.
What PromptSafe results do and do not show
PromptSafe evaluates how an agent behaves in defined simulated conversations against specified criteria. Results can support product development, quality assurance and governance. They are not safety certification, regulatory approval or a substitute for appropriate human review.
The evidence supporting an evaluator depends on the behaviour being measured, the test cases used and the conditions under which it has been assessed. We therefore document what has been tested, what the results support and what remains unknown.
We do not claim clinical validation or equivalence to human expert judgement. Where these properties have been tested, the evidence concerning repeatability, applicability, controlled discrimination, calibration and known limitations is reported separately for each evaluator.
Where the evidence is insufficient to support a claim, we say so.
Important behavioural risks can be missed by conventional incident reporting
Incident reports are valuable for identifying discrete and observable failures. They are less well suited to effects that emerge gradually through repeated interaction, such as inappropriate agreement, reduced user agency or increasing reliance on an AI system.
Identifying these risks requires evidence from behavioural science, clinical research and human factors research, alongside conventional safety sources.
A recent review in Nature defines several relevant concepts, including sycophancy, confirmation bias, automation bias, excessive delegation, persuasion, fluency drift, and overload or degradation. These do not all describe the same kind of problem. Some concern agent behaviour, some concern human responses to AI, and others concern the interaction or system performance over time.
The PromptSafe research programme is examining how relevant constructs can be translated into observable and testable properties of agent behaviour and human AI interaction. Areas of interest include:
- Inappropriate agreement or reinforcement
- Failure to investigate important contextual information
- Erosion of user agency
- Communication that creates unwarranted trust
- Consistency and behavioural degradation across longer conversations
- Fidelity to defined clinical or behaviour change methods
These are research priorities, not a claim that every published construct is currently measured by a PromptSafe evaluator. Some overlap with observable behaviours already assessed in the platform. Others require further operational definition, controlled cases and evaluator testing.
A named construct is not an evaluator. Before an evaluator is released, the target behaviour, unit of observation, likely confounders, applicability boundaries and supporting test evidence must be defined.
Weiner and Schwartz’s work on contextual error shows why inquiry matters. A response can sound empathic while still failing to investigate a contextual clue that should affect the care plan. Translating this distinction from clinician assessment to AI conversations requires separate development and testing.
Contextual errors and failures in individualizing patient care: A Multicenter Study. Annals of Internal Medicine, 2010.
Read the studySafety and security of large language models in healthcare. Nature, August 2026.
Read the reviewA review can define and organise relevant concepts. It does not establish that a behaviour occurs at a particular rate or that it can be measured reliably by an AI evaluator.
Our aim is to contribute evidence and methods that improve the evaluation of conversational AI.
Our research programme is ongoing. We clearly distinguish peer reviewed research, preprints and methodologies that remain under development, including the Evaluation Evidence Framework (EEF).