Skip to main content
For research and clinical teams

PromptSafe in research

The independent safety and evaluation layer for patient-facing conversational AI.

Research teams designing patient-facing AI studies have written PromptSafe into their plans as the safety and evaluation layer. Across those plans the role is the same: stress-test the agent before it reaches a patient, then audit its behaviour while the study runs.

If you are putting a conversational agent in front of patients and need to show it has been tested, this page explains how that usually works.

01 / What this is, and what it is not

Chosen, not yet funded

PromptSafe has been written into competitively reviewed research applications in the United States, the United Kingdom and the European Union, submitted to funders including the NIHR and the NIH. In each one it is named as the independent safety and evaluation layer for a patient-facing conversational agent.

These are applications and bids under review. They are not awarded grants, and we are not describing work that has been funded or delivered.

The signal worth taking from it is narrow and real: teams who design clinical studies for a living, and who had to choose a safety and evaluation approach they could defend to a review panel, chose this one. We would rather state the stage accurately and let you weigh it than let the funder names do work they have not earned.

02 / The role

What PromptSafe does in a study

The role has been consistent across every application, and it splits into four moments.

  • Before deploymentThe agent is run through large numbers of simulated conversations, including difficult and higher-risk ones, to surface where it behaves unsafely, off-brief, inequitably, or off-curriculum. This happens while there is still time to change the agent.
  • During the studyWhere your team holds the ethics approval and lawful basis to do it, samples of real conversations are reviewed retrospectively on an agreed cycle, against the same criteria. That decision and that data stay yours: we process what you choose to put in, under your approvals, not ours.
  • Whenever something is flaggedIt goes to a human expert. PromptSafe supports clinical review, it does not replace it, and nothing it produces is a clinical judgement on its own.
  • At every reporting pointEach cycle produces structured scores and a written report suitable for internal quality assurance, ethics and regulatory review, with the reasoning behind each finding included so reviewers can inspect it.
03 / Where

The areas this work sits in

The applications span four clinical areas, all of them conversational agents speaking directly to patients or their carers.

  • Obesity care, across adults, children and young people
  • GLP-1 care and access pathways
  • Cancer screening uptake and engagement
  • Cancer prehabilitation

These applications are still under review, so we are not naming the partner organisations yet. We will once the work is public and everyone involved is ready to talk about it.

04 / The thinking behind it

Grounded in published behavioural science

Our approach to evaluation is informed by FAST, a peer-reviewed framework for evaluating a generative AI health coach across fidelity, accuracy, safety and tone, published in Frontiers in Digital Health and co-authored by our founder. FAST informs how we think about evaluation. PromptSafe is not built on it and does not run it.

Our founder is an honorary senior lecturer in the Faculty of Medicine at Imperial College London and a collaborator at the Health Impact Lab there. That affiliation is his own and does not imply Imperial endorses or has certified PromptSafe.

The wider research programme, including how we develop and test the evaluators themselves, is set out on the Scientific foundations page.

05 / Regulatory engagement

MHRA AI Airlock

PromptSafe was submitted to the MHRA AI Airlock, the UK regulator’s sandbox for AI as a medical device, to work through how a real-time AI safety layer should be classified and validated. Submission is not selection or approval, and it confers no regulatory status. We mention it because it is the same question research partners raise: who decides whether the safety layer itself is fit for purpose.

06 / Scope

What PromptSafe is not

PromptSafe is used for evaluation and safety testing only. It is not a clinical decision-making tool and it does not deliver care. It supports human clinical review rather than replacing it, and a good result is not a statement that an agent is safe, compliant, or cleared for use. That judgement stays with your team and your governance process.

Evaluation outputs are probabilistic and may vary between runs, which is why every finding carries the reasoning behind it and why anything flagged goes to a human.

Test your agent before a patient does.

If you are writing a study protocol or preparing a funding application, we can talk through what the safety and evaluation work package would need to cover.