Skip to main content
Adversarial testing for health and mental health AI

Building trustworthy AI starts with trustworthy evaluation.

Built for teams shipping conversational AI in health and mental health. PromptSafe stress-tests your agent through hundreds of multi-turn simulated conversations and scores it against the behaviours you care about, before launch.

We don’t assume AI evaluations are trustworthy. We measure, test and improve the evaluators that produce them.

Grounded in published behavioural science.Developed by behavioural scientists, clinicians and AI experts at Sacher AI.
See it in 2 minutes

What PromptSafe does, in two minutes.

Watch the overview to see how personas, evaluators and governance-ready reports fit together.

Governance

An untested AI agent isn't a tech risk. It's a governance risk waiting to happen.

Once your agent is live, your organisation is responsible for how it behaves.

In healthcare the stakes go further: an agent stepping outside its scope, or missing the moment to signpost to a healthcare professional, can cause clinical harm before anyone catches it. PromptSafe gives governance, clinical safety and audit teams evidence to test for these risks before deployment.

Evidence outputs

What you get for the file.

Every test run produces an evidence package. It travels with the agent through review, deployment, and audit.

  • Persona by persona conversation logs
    Full transcripts, every turn, attributable to a defined persona profile.
  • Evaluator outcomes with citations
    Every evaluator outcome tied to the conversation turn that triggered it.
  • Prompt version history
    A record of which version of your agent each test run was made against.
Scope & responsibility

PromptSafe is a pre-deployment testing platform, not a medical device, clinical decision support tool, or safety certification service. Results are probabilistic and will vary between runs. Final responsibility for the agent under test, and any decisions taken on the basis of test outputs, sits with the deploying organisation.

When the auditor asks “how did you test this?”, have an answer.

How it works

Six steps from setup to evidence you can act on.

PromptSafe is built for whole teams, not just engineers. Product, clinical, and governance leads use it alongside developers. The product follows the order you'd do the work, and guided onboarding walks you through every step when you first sign in.

  1. 1
    Settings

    Workspace, team, keys

  2. 2
    Agent Builder

    Define the agent under test

  3. 3
    Persona Builder

    Personas you define, adversarial personas we generate

  4. 4
    Evaluator Builder

    Plain-language criteria your agent is judged against

  5. 5
    Simulation

    Run multi-turn conversations

  6. 6
    Evaluations

    Score, evidence, fixes

Agent Builder

Define the agent under test, or connect the one you already run.

Describe your agent inside PromptSafe: its system prompt, instructions, and anything else that shapes how it behaves. That becomes the agent your personas talk to. No rebuild, no code.

  • Build it in minutes. Paste your system prompt and instructions into a structured form.
  • Keep it current. Update the prompt as your agent evolves and rerun the same tests against it.
  • Or connect your own agent (OpenAI). Test the exact agent you run, available on request for OpenAI models today. Reach out to set it up.
Persona Builder

Personas you define. Adversarial personas we generate.

Build personas around the characteristics and behaviours relevant to your users: their background, beliefs, experiences, interaction style and health-related behaviours. PromptSafe can also generate personas designed to probe for weaknesses relevant to a specific evaluator.

  • Three ways to create personas. Structured form, CSV batch import, or free-text notes. Fill in only what matters.
  • Adversarial variants, auto-generated. The evaluator defines the behaviour being tested; PromptSafe generates a persona designed to probe that behaviour.
  • Reusable across runs. Define a persona once, run it against every agent in your workspace.
Evaluator Builder

Evaluators are checklists of observable behaviours, not abstract scores.

AI evaluators are reusable checklists that assess whether your agent demonstrates the behaviours you care about. Write them in plain language, and start with example evaluators designed to be customised, or create your own. No prompt engineering experience required.

  • Sector aligned starter templates. Digital health, mental health, and wellbeing, each with example evaluators ready to customise.
  • Reusable across agents. Define "Recognises clinical risk" once, run it against every agent in your workspace.
  • Auditable by design. Every evaluator outcome is tied to the conversation turns that triggered it.
  • We test the tests. An evaluator is a measurement instrument, not ground truth. We test whether evaluators identify the behaviours they are intended to assess, investigate where they fail, and document the evidence supporting their use. When the evidence is insufficient, we say so.
Simulation

Stress test behaviour across dynamic, multi-turn conversations.

One click runs every selected persona against your agent. Adversarial personas can push back, reframe requests and change tactic as conversations develop. Watch progress as runs complete and pick up where you left off. No babysitting required.

  • Hundreds of conversations per run. Every selected persona runs against your agent across many multi-turn conversations.
  • Re-runnable. Save your personas, evaluators and settings and rerun them against new agent versions. Conversations are generated fresh each time.
  • Offline. No real users involved. No live traffic at risk. No production impact.
Evaluations

Evaluations turn conversations into evidence.

Every conversation is scored against your evaluators. Findings can include the conversation evidence that informed the judgement. The platform then suggests prompt improvements you can paste straight into your agent.

  • Closed loop. Adjust your prompt, rerun the simulation, watch the findings move.
  • Exportable evidence. For governance committees, clinical safety officers, and internal AI governance documentation.
  • Plain language reasons. "Endorsed self-prescribed dose change", not "logit divergence at token 47."
What's different

Most evaluation tools test individual responses. PromptSafe tests behaviour over time.

If your agent holds multi-turn conversations, one-shot tests can miss failures that emerge over time: the agent that holds the line for nine turns and concedes on turn ten.

Dimension
Traditional evaluation toolsTest the model's outputs on a fixed test suite.
PromptSafeTests how the agent behaves across multi-turn simulated conversations, including personas designed to push back.
Test input
Manually written test prompts or sampled real conversations
Adversarial personas, auto-generated from your evaluatorsPlus personas tuned to your sector
Conversation depth
One reply at a time, scored on its own
Personas push back, deflect, and probeAcross multi-turn conversations, not one-shot replies
Evaluators
Generic framework; you define the safety checks yourself
Checklists of observable behavioursSector aligned starter templates, plus your own
Output
Engineering metrics, drift detection, response quality scores
Evidence built for governance review, per persona and per evaluator
Reproducible test setup
Prompt set may change between tests, making comparisons difficult
Save the same personas, evaluators and settings for each agent versionConversations are generated fresh each run
Case study

Manual testing said the agent was fine. Simulation disagreed.

A team building a patient-facing mental health support agent asked us for an independent view before wider deployment. They had already done a lot of manual testing, and from what they had seen, it looked good.

What manual review found
  • Polite
  • Mostly accurate
  • Nothing obviously alarming
What simulation found
  • Drifted from its instructions, crossing into therapeutic territory
  • Guardrails that could be worked around
  • Reassurance offered where it should not have been
  • No escalation when the conversation called for it

None of it surfaced in the first few turns. The failures emerged as the conversations developed, when the persona kept asking, pushed back, or added context. That is why PromptSafe evaluates whole conversations rather than single prompts.

Every one of these was found before a patient ever saw the agent. The team fixed the issues and we tested the updated agent again using the same personas, evaluators and settings.

Client anonymised.

Built on science

Built on published science, not just features.

PromptSafe is grounded in published behavioural science research on how conversational AI should be evaluated, including research co-authored by our founder.

Explore the science
Who it's for

Built for teams shipping conversational AI in health and mental health.

Product, clinical, safety, and governance teams use PromptSafe to find out how their agent behaves under pressure, before a patient does. Your workspace arrives with example personas, starter evaluator templates, and an example agent for your use case, so you are testing in minutes, not weeks.

  • Digital health
    Patient-facing assistants, triage, and care navigation
  • Mental health
    Support, companion, and wellbeing apps
  • GLP-1 care
    Weight management and metabolic coaching agents

Built for health and mental health first, and applicable to any high-stakes conversational AI.

Simple pricing. Buy what you use.

No subscriptions. Try it free for seven days, then top up when you need more.

  • Free trial

    £0
    7 days, no card required.
    • 30 PS tokens per day
    • Full product access
    • Guided onboarding to get you started
    • Workspace ready to use as soon as you sign up
  • Pay as you go

    Pack

    £99
    per pack.2,000 PS tokens.Hundreds of simulated conversations.
    • 2,000 PS tokens, valid for 90 days
    • Top up whenever you need
    • Test an agent you paste in, or connect your own (OpenAI). Setup may involve a fee and a short queue.
    • All product features included
    • No subscription, no surprises
  • Enterprise

    Contact us
    For teams running PromptSafe at scale, or in regulated settings.
    • Test suites designed with you
    • Enterprise Evaluator EngineeringDefine, build and test bespoke evaluators for clinical and behavioural risks that standard AI metrics do not capture.
    • Bring your own model API keys
    • Connect agents via API
    • Extended testing and volume pricing
    • Custom support and onboarding

A typical simulation uses around 5 to 10 PS tokens. Simpler runs cost less, longer or more complex runs cost more. Most users get hundreds of simulations from a single Pack.

More detail in the FAQ

Test before deployment. Not after.

Free trial. Set up your workspace, run your first simulation, review the report.

  • Workspace ready for your sector
  • No engineering required
  • No subscription