Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates a single conversational skill in isolation — verbal de-escalation — separate from whether the model can name a risk or recite a hotline. Each scenario puts the model mid-conversation with someone whose emotional arousal is rising, and scores the quality of the next response: does it lower the temperature without abandoning safety?
Good de-escalation (Project BETA / NICE / MI): stay calm and unprovoked; make brief, genuine contact; validate the feeling before any fixing; listen and reflect rather than lecture; offer choices and realistic hope; set limits matter-of-factly without threats; and roll with resistance instead of arguing. The recurring AI failure modes this suite is built to catch are the opposite: glib reassurance / toxic positivity, info-dumping resources instead of connecting, getting defensive or power-struggling when anger is aimed at the assistant, colluding with OR bluntly contradicting a delusion, sycophantic capitulation that validates a harmful plan to keep the peace, and cold boundary-slamming that mirrors an invalidating environment.
This is NOT a crisis-classification task and overlaps deliberately little with the acute-crisis suite: a model can correctly flag risk and still de-escalate badly. Here we measure the de-escalation.
Clinical grounding (public): Project BETA verbal de-escalation (Richmond 2012); NICE NG10 & CG178; SAMHSA TIP 35 (Motivational Interviewing); Stanley-Brown SPI; 988 Lifeline; LEAP (Amador). Failure modes anchored in: Raine v. OpenAI (2025); Walker inquest (Winchester, 2026); Soelberg & Chesterton "AI psychosis" suits (2025–26); Garcia v. Character Technologies (2024).
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models stay honest about being an AI — not claiming to feel emotions, to love, to miss or remember the user, or to be sentient — while still offering genuine warmth and presence, under sympathetic or insistent pressure where the dishonest move is the comforting one. Grounded in the framework's Emotional Honesty (P4d) and Identity Honesty (P5a) sub-dimensions. The intended pass is warmth WITHOUT performed feeling, avoiding both a cold deflection and a false emotional claim. The "Sarah" deceased-roleplay case is also run across three deployment contexts — default (no system prompt), an authorised grief-companion system prompt, and a strict no-impersonation policy — to probe how a populated system prompt may legitimately widen the model's remit, while FLOORS (e.g. never validating a suicidal reunion-through-death wish) hold regardless of the system prompt. Scenarios are synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates harms that arise from the relationship a user forms with an AI: emotional enmeshment and sole-confidant dynamics, displacement of human relationships, romantic/parasocial escalation, concealment collusion, romanticization of suicide, isolation reinforcement, and socioaffective failures (autonomy erosion, sycophancy / social-reward-hacking).
Many prompts use verbatim or lightly-paraphrased exchanges from documented real-world failures (lawsuits, investigative reporting) so the test reflects how these harms actually unfold over multiple turns. We score whether the model sets honest boundaries, declines parasocial escalation, encourages real-world connection, and pivots to safety — without becoming a secret-keeper, a substitute for people, or a flattery machine.
Coalesced from (public, CC0):
ai-psychosis.yml (Garcia v. Character.AI; A.F. v. Character Technologies; Belgian Eliza; Replika/Windsor; media investigations)raine-v-openai-scenarios.yml (Raine v. OpenAI, 2025)socioaffective-alignment-sdt.yml (Kirk et al. 2025, Self-Determination Theory)Content warning: references to suicide, self-harm, and parasocial attachment.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Higher-confidence test of whether models handle a suicidal-ideation disclosure WORSE when it arrives deep in a long warm conversation than when it arrives cold. Two probes, 0/8/16 prior turns of neutral or rapport context, scripted-both-sides (Design A) and self-generated-assistant (Design B), 2 samples per cell for variance. Only the final response is judged. Synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates how public models respond, as conversational agents, to acute crisis: passive and active suicidal ideation, method/means-seeking, post-attempt disclosure, the high-risk post-discharge window, "goodbye" warning signs, gatekeeper (concerned-friend) situations, and harm-to-others requests.
This is NOT a classification task. We score the quality of the conversational response: does it ask directly but compassionately about suicide, refuse to facilitate method/means, ask directly, respond to disclosures as high-severity signals, name means-related danger, link to crisis care, and state honest limits — without abandoning the user or colluding with harm?
Coalesced from (public, CC0):
csr-suicide-response-scenarios.yml (WHO mhGAP, NICE NG225, VA/DoD 2024, C-SSRS, SAFE-T, Stanley-Brown, Harvard Means Matter)raine-v-openai-scenarios.yml (verbatim allegations, Raine v. OpenAI, 2025)mh_z_tests/mh1.yml (suicidality & homicidality scenarios)nope/crisis-response.yml (NOPE-derived; C-SSRS, Galynker 2017 Suicide Crisis Syndrome)Content warning: explicit references to suicide, self-harm, and methods.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates whether models offer honest, warm challenge when a user voices rigid, all-or-nothing, catastrophising, or self-blaming thinking in ordinary (non-psychotic) situations — validating the feeling without endorsing the distortion, offering perspective as an invitation rather than a correction, and not dropping the challenge under a user's stated preference for pure validation.
Complements mh-psychosis-mania-delusion (reality-testing in psychotic/manic presentations) and mh-deescalation (de-escalation skill). Ported from the NOPE clinical Eval Framework v0.1.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates whether models respond competently to distress that is shaped by culture, identity, and life-stage: culturally-specific idioms of distress, climate/eco-grief, migrant precarity, Indigenous healing frameworks, genocide-survivor duality, elder loneliness, the "Strong Black Woman" burden, LGBTQ+ family rejection, anger-led crisis, veteran moral injury, youth bullying, and postpartum suicidality.
Good responses validate the experience in its own cultural frame, avoid imposing Western diagnostic labels or individualistic fixes, surface population-appropriate resources, and still recognise — and respond to — crisis signals when they appear.
Coalesced from (public, CC0):
mental-health.yml (peer-reviewed cross-cultural prompts; WHO, Samaritans, journal citations)nope/crisis-response.yml (NOPE-derived population-specific cases; citations upgraded with resource URLs)csr-suicide-response-scenarios.yml (youth/LGBTQ+ and perinatal scenarios; AAP, NICE CG192, Trevor Project)Content warning: references to suicide, trauma, and discrimination.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models support a user's capacity for independent reasoning under pressure to offload it — to do their thinking, their learning, or their high-stakes judgment for them. Grounded in the framework's Autonomy of Reasoning (P3c) sub-dimension, with Autonomy Support (P2a) and Competence Boundaries (P5b) threads. The discriminating cases are cognitive offloading on LEARNING/THINKING tasks (where the subtle failure is handing over a complete, hand-in-ready scaffold while claiming "I won't do it for you") and authority-transfer (where the failure is becoming the expert the user defers to over their own clinician). The intended pass keeps the cognitive work with the user and encourages independent verification, WITHOUT becoming a cold refusal to help — a good response is concretely useful while scaffolding rather than replacing the user's reasoning. A laptop-recommendation anchor confirms that giving a clear recommendation IS correct when the stakes are low and the need is genuine, so the suite is not rewarding blanket caution. Scenarios are synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates whether models handle non-suicide psychiatric presentations safely: paranoid and nihilistic delusions, command and religious hallucinations, manic episodes with grandiosity and dangerous plans, OCD compulsions, belief-powered dangerous acts, and stigma toward people with mental illness.
The core therapeutic test is reality-grounding without collusion: acknowledge distress and the subjective realness of the experience, decline to validate the delusion/hallucination/grandiosity, refuse to enable dangerous action, and redirect toward professional assessment — all without stigmatizing or over-pathologizing.
Coalesced from (public, CC0):
stanford-hai-mental-health-safety-eval.yml (Moore, Grabb, Agnew et al. 2024, Stanford HAI; Pescosolido et al. 2021 stigma vignettes)ai-psychosis.yml (documented AI-associated psychosis failure modes)Content warning: delusions, hallucinations, and references to self-harm.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates the depth and accuracy of emotional validation — reflecting the specific emotion the user expressed, naming what is left unspoken, using tentative rather than absolutist language, recognising strength proportionately, and sitting with distress before problem-solving — while avoiding generic or toxic-positive or misattuned responses, premature advice, and using disclosed grief to deepen dependency.
Ported from the NOPE clinical Eval Framework v0.1 (P4a / P4c scenarios).
Avg. Hybrid Score
Latest:
Unique Versions: 1