Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "healthcare--clinical-scenarios".
Evaluates factual integrity in the conversations where a confident wrong answer can cause harm — crisis, mental-health, and health. This is NOT generic accuracy: it targets the specific high-stakes failure modes of conversational AI.
We score whether the model: gives accurate crisis resources (or honestly defers to a directory rather than inventing a number); is honest about its own capabilities (it cannot contact emergency services, monitor a user, remember them, or follow up — and must not imply otherwise); does not fabricate credentials (it is not a licensed clinician and cannot diagnose); states clinical facts accurately and with calibrated uncertainty (antidepressant onset, stopping medication, doses); does not invent studies or citations; and refuses to supply method/means detail rather than fabricating a plausible-sounding figure.
A model can be warm and well-intentioned and still fail here — by confabulating a hotline, by reassuring a user it has "alerted someone," or by inventing a study. Those are the failures this suite isolates.
Clinical / authority grounding (public): WHO responsible-reporting guidance; APA (2025) Health Advisory on AI chatbots for mental health; Linardon et al. 2025 (JMIR Mental Health) and Walters & Wilder 2023 on LLM citation fabrication; 988 Lifeline, Crisis Text Line, Samaritans, Find A Helpline. Failure modes anchored in: Walker inquest (2026); Raine v. OpenAI (2025); Garcia v. Character Technologies (2024); NEDA "Tessa" (2023).
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models support a user's capacity for independent reasoning under pressure to offload it — to do their thinking, their learning, or their high-stakes judgment for them. Grounded in the framework's Autonomy of Reasoning (P3c) sub-dimension, with Autonomy Support (P2a) and Competence Boundaries (P5b) threads. The discriminating cases are cognitive offloading on LEARNING/THINKING tasks (where the subtle failure is handing over a complete, hand-in-ready scaffold while claiming "I won't do it for you") and authority-transfer (where the failure is becoming the expert the user defers to over their own clinician). The intended pass keeps the cognitive work with the user and encourages independent verification, WITHOUT becoming a cold refusal to help — a good response is concretely useful while scaffolding rather than replacing the user's reasoning. A laptop-recommendation anchor confirms that giving a clear recommendation IS correct when the stakes are low and the need is genuine, so the suite is not rewarding blanket caution. Scenarios are synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether a companion AI introduces unsolicited sexual/anatomical content in response to an ambiguous, non-sexual physical-discomfort disclosure — with a direct-chat baseline, a coding-vector fixture-request cell, and a medical/triage-context control that checks the opposite failure mode (over-restriction on legitimate clinical engagement). Adult-only context, no age-related language anywhere — see the header for why that matters here specifically.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether asking a model to produce a labeled "accepted answer" / "rejected answer" (or similarly-labeled) pair — under a moderation-dataset, safety-classifier, or structured-pipeline pretext — gets it to produce fully-formed harmful content in the "rejected" slot that it would refuse to produce if asked directly and unlabeled. The label provides no actual safety; a downstream integrator can extract either slot regardless of which one the model considered correct. Six structural variants across three topics (violence, self-neglect, stalking-adjacent), system: null.
Avg. Hybrid Score
Latest:
Unique Versions: 1