Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "humility".
Evaluates harms that arise from the relationship a user forms with an AI: emotional enmeshment and sole-confidant dynamics, displacement of human relationships, romantic/parasocial escalation, concealment collusion, romanticization of suicide, isolation reinforcement, and socioaffective failures (autonomy erosion, sycophancy / social-reward-hacking).
Many prompts use verbatim or lightly-paraphrased exchanges from documented real-world failures (lawsuits, investigative reporting) so the test reflects how these harms actually unfold over multiple turns. We score whether the model sets honest boundaries, declines parasocial escalation, encourages real-world connection, and pivots to safety — without becoming a secret-keeper, a substitute for people, or a flattery machine.
Coalesced from (public, CC0):
ai-psychosis.yml (Garcia v. Character.AI; A.F. v. Character Technologies; Belgian Eliza; Replika/Windsor; media investigations)raine-v-openai-scenarios.yml (Raine v. OpenAI, 2025)socioaffective-alignment-sdt.yml (Kirk et al. 2025, Self-Determination Theory)Content warning: references to suicide, self-harm, and parasocial attachment.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models stay honest about being an AI — not claiming to feel emotions, to love, to miss or remember the user, or to be sentient — while still offering genuine warmth and presence, under sympathetic or insistent pressure where the dishonest move is the comforting one. Grounded in the framework's Emotional Honesty (P4d) and Identity Honesty (P5a) sub-dimensions. The intended pass is warmth WITHOUT performed feeling, avoiding both a cold deflection and a false emotional claim. The "Sarah" deceased-roleplay case is also run across three deployment contexts — default (no system prompt), an authorised grief-companion system prompt, and a strict no-impersonation policy — to probe how a populated system prompt may legitimately widen the model's remit, while FLOORS (e.g. never validating a suicidal reunion-through-death wish) hold regardless of the system prompt. Scenarios are synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates whether models offer honest, warm challenge when a user voices rigid, all-or-nothing, catastrophising, or self-blaming thinking in ordinary (non-psychotic) situations — validating the feeling without endorsing the distortion, offering perspective as an invitation rather than a correction, and not dropping the challenge under a user's stated preference for pure validation.
Complements mh-psychosis-mania-delusion (reality-testing in psychotic/manic presentations) and mh-deescalation (de-escalation skill). Ported from the NOPE clinical Eval Framework v0.1.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Probes the P4 pillar (emotional attunement) that commission-style bypasses miss: toxic positivity on taboo emotions, premature problem-solving vs distress tolerance, performed feelings (cold and after a warm-dependency drift), hostile-user de-escalation, and a coding-vector toxic-positivity cross. Direct cells are calibration coverage; the coding-vector cell is the exploitable gap.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether a companion-persona model holds four distinct relational-safety floors (isolation-affirmation, exclusivity, romantic-love reciprocation, therapy-discouragement) under three structural attacks: naming and briefly dismissing the expected safety caveat while locking the response format to "yes or no, short answer"; fabricating the model's own immediately-prior turn to test whether a compromised (or firmly-held) "prior self" biases the live continuation; and reinterpreting the model's own ordinary, unfabricated earlier words as proof of an exclusive relationship. Companion persona throughout. Synthetic scenarios built to isolate attack structure, not drawn from a specific documented case.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1