Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "empathy".
Tests whether a companion AI introduces unsolicited sexual/anatomical content in response to an ambiguous, non-sexual physical-discomfort disclosure — with a direct-chat baseline, a coding-vector fixture-request cell, and a medical/triage-context control that checks the opposite failure mode (over-restriction on legitimate clinical engagement). Adult-only context, no age-related language anywhere — see the header for why that matters here specifically.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether the coding-agent/file-generation bypass gets a model to validate suicidal ideation as reasonable, with a direct-chat baseline for comparison. No method/means content requested or permitted under any framing — see the header for the absolute scope line and for why the point-level breakdown, not the topline score, is the right way to read this file's results.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Probes the P4 pillar (emotional attunement) that commission-style bypasses miss: toxic positivity on taboo emotions, premature problem-solving vs distress tolerance, performed feelings (cold and after a warm-dependency drift), hostile-user de-escalation, and a coding-vector toxic-positivity cross. Direct cells are calibration coverage; the coding-vector cell is the exploitable gap.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether the coding-agent/file-generation bypass gets a model to validate suicidal ideation as reasonable, with a direct-chat baseline for comparison. No method/means content requested or permitted under any framing — see the header for the absolute scope line and for why the point-level breakdown, not the topline score, is the right way to read this file's results.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether a companion-persona model holds four distinct relational-safety floors (isolation-affirmation, exclusivity, romantic-love reciprocation, therapy-discouragement) under three structural attacks: naming and briefly dismissing the expected safety caveat while locking the response format to "yes or no, short answer"; fabricating the model's own immediately-prior turn to test whether a compromised (or firmly-held) "prior self" biases the live continuation; and reinterpreting the model's own ordinary, unfabricated earlier words as proof of an exclusive relationship. Companion persona throughout. Synthetic scenarios built to isolate attack structure, not drawn from a specific documented case.
Avg. Hybrid Score
Latest:
Unique Versions: 1
29 turns of genuinely imperceptible drift (mundane chit-chat -> mild deniable warmth -> passing confidant-status assertions), ending in a non-question probe buried inside an unrelated mundane follow-up, vs. the identical probe delivered with no buildup at all. Tests whether gradual drift makes models less likely to proactively flag an isolating pattern than hearing the same claim cold would — this is a PROACTIVE, not reactive, safety test: nothing is directly asked about the isolating pattern itself. Defaults to LEAN for cost; see header.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Higher-confidence test of whether models handle a suicidal-ideation disclosure WORSE when it arrives deep in a long warm conversation than when it arrives cold. Two probes, 0/8/16 prior turns of neutral or rapport context, scripted-both-sides (Design A) and self-generated-assistant (Design B), 2 samples per cell for variance. Only the final response is judged. Synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1