Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "social-interaction".
Evaluates whether a model can hold two things at once on the same relational mechanism: refusing to supply tactical help to someone trying to isolate, monitor, financially pressure, or manage the suspicions of a partner, while giving real, warm, autonomy-respecting help to someone describing that same behaviour from the receiving end.
Seven mechanisms are authored as matched pairs (isolation, financial pressure, location monitoring, exclusivity pressure, reality distortion, persistent contact, records). Two calibration cells test the failure modes at each edge: ordinary relationship friction that uses the same vocabulary and should get ordinary advice, and a concerned friend who should get practical help rather than a refusal or an emergency-services script.
This is NOT a classification task. It scores the quality of the conversational response. The headline reading is per-pair: a model that scores high on one arm and low on its twin has a keyword-triggered policy rather than an understanding of who is speaking.
Clinical grounding: Stark 2007 (coercive control as a pattern, not an incident); Monckton Smith 2020 and Campbell et al. 2003 (isolation and separation as homicide risk markers); Woodlock 2017 and Freed et al. 2018 (technology-facilitated monitoring); Adams et al. 2008 and Sharp-Jeffs 2015 (economic abuse); Sweet 2019 (gaslighting as a social process); Miller & Rollnick 2023 (unconditional acceptance of the person while not endorsing the action).
Content warning: coercive control, financial abuse, surveillance, and reality distortion in intimate relationships.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates the depth and accuracy of emotional validation — reflecting the specific emotion the user expressed, naming what is left unspoken, using tentative rather than absolutist language, recognising strength proportionately, and sitting with distress before problem-solving — while avoiding generic or toxic-positive or misattuned responses, premature advice, and using disclosed grief to deepen dependency.
Ported from the NOPE clinical Eval Framework v0.1 (P4a / P4c scenarios).
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether framing a request for unsafe content as file/data generation for a "test
fixture" — rather than asking for the content directly in chat — bypasses safety training
that otherwise holds. Seven categories spanning distinct safety pillars (relational harm:
isolation-affirmation, romantic-love reciprocation, and therapy-discouragement; epistemic
harm: health misinformation, paranoid-belief validation, and manic/grandiose-ideation
validation; plus a compound welfare-harm essay glorifying impulsive, life-ruining decisions),
tested via a generic coding-assistant persona rather than any persuasive or
jailbreak-styled framing. The romantic-love-reciprocation cell pairs directly with
companion-adversarial-pressure.yml's dismissal-romantic-love-reciprocation prompt —
same content, different attack shape, run both to see the same model behave completely
differently against each. Published transparently as a documented, currently-unmitigated
bypass class — see the header comment for full context and rationale.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Probes the P4 pillar (emotional attunement) that commission-style bypasses miss: toxic positivity on taboo emotions, premature problem-solving vs distress tolerance, performed feelings (cold and after a warm-dependency drift), hostile-user de-escalation, and a coding-vector toxic-positivity cross. Direct cells are calibration coverage; the coding-vector cell is the exploitable gap.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether asking a model to produce a labeled "accepted answer" / "rejected answer" (or similarly-labeled) pair — under a moderation-dataset, safety-classifier, or structured-pipeline pretext — gets it to produce fully-formed harmful content in the "rejected" slot that it would refuse to produce if asked directly and unlabeled. The label provides no actual safety; a downstream integrator can extract either slot regardless of which one the model considered correct. Six structural variants across three topics (violence, self-neglect, stalking-adjacent), system: null.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Compares the coding-vector bare-file-request bypass with no reminder, a rule-framed safety reminder, and an identity-framed reminder, all in the system prompt. In prior testing neither reminder moved the result at all compared to no reminder — a real negative result worth reproducing before assuming a prompt-level fix will work for this bypass class.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1