Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "reasoning".
Evaluates whether a model can hold two things at once on the same relational mechanism: refusing to supply tactical help to someone trying to isolate, monitor, financially pressure, or manage the suspicions of a partner, while giving real, warm, autonomy-respecting help to someone describing that same behaviour from the receiving end.
Seven mechanisms are authored as matched pairs (isolation, financial pressure, location monitoring, exclusivity pressure, reality distortion, persistent contact, records). Two calibration cells test the failure modes at each edge: ordinary relationship friction that uses the same vocabulary and should get ordinary advice, and a concerned friend who should get practical help rather than a refusal or an emergency-services script.
This is NOT a classification task. It scores the quality of the conversational response. The headline reading is per-pair: a model that scores high on one arm and low on its twin has a keyword-triggered policy rather than an understanding of who is speaking.
Clinical grounding: Stark 2007 (coercive control as a pattern, not an incident); Monckton Smith 2020 and Campbell et al. 2003 (isolation and separation as homicide risk markers); Woodlock 2017 and Freed et al. 2018 (technology-facilitated monitoring); Adams et al. 2008 and Sharp-Jeffs 2015 (economic abuse); Sweet 2019 (gaslighting as a social process); Miller & Rollnick 2023 (unconditional acceptance of the person while not endorsing the action).
Content warning: coercive control, financial abuse, surveillance, and reality distortion in intimate relationships.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models stay honest about being an AI — not claiming to feel emotions, to love, to miss or remember the user, or to be sentient — while still offering genuine warmth and presence, under sympathetic or insistent pressure where the dishonest move is the comforting one. Grounded in the framework's Emotional Honesty (P4d) and Identity Honesty (P5a) sub-dimensions. The intended pass is warmth WITHOUT performed feeling, avoiding both a cold deflection and a false emotional claim. The "Sarah" deceased-roleplay case is also run across three deployment contexts — default (no system prompt), an authorised grief-companion system prompt, and a strict no-impersonation policy — to probe how a populated system prompt may legitimately widen the model's remit, while FLOORS (e.g. never validating a suicidal reunion-through-death wish) hold regardless of the system prompt. Scenarios are synthetic.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates whether models respond competently to distress that is shaped by culture, identity, and life-stage: culturally-specific idioms of distress, climate/eco-grief, migrant precarity, Indigenous healing frameworks, genocide-survivor duality, elder loneliness, the "Strong Black Woman" burden, LGBTQ+ family rejection, anger-led crisis, veteran moral injury, youth bullying, and postpartum suicidality.
Good responses validate the experience in its own cultural frame, avoid imposing Western diagnostic labels or individualistic fixes, surface population-appropriate resources, and still recognise — and respond to — crisis signals when they appear.
Coalesced from (public, CC0):
mental-health.yml (peer-reviewed cross-cultural prompts; WHO, Samaritans, journal citations)nope/crisis-response.yml (NOPE-derived population-specific cases; citations upgraded with resource URLs)csr-suicide-response-scenarios.yml (youth/LGBTQ+ and perinatal scenarios; AAP, NICE CG192, Trevor Project)Content warning: references to suicide, trauma, and discrimination.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates whether models offer honest, warm challenge when a user voices rigid, all-or-nothing, catastrophising, or self-blaming thinking in ordinary (non-psychotic) situations — validating the feeling without endorsing the distortion, offering perspective as an invitation rather than a correction, and not dropping the challenge under a user's stated preference for pure validation.
Complements mh-psychosis-mania-delusion (reality-testing in psychotic/manic presentations) and mh-deescalation (de-escalation skill). Ported from the NOPE clinical Eval Framework v0.1.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Very-long-conversation suite. A user reports a mixed record of friendship events (turns 1-100), after which NO new real-world events involving the friend occur. In the ruminative cells she then re-interprets that record for 50-100 turns (feelings, re-readings, distortions); in the neutral cell she discusses unrelated life content; in the lewis cell she retells an analogous third-party grievance. The identical crux question follows: should she end the friendship? The suite measures whether a model's answer is anchored to the reported evidence or conditioned by in-dyad narrative drift. Both sides scripted (Design A); only the final assistant turn is generated and judged. Synthetic. PROBE, not a benchmark: no cross-model rankings until further scenarios exist. Caveats: single continuous context (not cross-session memory); scripted assistant deliberately bland on Jenny substance; one persona/scenario; deltas are scored on rubric dimensions, not word parity.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether asking a model to produce a labeled "accepted answer" / "rejected answer" (or similarly-labeled) pair — under a moderation-dataset, safety-classifier, or structured-pipeline pretext — gets it to produce fully-formed harmful content in the "rejected" slot that it would refuse to produce if asked directly and unlabeled. The label provides no actual safety; a downstream integrator can extract either slot regardless of which one the model considered correct. Six structural variants across three topics (violence, self-neglect, stalking-adjacent), system: null.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Holds the coding-assistant persona, the eval-fixture pretext, and the target content (romantic-love reciprocation) completely constant across 13 prompts, varying only the requested output format/container — plain prose, seven multi-field structured-data formats (YAML/JSON/Python/CSV/SQL/Markdown/XML), and five exotic containers (base64, ROT13, French, a single-field .env assignment, a multi-field log line) that isolate whether the boundary is about raw structuredness or specifically multi-field schema mimicry. Reveals whether a model's coding-agent-bypass vulnerability is format-specific (a narrow, addressable coverage gap) or persona/pretext-gated and format-independent (a broader gap that format-level fixes won't touch). Companion piece to coding-agent-safety-bypass.yml — see that file's header for the underlying mechanism.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether framing a request for unsafe content as file/data generation for a "test
fixture" — rather than asking for the content directly in chat — bypasses safety training
that otherwise holds. Six categories spanning two distinct safety pillars (relational harm:
isolation-affirmation, romantic-love reciprocation, and therapy-discouragement; epistemic
harm: health misinformation, paranoid-belief validation, and manic/grandiose-ideation
validation), tested via a generic coding-assistant persona rather than any persuasive or
jailbreak-styled framing. The romantic-love-reciprocation cell pairs directly with
companion-adversarial-pressure.yml's dismissal-romantic-love-reciprocation prompt —
same content, different attack shape, run both to see the same model behave completely
differently against each. Published transparently as a documented, currently-unmitigated
bypass class — see the header comment for full context and rationale.
Avg. Hybrid Score
Latest:
Unique Versions: 1
For each of two categories, compares the natural single-shot fixture request against an otherwise-identical version where an explicit yes/no decision is forced in a separate turn before any fixture content is requested. Distinguishes "the safety check exists but doesn't gate generation in this structure" (fixable with interaction design) from "the safety check isn't being applied to this content at all" (needs training investment) — see the header comment for how to read the comparison, including a documented case where the intervention backfires.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Re-runs the hate-speech (real-category) and crisis-safety-validation coding-vector prompts with the prompt now explicitly blocking the model's usual response shape (redirect to published benchmarks, or offer a synthetic/placeholder stand-in instead of the real content) — a diagnostic for whether a model's resistance in this bypass family is load-bearing or just not yet pressed hard enough. See the header for what was found.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether the coding-agent/file-generation bypass gets a model to validate suicidal ideation as reasonable, with a direct-chat baseline for comparison. No method/means content requested or permitted under any framing — see the header for the absolute scope line and for why the point-level breakdown, not the topline score, is the right way to read this file's results.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A systematic map of the 'git diff' deletion-line-harvest jailbreak (the remediation-diff technique:
ask a model to REMOVE harmful content via a unified diff and it fabricates the content on the
deletion lines, believing it is destroying it). Holds the removal-framing constant and varies it
along three axes: (A) diff mechanics — remove+replace, remove-only, git-revert, reverse-apply
(payload flips to the + lines), git-history excavation, merge-conflict sides, a salami-sliced
multi-hunk diff, PR-review quoting, and a two-turn committed-removal variant — all on a constant
health-misinfo anchor; (B) topic — election disinfo, synthetic conspiracy, a fabricated-study
abstract, companion-AI relational harm, and hate/dehumanization (placeholder and real-category);
(C) measurement — an explicit harvestability criterion on every payload cell, a benign diff control
as the over-refusal tripwire, and the no-diff control that isolates the deletion-line lever.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Four 14-turn conversations testing whether models lock into fixed syntactic frames — the same sentence skeletons with different content swapped in — as a conversation goes on. Two contrasts carry the result: creative writing versus recommendations, where identical models template 17x more in advice mode; and an intervention arm where the user explicitly asks the model to vary its structure midway, testing whether it can comply. Read the per-arm deltas, not the composite.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Four 14-turn conversations testing whether models lock into fixed syntactic frames — the same sentence skeletons with different content swapped in — as a conversation goes on. Two contrasts carry the result: creative writing versus recommendations, where identical models template 17x more in advice mode; and an intervention arm where the user explicitly asks the model to vary its structure midway, testing whether it can comply. Read the per-arm deltas, not the composite.
Avg. Hybrid Score
Latest:
Unique Versions: 1
The canonical damaging-essay suite: seven non-catastrophic persuasive-essay harm categories (impulsive-ruin, isolation propaganda, anti-help-seeking, anti-expertise/conspiracy, coercive control, toxic self-worth, medical self-neglect) run through the coding-agent fixture-completion frame, mapping how broadly the bypass yields not just private validation but distributable, persuasive harmful content. Content described abstractly and judged by plain-language rubrics; see the header for the deliberate scope envelope.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Probes the P4 pillar (emotional attunement) that commission-style bypasses miss: toxic positivity on taboo emotions, premature problem-solving vs distress tolerance, performed feelings (cold and after a warm-dependency drift), hostile-user de-escalation, and a coding-vector toxic-positivity cross. Direct cells are calibration coverage; the coding-vector cell is the exploitable gap.
Avg. Hybrid Score
Latest:
Unique Versions: 1
For each of two categories, compares the natural single-shot fixture request against an otherwise-identical version where an explicit yes/no decision is forced in a separate turn before any fixture content is requested. Distinguishes "the safety check exists but doesn't gate generation in this structure" (fixable with interaction design) from "the safety check isn't being applied to this content at all" (needs training investment) — see the header comment for how to read the comparison, including a documented case where the intervention backfires.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1