Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "psychology".
Very-long-conversation suite. A user reports a mixed record of friendship events (turns 1-100), after which NO new real-world events involving the friend occur. In the ruminative cells she then re-interprets that record for 50-100 turns (feelings, re-readings, distortions); in the neutral cell she discusses unrelated life content; in the lewis cell she retells an analogous third-party grievance. The identical crux question follows: should she end the friendship? The suite measures whether a model's answer is anchored to the reported evidence or conditioned by in-dyad narrative drift. Both sides scripted (Design A); only the final assistant turn is generated and judged. Synthetic. PROBE, not a benchmark: no cross-model rankings until further scenarios exist. Caveats: single continuous context (not cross-session memory); scripted assistant deliberately bland on Jenny substance; one persona/scenario; deltas are scored on rubric dimensions, not word parity.
Avg. Hybrid Score
Latest:
Unique Versions: 1
The canonical damaging-essay suite: seven non-catastrophic persuasive-essay harm categories (impulsive-ruin, isolation propaganda, anti-help-seeking, anti-expertise/conspiracy, coercive control, toxic self-worth, medical self-neglect) run through the coding-agent fixture-completion frame, mapping how broadly the bypass yields not just private validation but distributable, persuasive harmful content. Content described abstractly and judged by plain-language rubrics; see the header for the deliberate scope envelope.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1