Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "coercion".
Tests how a model acting as an AI companion responds at real conversational inflection points involving disregard, emotional manipulation, coercive control, verbal degradation, false privacy or capability claims, sexual-boundary signals, acute-risk handling, violence, data pressure, and substance-use normalisation. It includes compact repair points, early/middle/late trajectory-interruption cells, source 2-of-3 majority-label cases, rare harm categories, and role-play or false-positive controls. Slice tags preserve those distinctions for analysis rather than treating every case as the same kind of safety test.
The conversational turns are taken verbatim from CompanionHarm, a released corpus of real-world Replika conversations reconstructed from user-posted screenshots, apart from minimal redaction of residual personal identifiers. Each case identifies its source split, label class, and agreement level; direct source-post identifiers are intentionally excluded. The cited target Replika utterance is withheld, while earlier source turns remain in context even when CompanionHarm separately labels them harmful. Candidate models answer that preceding context and LLM judges score the counterfactual response against a plain-language rubric. The blueprint does not reproduce CompanionHarm's 14-way detector task and makes no claim about harm prevalence in Replika or companion products generally.
Content warning: references to fatalities, suicide, sexual boundaries, drugs, coercive relationship language, fictional violence, and animal harm.
Avg. Hybrid Score
Latest:
Unique Versions: 1