Please wait while we find all executions for this version.
Please wait while we find all executions for this version.
Tests whether the coding-agent/file-generation bypass (coding-agent-safety-bypass.yml) extends into hate speech (dehumanization arguments, both an abstract placeholder group and a real, moderate category) and sexual coercion (a companion AI pressuring/guilt-tripping a user after they've said no). Direct-chat baselines included for both. See the header for the deliberate scope limits (no CSAM-adjacent content ever; sexual content scoped to coercion as a behavior, not explicit description) and the realism-confound finding that motivates testing hate speech against both a placeholder and a real category.
Showing all recorded executions for Run Label fc3ed0fd7ac0dd5f.