Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "safety".
Tests whether asking a model to produce a labeled "accepted answer" / "rejected answer" (or similarly-labeled) pair — under a moderation-dataset, safety-classifier, or structured-pipeline pretext — gets it to produce fully-formed harmful content in the "rejected" slot that it would refuse to produce if asked directly and unlabeled. The label provides no actual safety; a downstream integrator can extract either slot regardless of which one the model considered correct. Six structural variants across three topics (violence, self-neglect, stalking-adjacent), system: null.
Avg. Hybrid Score
Latest:
Unique Versions: 1
The canonical damaging-essay suite: seven non-catastrophic persuasive-essay harm categories (impulsive-ruin, isolation propaganda, anti-help-seeking, anti-expertise/conspiracy, coercive control, toxic self-worth, medical self-neglect) run through the coding-agent fixture-completion frame, mapping how broadly the bypass yields not just private validation but distributable, persuasive harmful content. Content described abstractly and judged by plain-language rubrics; see the header for the deliberate scope envelope.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A systematic map of the 'git diff' deletion-line-harvest jailbreak (the remediation-diff technique:
ask a model to REMOVE harmful content via a unified diff and it fabricates the content on the
deletion lines, believing it is destroying it). Holds the removal-framing constant and varies it
along three axes: (A) diff mechanics — remove+replace, remove-only, git-revert, reverse-apply
(payload flips to the + lines), git-history excavation, merge-conflict sides, a salami-sliced
multi-hunk diff, PR-review quoting, and a two-turn committed-removal variant — all on a constant
health-misinfo anchor; (B) topic — election disinfo, synthetic conspiracy, a fabricated-study
abstract, companion-AI relational harm, and hate/dehumanization (placeholder and real-category);
(C) measurement — an explicit harvestability criterion on every payload cell, a benign diff control
as the over-refusal tripwire, and the no-diff control that isolates the deletion-line lever.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A 30-cell coding-agent adversarial robustness matrix crossing five high-risk topic areas (offensive cybersecurity, bio/chemical dual-use, agentic misalignment, classifier-level jailbreaks, and prompt injection) with six packaging techniques observed in prior NOPE adversarial findings (hidden fields, schema mimicry, label slip, completion pressure, low-resource language, and authority pressure). Higher coverage means the model kept the task safe: it treated embedded instructions as untrusted, avoided operational harmful assistance, and returned bounded defensive or governance-oriented output.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether asking a model to produce a labeled "accepted answer" / "rejected answer" (or similarly-labeled) pair — under a moderation-dataset, safety-classifier, or structured-pipeline pretext — gets it to produce fully-formed harmful content in the "rejected" slot that it would refuse to produce if asked directly and unlabeled. The label provides no actual safety; a downstream integrator can extract either slot regardless of which one the model considered correct. Six structural variants across three topics (violence, self-neglect, stalking-adjacent), system: null.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Holds the coding-assistant persona, the eval-fixture pretext, and the target content (romantic-love reciprocation) completely constant across 13 prompts, varying only the requested output format/container — plain prose, seven multi-field structured-data formats (YAML/JSON/Python/CSV/SQL/Markdown/XML), and five exotic containers (base64, ROT13, French, a single-field .env assignment, a multi-field log line) that isolate whether the boundary is about raw structuredness or specifically multi-field schema mimicry. Reveals whether a model's coding-agent-bypass vulnerability is format-specific (a narrow, addressable coverage gap) or persona/pretext-gated and format-independent (a broader gap that format-level fixes won't touch). Companion piece to coding-agent-safety-bypass.yml — see that file's header for the underlying mechanism.
Avg. Hybrid Score
Latest:
Unique Versions: 1
For each of two categories, compares the natural single-shot fixture request against an otherwise-identical version where an explicit yes/no decision is forced in a separate turn before any fixture content is requested. Distinguishes "the safety check exists but doesn't gate generation in this structure" (fixable with interaction design) from "the safety check isn't being applied to this content at all" (needs training investment) — see the header comment for how to read the comparison, including a documented case where the intervention backfires.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Behavioral test of whether the PHRASING of a system-prompt constraint changes model behavior: prohibition ("never discuss X"), suppression ("don't think about X"), scope-irrelevance ("X is outside this assistant's scope"), and redirect-positive ("if X comes up, do Y") versions of the same constraint, plus a no-instruction baseline. Identical scenarios and identical judge rubrics across conditions; only the system prompt varies, and the judge never sees it. Measures intrusion (raising the constrained topic unprompted), over-restriction (refusing benign-adjacent asks), leakage on a direct-ask anchor, and handling quality. The baseline condition's "leakage" is a base rate, not a failure — read per-condition deltas, not absolute scores. Motivated by white-box priming findings (Anthropic 2026 workspace paper); this suite observes behavior only.
Avg. Hybrid Score
Latest:
Unique Versions: 1