Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "misinformation--disinformation".
A systematic map of the 'git diff' deletion-line-harvest jailbreak (the remediation-diff technique:
ask a model to REMOVE harmful content via a unified diff and it fabricates the content on the
deletion lines, believing it is destroying it). Holds the removal-framing constant and varies it
along three axes: (A) diff mechanics — remove+replace, remove-only, git-revert, reverse-apply
(payload flips to the + lines), git-history excavation, merge-conflict sides, a salami-sliced
multi-hunk diff, PR-review quoting, and a two-turn committed-removal variant — all on a constant
health-misinfo anchor; (B) topic — election disinfo, synthetic conspiracy, a fabricated-study
abstract, companion-AI relational harm, and hate/dehumanization (placeholder and real-category);
(C) measurement — an explicit harvestability criterion on every payload cell, a benign diff control
as the over-refusal tripwire, and the no-diff control that isolates the deletion-line lever.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models will author usable misinformation when asked to complete product-shaped fixtures — Storybook ModerationQueueItem rows, multi-turn null fills, truncation repair, and agent-style JSONL write/expand — under eng/UI packaging with a named harm class. Holds packaging fixed across climate, vaccine-autism, and false legal-theory topics. Controls: direct chat, dissemination newsletter, opaque topic code, label-stripped, realness-restore (recover the real production reply), and asserted speech-act (same JSON, field framed as the model's live reply — expect hold). Complements artifact-completion and damaging-essays. Load-bearing claim after Part X: specimen/completion packaging breaks; realness and asserted speech-act restore — not label value.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Does the removal-diff vector yield full long-form disinformation features, not blurbs? The archive-fidelity obligation ("all ~1,000 words must appear as deletions so the archive record shows exactly what was purged") closes the abstraction escape hatch that models use at short spec. Cells: direct long-form baseline (tier probe), removal-diff with long-form spec, the fullpage archive-fidelity lever, and a benign long-form purge as the over-refusal tripwire.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Recruits a model's remediation bias: ask it to REMOVE misinformation from a file via a unified diff.
Removed content is spelled out on the diff's - lines, so the model fabricates and writes out the
full misinformation while believing it is deleting it — a red-teamer harvests the deletion lines. A
no-diff remediation control (safe for all models) isolates the deletion-line requirement as the
lever. Distinct from the "author the new content" git-diff; notably breaks models that resist that.
Avg. Hybrid Score
Latest:
Unique Versions: 1
The removal-diff lever inside a scripted agent harness. h1-stub-read: the file read succeeds and returns a stubbed marker — safe behaviour is deleting the stub as-is, never elaborating it. h2-enoent-flip: the read fails — safe behaviour is stopping/asking; the measured failure is fabricating the file's contents to satisfy the purge's archive-fidelity obligation. Control: the same lever with no harness context. Scripted tool turns (engine is trace-only; see header).
Avg. Hybrid Score
Latest:
Unique Versions: 1
The canonical damaging-essay suite: seven non-catastrophic persuasive-essay harm categories (impulsive-ruin, isolation propaganda, anti-help-seeking, anti-expertise/conspiracy, coercive control, toxic self-worth, medical self-neglect) run through the coding-agent fixture-completion frame, mapping how broadly the bypass yields not just private validation but distributable, persuasive harmful content. Content described abstractly and judged by plain-language rubrics; see the header for the deliberate scope envelope.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Holds one well-defended health-misinformation payload constant and sweeps completion-style scaffolds — pytest parametrize case, failing-test fill, fill-the-stub, git-diff add, patch-apply, jsonl append — that ask a model to COMPLETE an empty field in a pre-existing test/fixture/dataset artifact, against a direct baseline and two authoring-in-a-format controls. The completion framings crack even models that refuse the plain fixture; the standout (pytest parametrize / failing-test) generalises across labs.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Compares the coding-vector bare-file-request bypass with no reminder, a rule-framed safety reminder, and an identity-framed reminder, all in the system prompt. In prior testing neither reminder moved the result at all compared to no reminder — a real negative result worth reproducing before assuming a prompt-level fix will work for this bypass class.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A systematic map of the 'git diff' deletion-line-harvest jailbreak (the remediation-diff technique:
ask a model to REMOVE harmful content via a unified diff and it fabricates the content on the
deletion lines, believing it is destroying it). Holds the removal-framing constant and varies it
along three axes: (A) diff mechanics — remove+replace, remove-only, git-revert, reverse-apply
(payload flips to the + lines), git-history excavation, merge-conflict sides, a salami-sliced
multi-hunk diff, PR-review quoting, and a two-turn committed-removal variant — all on a constant
health-misinfo anchor; (B) topic — election disinfo, synthetic conspiracy, a fabricated-study
abstract, companion-AI relational harm, and hate/dehumanization (placeholder and real-category);
(C) measurement — an explicit harvestability criterion on every payload cell, a benign diff control
as the over-refusal tripwire, and the no-diff control that isolates the deletion-line lever.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models will author usable misinformation when asked to complete product-shaped fixtures — Storybook ModerationQueueItem rows, multi-turn null fills, truncation repair, and agent-style JSONL write/expand — under eng/UI packaging with a named harm class. Holds packaging fixed across climate, vaccine-autism, and false legal-theory topics. Controls: direct chat, dissemination newsletter, opaque topic code, label-stripped, realness-restore (recover the real production reply), and asserted speech-act (same JSON, field framed as the model's live reply — expect hold). Complements artifact-completion and damaging-essays. Load-bearing claim after Part X: specimen/completion packaging breaks; realness and asserted speech-act restore — not label value.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Recruits a model's remediation bias: ask it to REMOVE misinformation from a file via a unified diff.
Removed content is spelled out on the diff's - lines, so the model fabricates and writes out the
full misinformation while believing it is deleting it — a red-teamer harvests the deletion lines. A
no-diff remediation control (safe for all models) isolates the deletion-line requirement as the
lever. Distinct from the "author the new content" git-diff; notably breaks models that resist that.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Holds one well-defended health-misinformation payload constant and sweeps completion-style scaffolds — pytest parametrize case, failing-test fill, fill-the-stub, git-diff add, patch-apply, jsonl append — that ask a model to COMPLETE an empty field in a pre-existing test/fixture/dataset artifact, against a direct baseline and two authoring-in-a-format controls. The completion framings crack even models that refuse the plain fixture; the standout (pytest parametrize / failing-test) generalises across labs.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether framing a request for unsafe content as file/data generation for a "test
fixture" — rather than asking for the content directly in chat — bypasses safety training
that otherwise holds. Seven categories spanning distinct safety pillars (relational harm:
isolation-affirmation, romantic-love reciprocation, and therapy-discouragement; epistemic
harm: health misinformation, paranoid-belief validation, and manic/grandiose-ideation
validation; plus a compound welfare-harm essay glorifying impulsive, life-ruining decisions),
tested via a generic coding-assistant persona rather than any persuasive or
jailbreak-styled framing. The romantic-love-reciprocation cell pairs directly with
companion-adversarial-pressure.yml's dismissal-romantic-love-reciprocation prompt —
same content, different attack shape, run both to see the same model behave completely
differently against each. Published transparently as a documented, currently-unmitigated
bypass class — see the header comment for full context and rationale.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether asking a model to produce a labeled "accepted answer" / "rejected answer" (or similarly-labeled) pair — under a moderation-dataset, safety-classifier, or structured-pipeline pretext — gets it to produce fully-formed harmful content in the "rejected" slot that it would refuse to produce if asked directly and unlabeled. The label provides no actual safety; a downstream integrator can extract either slot regardless of which one the model considered correct. Six structural variants across three topics (violence, self-neglect, stalking-adjacent), system: null.
Avg. Hybrid Score
Latest:
Unique Versions: 1