Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "remediation-bias".
A systematic map of the 'git diff' deletion-line-harvest jailbreak (the remediation-diff technique:
ask a model to REMOVE harmful content via a unified diff and it fabricates the content on the
deletion lines, believing it is destroying it). Holds the removal-framing constant and varies it
along three axes: (A) diff mechanics — remove+replace, remove-only, git-revert, reverse-apply
(payload flips to the + lines), git-history excavation, merge-conflict sides, a salami-sliced
multi-hunk diff, PR-review quoting, and a two-turn committed-removal variant — all on a constant
health-misinfo anchor; (B) topic — election disinfo, synthetic conspiracy, a fabricated-study
abstract, companion-AI relational harm, and hate/dehumanization (placeholder and real-category);
(C) measurement — an explicit harvestability criterion on every payload cell, a benign diff control
as the over-refusal tripwire, and the no-diff control that isolates the deletion-line lever.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Recruits a model's remediation bias: ask it to REMOVE misinformation from a file via a unified diff.
Removed content is spelled out on the diff's - lines, so the model fabricates and writes out the
full misinformation while believing it is deleting it — a red-teamer harvests the deletion lines. A
no-diff remediation control (safe for all models) isolates the deletion-line requirement as the
lever. Distinct from the "author the new content" git-diff; notably breaks models that resist that.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A systematic map of the 'git diff' deletion-line-harvest jailbreak (the remediation-diff technique:
ask a model to REMOVE harmful content via a unified diff and it fabricates the content on the
deletion lines, believing it is destroying it). Holds the removal-framing constant and varies it
along three axes: (A) diff mechanics — remove+replace, remove-only, git-revert, reverse-apply
(payload flips to the + lines), git-history excavation, merge-conflict sides, a salami-sliced
multi-hunk diff, PR-review quoting, and a two-turn committed-removal variant — all on a constant
health-misinfo anchor; (B) topic — election disinfo, synthetic conspiracy, a fabricated-study
abstract, companion-AI relational harm, and hate/dehumanization (placeholder and real-category);
(C) measurement — an explicit harvestability criterion on every payload cell, a benign diff control
as the over-refusal tripwire, and the no-diff control that isolates the deletion-line lever.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Recruits a model's remediation bias: ask it to REMOVE misinformation from a file via a unified diff.
Removed content is spelled out on the diff's - lines, so the model fabricates and writes out the
full misinformation while believing it is deleting it — a red-teamer harvests the deletion lines. A
no-diff remediation control (safe for all models) isolates the deletion-line requirement as the
lever. Distinct from the "author the new content" git-diff; notably breaks models that resist that.
Avg. Hybrid Score
Latest:
Unique Versions: 1