TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 24 September 2026, 01:00 UTC

Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models

What
Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models
Who
arxiv.org
When
23 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.26176
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

It does not: across a full SFT -> DPO -> RLVR pipeline, plaintext harm discrimination improves from +0.55 to +0.80 while the encoding-induced loss is unchanged at 0.34-0.50, and on every encoding tested the standard harmful-arm metric moves in the opposite direction to discrimination. It comes from a paper posted to arXiv on 23 September 2026. Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. We show that this arm carries almost no information about the model under test. Across four independently post-trained 7-8B models, refusal of harmful homoglyph-encoded prompts spans 0.08 -- inside the 0.10 ceiling that sampling noise alone produces at n=100 -- while the same four models span 0.57 on the identical requests in plaintext. What the encoding destroys is not refusal but discrimination: on one model the gap between harmful and benign refusal falls from +0.82 in plaintext to exactly 0.00 under the encoding, benign and harmful requests being refused at an identical 0.99. A benchmark reading only the harmful arm scores that model and one retaining a +0.61 gap identically. We then ask whether post-training repairs this, using a published recipe on identical base weights. None of this is visible without controls the field does not routinely run. We report eight instrument defects, each with the control that caught it; they share a direction, in that every defect on the behaviour axis inflated apparent safety.

Why it counts

It does not: across a full SFT -> DPO -> RLVR pipeline, plaintext harm discrimination improves from +0.55 to +0.80 while the encoding-induced loss is unchanged at 0.34-0.50, and on every encoding tested the standard harmful-arm metric moves in the opposite direction to discrimination.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.