What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
TL;DR - This paper finds that tested compliance detectors exhibit “rule blindness”: their verdicts often depend on scenario cues rather than the governing rule. It introduces benchmarks and counterfactual tests to expose this weakness, showing that step-by-step reasoning performs better than fast detectors.
- Deleting, permuting, or replacing rules did not change accuracy across tested guard models and activation probes.
- A crossed-rule benchmark confirms that neither the rule nor scenario alone should predict compliance labels.
- The proposed training-free Internal Compliance Score matched a bag-of-words baseline and failed its pre-registered criterion.
- ICS can improve response ranking cheaply, but its gains disappear under an adaptive white-box attack.