🛰️ Daily AI Frontier
‹ back to 2026-08-18

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

arXiv cs.AI LLMs & Foundation Models Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth 2026-08-17
Representative image for What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

TL;DR - This paper finds that tested compliance detectors exhibit “rule blindness”: their verdicts often depend on scenario cues rather than the governing rule. It introduces benchmarks and counterfactual tests to expose this weakness, showing that step-by-step reasoning performs better than fast detectors.

  • Deleting, permuting, or replacing rules did not change accuracy across tested guard models and activation probes.
  • A crossed-rule benchmark confirms that neither the rule nor scenario alone should predict compliance labels.
  • The proposed training-free Internal Compliance Score matched a bag-of-words baseline and failed its pre-registered criterion.
  • ICS can improve response ranking cheaply, but its gains disappear under an adaptive white-box attack.

view merged work →