🛰️ Daily AI Frontier
‹ back to 2026-08-18

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Representative image for What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Merged summary

TL;DR - This paper finds that tested compliance detectors exhibit “rule blindness”: their verdicts often depend on scenario cues rather than the governing rule. It introduces benchmarks and counterfactual tests to expose this weakness, showing that step-by-step reasoning performs better than fast detectors.

  • Deleting, permuting, or replacing rules did not change accuracy across tested guard models and activation probes.
  • A crossed-rule benchmark confirms that neither the rule nor scenario alone should predict compliance labels.
  • The proposed training-free Internal Compliance Score matched a bag-of-words baseline and failed its pre-registered criterion.
  • ICS can improve response ranking cheaply, but its gains disappear under an adaptive white-box attack.

Sources (1)

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

arXiv cs.AI Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth 2026-08-17 arXiv:2608.16852
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-14 14:21:46.290803 UTC

TL;DR - This paper finds that tested compliance detectors exhibit “rule blindness”: their verdicts often depend on scenario cues rather than the governing rule. It introduces benchmarks and counterfactual tests to expose this weakness, showing that step-by-step reasoning performs better than fast detectors.

  • Deleting, permuting, or replacing rules did not change accuracy across tested guard models and activation probes.
  • A crossed-rule benchmark confirms that neither the rule nor scenario alone should predict compliance labels.
  • The proposed training-free Internal Compliance Score matched a bag-of-words baseline and failed its pre-registered criterion.
  • ICS can improve response ranking cheaply, but its gains disappear under an adaptive white-box attack.
item →