🛰️ Daily AI Frontier
‹ back to 2026-08-16

Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

arXiv cs.CL LLMs & Foundation Models Ping Wu, Haibo Tong, Feifei Zhao, Han Shen, Yu Shi, Yilin Zhao, Sicheng Shen, Guobin Shen, Yun Luo, Yi Zeng 2026-08-13
Representative image for Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

TL;DR - WIFA trains LLMs to base refusal decisions on harmful intent rather than prompt formatting by pairing wrapped harmful prompts with structurally matched benign examples. It improves resistance to wrapper-based safety bypasses while reducing benign over-refusal.

  • WIFA automatically creates intent-group supervision without external teachers or per-wrapper intent labels.
  • WIFA-Boost achieves the strongest transformed-harmful refusal in the reported Qwen experiments.
  • A-GCRT enforces consistent decisions across same-intent wrappers and separates harmful and benign groups with a margin.
  • A-GCRT lowers Qwen’s OR-Bench over-refusal from 25.7% to 17.4%, with Llama experiments and ablations supporting the approach.

view merged work →