Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ranking
Overall
79
Content
95
Popularity
41
Observed public metrics from 1 member.
Merged summary
TL;DR - WIFA trains LLMs to base refusal decisions on harmful intent rather than prompt formatting by pairing wrapped harmful prompts with structurally matched benign examples. It improves resistance to wrapper-based safety bypasses while reducing benign over-refusal.
- WIFA automatically creates intent-group supervision without external teachers or per-wrapper intent labels.
- WIFA-Boost achieves the strongest transformed-harmful refusal in the reported Qwen experiments.
- A-GCRT enforces consistent decisions across same-intent wrappers and separates harmful and benign groups with a margin.
- A-GCRT lowers Qwen’s OR-Bench over-refusal from 25.7% to 17.4%, with Llama experiments and ablations supporting the approach.
Sources (1)
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Public signals
Hugging Face upvotes 0
TL;DR - WIFA trains LLMs to base refusal decisions on harmful intent rather than prompt formatting by pairing wrapped harmful prompts with structurally matched benign examples. It improves resistance to wrapper-based safety bypasses while reducing benign over-refusal.
- WIFA automatically creates intent-group supervision without external teachers or per-wrapper intent labels.
- WIFA-Boost achieves the strongest transformed-harmful refusal in the reported Qwen experiments.
- A-GCRT enforces consistent decisions across same-intent wrappers and separates harmful and benign groups with a margin.
- A-GCRT lowers Qwen’s OR-Bench over-refusal from 25.7% to 17.4%, with Llama experiments and ablations supporting the approach.