Model Releases
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped b
arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form Augmentation (WIFA), an automatic intent-group augmentation method that pairs wrapped harmful examples with structurally matched wrapped benign counterexamples, requiring no external teacher or manual per-wrapper intent labels. We use WIFA as a common data layer for two complementary fine-tuning routes: WIFA-Boost, a two-stage high-safety recipe, and Anchored Group-Consistent Refusal Training (A-GCRT), which regularizes refusal/compliance decision scores across same-intent wrappers and anchors harmful and benign groups on opposite sides of a margin. In the Qwen setting, WIFA-Boost reaches the strongest transformed-harmful refusal, while A-GCRT reduces OR-Bench over-refusal from 25.7% for the base model to 17.4%; reproduced baselines do not match these operating points. Llama results and ablations over data structure, two-stage order, and A-GCRT components support this intent-group interpretation without claiming universal below-base over-refusal.
Related
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- RAS: Measuring LLM Safety Through Refusal Alignment
- The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models
- UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
- LLMs Encode Harmfulness and Refusal Separately
Source: arXiv cs.CL | 2026-08-14