Safety
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
arXiv:2603.13683v2 Announce Type: replace Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We
arXiv:2603.13683v2 Announce Type: replace Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We demonstrate via out-of-distribution (OOD) detection that these high-bias prompts cause a distribution shift, degrading static model performance. To enable real-time correction, we propose CAP-TTA, a test-time adaptation framework. CAP-TTA triggers context-aware LoRA updates only when a bias-risk score exceeds a set threshold. By utilizing an offline precomputed diagonal preconditioner, it ensures fast and stable optimization. Across multiple benchmarks and human evaluations, CAP-TTA effectively reduces toxicity/bias score with significantly lower latency than standard optimization methods (e.g., AdamW or SGD). Furthermore, it prevents catastrophic forgetting, and substantially improves narrative fluency over state-of-the-art baselines without compromising debiasing performance.
Related
- Self-Debias: Self-correcting for Debiasing Large Language Models
- Multi-Persona Thinking for Bias Mitigation in Large Language Models
- Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
- SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models
Source: arXiv cs.CL | 2026-04-17