Research
ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection
arXiv:2604.16749v1 Announce Type: cross Abstract: Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-w
arXiv:2604.16749v1 Announce Type: cross Abstract: Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-wild deepfakes. We introduce a novel extbf{I}n-extbf{C}ontext extbf{L}earning paradigm with comparison-guidance for extbf{A}udio extbf{D}eepfake detection (extbf{ICLAD}). The framework enables the use of audio language models (ALMs) for training-free generalization to unseen deepfakes and provides textual rationales on the detection outcome. At the core of ICLAD is a pairwise comparative reasoning strategy that guides the ALM to discover and filter hallucinations and deepfake-irrelevant acoustic attributes. The ALM works alongside a specialized deepfake detector, whereby a routing mechanism feeds out-of-distribution samples to the ALM. On in-the-wild datasets, ICLAD improves macro F1 over the specialized detector, with up to 2imes relative improvement. Further analysis demonstrates the flexibility of ICLAD and its potential for deployment on recent open-source ALMs.
Related
- Quantum Vision Theory Applied to Audio Classification for Deepfake Speech Detection
- Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages
- UCS: Estimating Unseen Coverage for Improved In-Context Learning
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
Source: arXiv cs.CL | 2026-04-21