Research
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
arXiv:2608.12623v1 Announce Type: new Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing wh
arXiv:2608.12623v1 Announce Type: new Abstract: Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information but can ask the classifier for a label plus a short rationale or quoted evidence. We introduce Groundedness Drift, a lightweight score measuring whether the answer summary remains grounded in the input. Across two 7B backbones, five datasets, and four common non-adaptive OpenBackdoor-style attack families, Groundedness Drift achieves higher AUROC and lower residual target ASR than every compared detector in all cases at a nominal 5% clean-FPR budget. We then evaluate Unsupported Groundedness, a multi-probe escalation for explanation-camouflage stress cases. Unsupported Groundedness improves signals but does not close the adaptive gap.
Related
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
- AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
- Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models
Source: arXiv cs.CL | 2026-08-14