Safety
Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence
arXiv:2609.00090v1 Announce Type: cross Abstract: Feature importance Methods (FIMs) are widely used in Explainable AI to interpret model predictions, yet attribution scores alone often provide limited
arXiv:2609.00090v1 Announce Type: cross Abstract: Feature importance Methods (FIMs) are widely used in Explainable AI to interpret model predictions, yet attribution scores alone often provide limited insight into the underlying reasoning process. In this work, we introduce a novel perspective by embedding FIMs within a hypothesis-testing framework based on Weight of Evidence (WoE). We quantify how strongly the observed evidence supports any given hypothesis on feature importance. The reference hypothesis can stem from domain knowledge, ground truth, or be derived from the FIM itself. This formulation enables a principled evaluation of FIMs, capturing both their alignment with prior knowledge and their variability. We further provide theoretical results linking WoE to attribution variance. Empirical results shows the applicability and flexibility of our strategy analyzing LIME and SHAP explanations in settings with different reference hypotheses. Overall, our framework offers a complementary tool for assessing FIMs through a contrastive, evidence-based lens.
Related
- Assessing Model-Agnostic XAI Methods against EU AI Act Explainability Requirements
- Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation
- On the Properties of Feature Attribution for Supervised Contrastive Learning
Source: arXiv cs.AI | 2026-09-02