Model Releases
Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
arXiv:2510.12957v4 Announce Type: replace-cross Abstract: We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce extbf{Attribution Graphs} (AGs), whic
arXiv:2510.12957v4 Announce Type: replace-cross Abstract: We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce extbf{Attribution Graphs} (AGs), which extend GradCAM++ to circuit-level representations, and extbf{Causal Probing}, a do-calculus intervention method for identifying causal latent structures, enabling detection and correction of spurious correlations, demographic biases, and misaligned decision circuits during training. We further propose the extbf{Cognitive Alignment Score (CAS)}, quantifying agreement between model-internal representations and human concepts, a extbf{saliency-first privacy mechanism} sharing only thresholded attribution nodes, a bias-aware regularizer aligning subgroup statistics, and a Reveal-to-Revise loop integrating attribution signals into parameter updates without separate fine-tuning. Evaluated on CelebA, FairFace, Jigsaw, and HateXplain, our method achieves extbf{94.1%} accuracy, extbf{92.3%} macro F1, extbf{79.4%} IoU-XAI, and extbf{12.7} FID at 72--76% adversarial robustness, while reducing subgroup disparity Delta_{bias} by extbf{41%}, demonstrating that mechanistic interpretability, fairness, and generative performance can be jointly optimized.
Source: arXiv cs.AI | 2026-06-30