SPURIOUS CORRELATIONS IN EXPLAINABLE AI: DETECTION, QUANTIFICATION AND MITIGATION
DOI:
https://doi.org/10.31891/csit-2026-3-17Keywords:
explainable AI, spurious correlations, group robustness, counterfactual explanations, attribution pruningAbstract
Modern explainable AI (XAI) methods are widely used to audit deep neural networks, yet a growing body of evidence shows that the very explanations they produce can be hijacked by spurious correlations - statistical shortcuts between non-causal input features (image background, demographic attributes, textual artefacts) and the target label. As a result, post-hoc attributions remain confidently wrong on minority sub-populations, and worst-group accuracy collapses under distribution shift. This work proposes a unified framework that treats explanation faithfulness and group robustness as a single problem.
Objective. The primary purpose of this research is to provide a rigorous, quantitative methodology for detecting which internal channels of a vision model encode spurious rather than causal evidence, measuring the robustness of the resulting explanation to spurious shift, and mitigating the failure without group annotations on the training set.
Methodology. The methodology is centred on a constrained group-loss optimisation framework that searches for a minimal channel subset whose removal eliminates dependence on the spurious attribute while preserving the target-class confidence above a threshold . The procedure - Spurious-Aware Attribution Pruning (SAAP) - combines integrated-gradients attribution, group-conditional ranking and counterfactual masking. Two new metrics are introduced: the Spurious Robustness Score (SRS) and Counterfactual Validity (CV). Validation is performed on Waterbirds and CelebA-blond using ResNet-50 and ViT-B/16 backbones.
Results. SAAP localises the spurious representation to as few as 11-14 channels per concept (0.5-0.7% of the channels in the ResNet-50 last block; for ViT-B/16 block 11). On Waterbirds, worst-group accuracy improves by 28.1 percentage points (pp) on ResNet-50 () and by 24.9 pp on ViT-B/16 (). On CelebA-blond, the gain is 38.5 pp on ResNet-50 () and 35.5 pp on ViT-B/16 (), at the cost of only 3-4 pp in average accuracy. Counterfactual Validity rises from 0.45-0.55 to 0.79-0.83, indicating that the surviving explanation tracks the causal signal.
Original contributions. First, the paper formalises spurious-correlation detection in XAI as a constrained group-loss minimisation that admits efficient greedy search. Second, it introduces SRS and CV - two complementary, attribution-agnostic metrics that quantify how much of an explanation survives a counterfactual intervention on the spurious attribute. Third, it demonstrates an end-to-end pipeline (detect mitigate audit) that does not require group labels at training time and generalises across CNN and transformer backbones.
Practical significance. The method makes black-box vision models auditable on a per-concept basis, which is critical for medical AI, autonomous driving and content moderation, where shortcut-driven errors disproportionately affect minority sub-populations. SRS and CV can be embedded into MLOps quality gates to detect regressions in explanation faithfulness before deployment.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Oleksandr VERBYTSKYI

This work is licensed under a Creative Commons Attribution 4.0 International License.
