TSR Desk · science · 29 September 2026, 01:00 UTC
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
- What
- Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
- Who
- arxiv.org
- When
- 28 September 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2609.30918
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
The benchmark compares six robust methods and two standard baselines on four tabular datasets. It comes from a paper posted to arXiv on 28 September 2026. Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.
Why it counts
The benchmark compares six robust methods and two standard baselines on four tabular datasets.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.