TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 29 September 2026, 01:00 UTC

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

What
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Who
arxiv.org
When
28 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.30918
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

The benchmark compares six robust methods and two standard baselines on four tabular datasets. It comes from a paper posted to arXiv on 28 September 2026. Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.

Why it counts

The benchmark compares six robust methods and two standard baselines on four tabular datasets.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.