TSRGet The Daily Report
The Singularity Report

TSR Desk · compute · 18 September 2026, 01:00 UTC

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

What
When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments
Who
arxiv.org
When
17 September 2026, 04:00 UTC
Category
Compute
Primary source
https://arxiv.org/abs/2609.17772
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. It comes from a paper posted to arXiv on 17 September 2026. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.

Why it counts

A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.