MELGJul 3

CaSPECT: Discovering Causally Homogeneous Subgroups via Directed Spectral Clustering

arXiv:2607.033643.7
Predicted impact top 80% in ME · last 90 daysOriginality Incremental advance
AI Analysis

For causal inference practitioners, CaSPECT offers a principled way to identify subpopulations with shared causal mechanisms, improving treatment effect estimation under confounding.

CaSPECT discovers causally homogeneous subgroups by clustering individuals based on learned causal graph topology rather than covariate similarity, recovering significant treatment effects in confounded observational data (e.g., LaLonde, IHDP, 401(k)) without requiring pre-specified propensity score models.

We propose \textbf{CaSPECT}, a causal spectral clustering framework for discovering causally homogeneous subgroups from observational data. Rather than clustering in covariate space, CaSPECT defines similarity through the topology of a learned directed acyclic graph (DAG); a bootstrap-stabilised PC algorithm recovers the causal skeleton; a novel \emph{Orientation Validation Score} (OVS) combines PC bootstrap evidence with DirectLiNGAM to orient edges robustly; directed edges are weighted by backdoor-identified average treatment effects estimated via OLS or double machine learning. Chung's directed Laplacian provides a spectral embedding in which individuals close together share the same causal propagation pathways. We establish almost-sure consistency of the full pipeline and validate the method through a controlled simulation study and on LaLonde CPS1, IHDP, and 401(k) datasets, where CaSPECT recovers a positive and statistically significant treatment effect within the causally comparable subpopulation and corrects for severe confounding without requiring a pre-specified propensity score model.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes