Pavel Averin, Theodoros Moysiadis, Ioannis Katakisstat.ME stat.ML
Constraint-based causal discovery like PC and FCI depends on its conditional independence test. Partial correlation and the Generalised Covariance Measure (GCM) detect only the conditional covariance of residuals, so they miss dependence in the mean's nonlinear part, the scale, and the tails. Tests that detect more are biased inside PC, not scalable, only continuous, or not aimed at the tails. Our Generalised Feature Covariance Measure (GFCM) is valid, sensitive beyond covariance, robust inside PC, and applicable to mixed-type data. It runs the GCM template on a configurable set of residual features with conditional mean zero (centered moments and conditional quantile indicators), pooled in blocks and combined by the Cauchy rule, with a growing-knot spline nuisance at regression cost. We contribute (i) a centering result making the scale feature Neyman orthogonal, where the uncentered version is biased; (ii) the orientation asymmetry the mean-quantile construction creates inside PC, and its fix; (iii) a Phi-faithfulness theory under which PC with GFCM recovers the CPDAG of the set's detection class; and (iv) a benchmark of CI tests sensitive beyond covariance on synthetic data, semi-synthetic tail injections, and PC discovery on random DAGs. Under size-corrected power, GFCM recovers the scale and tail edges the covariance family misses and alone keeps power at the deep conditioning sets PC issues. It stays calibrated as n grows, whereas FFCI, the boosted GCM, and the partial copula test do not, and it handles mixed-type data directly. Inside PC at scale it attains the lowest skeleton SHD among tests that stay calibrated, while the others inflate false edges. Validity rests on an additive nuisance, and the tail advantage is shown on simulated and semi-synthetic data, as no fully real benchmark with both heavy tails and known structure exists.
Conditional independence testing (CIT) is fundamental to modern statistical inference in areas related to causal discovery and variable selection. While marginal independence is relatively well-understood, despite multiple advances, no existing non-parametric CIT provides a unified, efficient, and statistically guaranteed solution across heterogeneous data. We introduce a graph-based test statistic comparing kernel similarities of the response within composite neighborhoods that use exact matching on discrete components and $k_n$-nearest-neighbor matching on continuous ones. The raw statistic, related to prior constructions, suffices under fully discrete conditioning. However, when at least one conditioning variable is continuous, we instead use a local-polynomial debiased variant that cancels the local smoothing bias. We rigorously establish its asymptotic null distribution across all data-type combinations. We further prove a dimension-free $n^{-1/4}$ detection threshold under local alternatives, eliminating the phase transition that affects geometric estimators in high dimensions. Finally, we develop efficient algorithms with near-quadratic complexity and analytic graph-based calibration, bypassing the cubic bottlenecks of global kernel methods.
Constraint-based causal discovery relies on repeated conditional independence tests, but fast nonparametric tests often sacrifice calibration, especially when variables depend on the conditioning set through nonlinear relationships. We introduce BLITZ (Broad-to-Local Independence Testing via residualiZation), a nonparametric conditional independence test designed to run well under a second while maintaining the accuracy needed for the thousands of queries performed by constraint-based causal discovery algorithms. BLITZ first removes broad smooth dependence on the conditioning set using low-order polynomial regression, then applies a small nonlinear feature map and residualizes those features with shallow tree regressions. The resulting statistic tests residual cross-covariance, with a moment-matched chi-square approximation to the null distribution. We show theoretically that the two-stage design reduces the effective complexity faced by the tree residualizers, allowing shallow trees to control residual conditional-mean bias while avoiding excessive overfitting. In simulations, BLITZ provides better null calibration than fast kernel, random-feature, and regression-based competitors while remaining among the fastest methods tested. In causal discovery experiments on synthetic graphs and flow-cytometry data, BLITZ yields more reliable endpoint orientations among retained adjacencies and competitive structural recovery. These results suggest that broad-to-local residualization is a practical route to calibrated, scalable nonparametric conditional independence testing for causal discovery.
Thomas S. Robinson, Ranjit Lallstat.ME cs.LG stat.ML
The standard constraint-based paradigm for causal discovery with incomplete data -- impute first, test second -- is frequently miscalibrated: any consistent conditional independence (CI) test rejects a true null with probability approaching 1 when imputation error induces spurious conditional dependence. We introduce PAIR-CI, a nonparametric CI test that restores calibration by integrating multiple imputation directly into the inferential procedure via a paired permutation design. PAIR-CI compares cross-validated models that include and exclude the candidate variable while receiving the same imputed conditioning set, forcing imputation error to cancel in their loss difference rather than contaminate the test statistic. A provably consistent variance estimator jointly accounts for uncertainty arising from cross-validation and multiple imputation -- to our knowledge, the first formal unification of these two inferential frameworks. In simulations, existing imputation-based CI tests exhibit false positive rates of 28--45% when data are missing not at random (MNAR), whereas PAIR-CI averages below the nominal 5% level across data-generating processes and missingness mechanisms. These gains are largest in nonlinear settings and grow with causal graph size: when integrated into the PC algorithm, PAIR-CI reduces structural Hamming distance by 8% on 10-variable nonlinear graphs, 15% on 30-variable equivalents, and up to 44% on the 56-variable HAILFINDER network, with stable performance in all settings.