Observational datasets frequently contain many baseline variables, yet investigators estimating causal effects may not know which variables to include in the adjustment set. Confounding information may also be distributed weakly across many variables. Propensity scores can simplify adjustment by reducing high-dimensional covariates to a scalar with binary treatment. Although the propensity score is the coarsest balancing score, this distributional optimality does not imply maximal specificity over causal graphs. We instead examine all causal graphs among a candidate score, treatment, and outcome while allowing latent variables. Under faithfulness, we identify the largest set of unconditional and conditional dependence relations whose truth is invariant to whether treatment causes the outcome, leaving treatment-effect estimation to the downstream analysis. This criterion defines the maximally specific graph class expressible through these relations. We then develop the proposed algorithm, which operationalizes the criterion through a generalized eigenvalue problem whose score space targets the span of a balancing coordinate and an outcome-guided coordinate. We show that sufficiently informative proxies can recover this span without direct observation of the adjustment variables, characterize the resulting estimation and causal errors, and establish bootstrap validity for the complete procedure. Simulations and a real-data application demonstrate superior performance over several alternatives.
Seong Woo Ahn, Alessandro Leite, José Lucas De Melo Costa +3cs.LG stat.ML
Local causal discovery is a scalable alternative to global structure learning. However, it can struggle to identify valid adjustment sets in data-scarce settings because of finite-sample uncertainty, incomplete local neighborhoods, and unresolved Markov equivalence. Although many application domains provide structured background knowledge, its integration into local causal discovery remains limited. We propose b-LOAD, a knowledge-informed extension of the LOAD algorithm for local discovery of optimal adjustment sets. b-LOAD incorporates prior edge constraints directly into the local structure-learning procedure and uses Meek's rules to expand the discovery frontier dynamically, yielding a knowledge-constrained partially directed graph over the relevant local subgraph. This strategy helps prevent structurally relevant nodes introduced by prior knowledge from being excluded by local search. We prove that, under sound background knowledge, the procedure monotonically refines the admissible equivalence class and can enlarge the set of identifiable causal queries, enabling recovery of optimal adjustment sets that are not identifiable from observational conditional-independence information alone. Empirically, b-LOAD improves downstream causal effect estimation relative to purely data-driven and standard knowledge-augmented baselines, particularly in data-scarce and structurally complex regimes. Results on real-world biological networks show that locally targeted prior knowledge provides the largest gains and remains beneficial under moderate structural noise. These findings position b-LOAD as a scalable approach for converting fragmented domain knowledge into more reliable causal-effect estimation.