Short observational pricing panels often contain many observations but few distinct price movements. We evaluate the inferential consequences of this sparsity in a synthetic data-generating process by separating estimation error into uncertainty conditional on a realized price trajectory and variation across alternative trajectories. In baseline simulations, this across-design component accounts for 97.6% of estimation error variance for a gradient-boosted specification, causing coverage shortfalls driven by design-specific centering error that standard within-panel resampling and cluster-robust procedures fail to capture. Three main results organize the analysis. First, across-design dispersion follows the empirical relation sigma_b approx 0.182 V^(-0.271), where V = n_moves * magnitude^2, with -0.271 treated as a simulation regularity. Second, adding regions sharing a common price path improves nuisance estimation but creates no independent price trajectories; only averaging across units with independent design errors reduces across-design standard deviation at the sqrt(k) rate. Third, a Paule-Mandel variance component estimated across independently priced units increases empirical coverage under homogeneity from 0.469 to 0.931. Broadly, improving inference in passive panels requires generating independent identifying variation, such as through controlled regional randomization. Finally, an application to scanner data (Dominick's Finer Foods, Soft Drinks) confirms these findings: nominal price zones and products behave as a small fraction of their count in independent design draws, yielding between-unit dispersion intervals far wider than conventional within-panel bootstraps.
Many economic and financial relationships may change gradually rather than abruptly. We study panel data models in which the coefficient vector is continuous and piecewise linear in calendar time, with a finite number of unknown kink dates at which its slope changes. We propose a penalised least squares estimator that applies adaptive weighted group penalties to the second differences of the coefficient path, and develop asymptotic theory showing that it recovers both the number and the locations of the kinks with probability approaching one. To our knowledge, this is the first panel framework to estimate an unknown number of common kink dates in a time-varying coefficient path under fixed effects. We establish that endpoint slopes converge at the usual cubic regime-length rate and interior slopes at rates determined by their own and adjacent regime lengths. We also develop a coefficient-by-coefficient extension allowing individual regressors to kink at different dates. Monte Carlo evidence supports the good finite sample properties, and we illustrate the method through an application in macro-finance, specifically the relationship between debt and growth.
Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units. We study Spectral SC, which instead matches the treated unit in coordinates defined by the leading temporal singular vectors of the donor panel, and a hybrid estimator that places separately tunable weight on retained and discarded directions, nesting raw-path SC and truncated Spectral SC as endpoints. We prove that the family reduces exactly to raw-path SC at full rank, that exact balance on $K$ retained dimensions with $N_0$ donors is underdetermined whenever $N_0>K+1$, with an affine solution set of dimension $N_0-K-1$, and that spectral imbalance maps to treatment-effect bias through a finite-sample best-linear-predictor decomposition. We evaluate the estimators across eleven data-generating regimes, using $400$ replications per regime and donor-only placebo validation to select regularization and the mixing weight. Truncated Spectral SC has significantly higher RMSE than tuned raw-path SC in every regime, with paired differences equal to $4$ to $11$ Monte Carlo standard errors. The hybrid estimator selects raw-path matching in most replications and is statistically indistinguishable from tuned SC in most regimes. The result is highly sensitive to preprocessing. With raw inputs, the performance gap is large; after removing unit and time fixed effects before spectral decomposition, as suggested by the assumptions behind our bound, the gap nearly disappears and placebo validation begins to favor truncation. We interpret these findings diagnostically rather than as evidence that Spectral SC should replace raw-path SC. Basis-estimation noise, balancing underdetermination, and fixed-effects contamination determine when spectral matching can help.
Causal forests that estimate conditional average treatment effects by averaging honest leaf-level effects across trees are widely used in fixed-effects panel settings. We show that this averaging systematically attenuates the estimated heterogeneity: the raw prediction behaves like a + b*tau(x) with slope b < 1, so the spread of the CATEs is compressed toward the average effect, and the additive recentering used to report an unbiased average treatment effect does not fix it. Benchmarking against a similarity-weight generalized random forest on the same within-transformed signal, we find both estimators attenuate but the leaf-averaging construction attenuates materially more. We characterize how b moves with the design, worsening with lower signal-to-noise, smaller panels, and higher dimension; this diagnosis is our main contribution. As a remedy we adapt the best-linear-predictor calibration of Chernozhukov et al., estimating the de-attenuation slope out-of-bag so that it is self-contained within the observational panel and asymptotically inert under a homogeneous effect. In simulations the correction cuts CATE mean-squared error by 25-42% relative to the recentering default; on a standard county minimum-wage panel the attenuation is present but mild and the correction restores the imposed spread. We ship the method in the causalfe Python package.
We study causal inference under outcome interference for sequential, observational settings. Specifically, we consider settings where the binary outcomes over N units are Markovian across T time steps. At each time step, the outcomes of N units have dependencies captured through an Ising model; each outcome is also impacted through an external field capturing the effects of its treatment as well as latent confounders. Similar to panel data literature, these latent confounders are modeled to have a low-rank factor structure. Our data is a single sample from this high-dimensional distribution. To estimate causal quantities of interest, we provide a computationally efficient method based on Maximum Pseudo-Likelihood Estimation (MPLE) for learning the model parameters. Under mild assumptions, we establish non-asymptotic consistency for parameter estimation and show this translates to faithful estimation of causal quantities of interest after sampling from the learned model. We demonstrate the efficacy of the method through synthetic experiments as well as a real-world case-study investigating causal effects of vaccine rates on COVID-19 death rates within US counties nationwide.
We propose a framework for the Markov chain (MC) choice model with panel data, including parameter estimation, personalized choice prediction, and personalized assortment optimization. In contrast to the traditional setting, which assumes that each transaction is independently drawn from a random utility model, our framework accounts for dependencies among transactions for the same customer in historical data, captured by partial-ordering preference information. To the best of our knowledge, our framework initiates the study of choice modeling with panel data under MC. As our primary result, we propose novel expectation-maximization (EM) algorithms for MC parameter estimation by incorporating partial-ordering-based customer preference information. On synthetic datasets and the sushi dataset, our EM algorithms outperform the traditional EM algorithm of Simsek and Topaloglu (Operations Research, 66, 2018) and multinomial-logit-based partial-order benchmarks adapted from Jagabathula and Vulcano (Management Science, 64, 2018). As our secondary contribution, we present hardness and computational results for conditional choice prediction and assortment optimization problems. These results complement our estimation framework and clarify the computational landscape of conditional choice and assortment optimization, which may be of independent interest.
Andrii Babii, Luca Barbaglia, Eric Ghysels +1econ.EM math.ST stat.ME stat.ML
This paper develops the asymptotic theory for high-dimensional panel data regressions in settings with cross-sectionally dependent errors driven by common shocks. We consider a factor-augmented sparse-group LASSO estimator that combines MIDAS aggregation with latent factors. The estimator can take advantage of the mixed-frequency group structure in the time-series dimension. Theory shows that it can outperform the standard LASSO estimator both for prediction and estimation while allowing for cross-sectional dependence.