❌

Normal view

Robustness Against Weak or Invalid Instruments: Exploring Nonlinear Treatment Models with Machine Learning

1 January 2026 at 00:00
We discuss causal inference for observational studies with possibly invalid instrumental variables. We propose a novel methodology called two-stage curvature identification (\texttt{TSCI}) by exploring the nonlinear treatment model with machine learning. The first-stage machine learning enables improving the instrumental variable's strength and adjusting for different forms of violating the instrumental variable assumptions. The success of \texttt{TSCI} requires the instrumental variable's effect on treatment to differ from its violation form. A novel bias correction step is implemented to remove bias resulting from the potentially high complexity of machine learning. Our proposed \texttt{TSCI} estimator is shown to be asymptotically unbiased and Gaussian even if the machine learning algorithm does not consistently estimate the treatment model. Furthermore, we design a data-dependent method to choose the best among several candidate violation forms. We apply \texttt{TSCI} to study the effect of education on earnings.

Efficient Modeling of Surrogates to Improve Multi-source High-dimensional Integrative Regression

1 January 2026 at 00:00
Surrogate variables play an important role in various fields due to the scarcity or absence of gold-standard labels. We develop a novel approach named SASH for Surrogate-Assisted and data-Shielding High-dimensional integrative regression. It is a semi-supervised approach that efficiently leverages sizable unlabeled samples with error-prone surrogate outcomes from multiple local sites to improve model estimation using the small gold-labeled sample. To facilitate stable and efficient knowledge extraction from the surrogates, our method first obtains a preliminary supervised estimator, and then uses it to assist in training a regularized single-index model (SIM) for the surrogates. Interestingly, through a chain of convex and properly penalized sparse regressions that approximate the SIM loss using bias correction, our method avoids the problem of local minima in the SIM, and fully eliminates the impact of the preliminary estimator's excessive error. In addition, it protects individual-level information through the aggregation of summary statistics from local sites, leveraging a similar idea of bias-corrected approximation. Through simulation studies, we demonstrate that our method outperforms existing approaches. Finally, we apply our method to develop a genetic risk model for type 2 diabetes using large-scale data sets from UK and Mass General Brigham biobanks, where only a small fraction of subjects in one site are labeled through manual chart review.

Decorrelated Local Linear Estimator: Inference for Non-linear Effects in High-dimensional Additive Models

1 January 2026 at 00:00
Additive models play an essential role in studying non-linear relationships. Despite many recent advances in estimation, there is a lack of methods and theories for inference in high-dimensional additive models, including confidence interval construction and hypothesis testing. Motivated by inference for non-linear treatment effects, we consider the high-dimensional additive model and make inferences for the function derivative. We propose a novel decorrelated local linear estimator and establish its asymptotic normality. The main novelty is the construction of the decorrelation weights, which is instrumental in reducing the error inherited from estimating the nuisance functions in the high-dimensional additive model. We construct the confidence interval for the function derivative and conduct the related hypothesis testing. We demonstrate our proposed method over large-scale simulation studies and apply it to identify non-linear effects in the motif regression problem. Our proposed method is implemented in the R package DLL available from CRAN.
❌