Beyond Propensity Score: The SAHR, Subgroup-Aware Holistic Ridge for Causal Inference in Observational Studies
Abstract
Causal inference is critical for decision-making in healthcare, economics, and policy, but it faces challenges such as confounding, heterogeneous treatment effects, and scalability in high-dimensional settings. Traditional methods like propensity scores often struggle with these complexities, particularly when treatment effects vary across subpopulations or datasets are large. To address these limitations, we propose Subgroup-Aware Holistic Ridge (SAHR), a novel framework that utilize subgroup analysis, kernel approximation, and adaptive covariate balancing based on BNN or regression covariates for robust causal inference. SAHR identifies subgroups with heterogeneous treatment effects using Gaussian Mixture Models (GMM), ensures scalability via the Nyström method, and balances covariates adaptively through a Holistic ridge regression framework. Experiments on benchmark datasets demonstrate that SAHR outperforms state-of-the-art methods in predictive accuracy, subgroup identification, and covariate balancing. By addressing key challenges in observational studies, SAHR advances causal inference methodology, offering an interpretable tool for applications in personalized medicine, and policy evaluation.
Keywords:
Causal inference, Subgroup analysis, Covariate balancing, Gaussian mixture models, BNN, Propensity scores, RegressionReferences
- [1] Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41–55. https://doi.org/10.1093/biomet/70.1.41
- [2] Austin, P. C. (2011). An Introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate behavioral research, 46(3), 399–424. https://doi.org/10.1080/00273171.2011.568786
- [3] Chipman, H. A., George, E. I., & McCulloch, R. E. (2010). Bart: Bayesian additive regression trees. Annals of applied statistics, 4(1), 266–298. https://doi.org/10.1214/09-AOAS285
- [4] Bang, H., & Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics, 61(4), 962–973. https://doi.org/10.1111/j.1541-0420.2005.00377.x
- [5] Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., & Smola, A. (2012). A kernel two-sample test. The journal of machine learning research, 13, 723–773. https://dl.acm.org/doi/10.5555/2188385.2188410
- [6] Reynolds, D. (2009). Gaussian mixture models. In Encyclopedia of biometrics (pp. 659–663). Boston, MA: Springer US. https://doi.org/10.1007/978-0-387-73003-5_196
- [7] Williams, C. K. I., & Seeger, M. (2000). Using the nyström method to speed up kernel machines. Proceedings of the 14th international conference on neural information processing systems (pp. 661–667). Cambridge, MA, USA: MIT Press. https://doi.org/10.5555/3008751.3008847
- [8] Hoerl, A. E., & Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1), 55–67. https://doi.org/10.1080/00401706.1970.10488634
- [9] Tesei, G., Giampanis, S., Shi, J., & Norgeot, B. (2023). Learning end-to-end patient representations through self-supervised covariate balancing for causal treatment effect estimation. Journal of biomedical informatics, 140, 104339. https://doi.org/10.1016/j.jbi.2023.104339
- [10] Shalit, U., Johansson, F. D., & Sontag, D. (2017). Estimating individual treatment effect: Generalization bounds and algorithms. Proceedings of the 34th international conference on machine learning (Vol. 70, pp. 3076–3085). PMLR. https://doi.org/10.5555/3305890.3305999
- [11] Villani, C. 2009). Optimal transport: Old and new (Vol. 338). Springer. https://doi.org/10.1007/978-3-540-71050-9
- [12] Imai, K., & Ratkovic, M. (2014). Covariate balancing propensity score. Journal of the royal statistical society series b: Statistical methodology, 76(1), 243–263. https://doi.org/10.1111/rssb.12027
- [13] Robins, J., Sued, M., Lei-Gomez, Q., & Rotnitzky, A. (2007). Comment: Performance of double-robust estimators when “inverse probability” weights are highly variable. Statistical science, 22(4), 544–559. http://doi.org/10.1214/07-STS227D
- [14] Wager, S., & Athey, S. (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American statistical association, 113(523), 1228–1242. https://doi.org/10.1080/01621459.2017.1319839
- [15] Yoon, J., Jordon, J., & Van Der Schaar, M. (2018). GANITE: Estimation of individualized treatment effects using generative adversarial nets. International conference on learning representations. ICLR 2018. (pp. 3076–3085). https://dblp.org/rec/conf/iclr/YoonJS18.html
- [16] Su, X., Tsai, C. L., Wang, H., Nickerson, D. M., & Li, B. (2009). Subgroup Analysis via Recursive Partitioning. The journal of machine learning research, 10, 141–158. https://doi.org/10.5555/1577069.1577074
- [17] LaLonde, R. J. (1986). Evaluating the econometric evaluations of training programs with experimental data. The American economic review, 76(4), 604–620. http://www.jstor.org/stable/1806062
- [18] Schochet, P. Z. (2010). Is regression adjustment supported by the Neyman model for causal inference? Journal of statistical planning and inference, 140(1), 246–259. https://doi.org/10.1016/j.jspi.2009.07.008
- [19] Belthangady, C., Stedden, W., & Norgeot, B. (2021). Minimizing bias in massive multi-arm observational studies with BCAUS: Balancing covariates automatically using supervision. BMC medical research methodology, 21(1), 190. https://doi.org/10.1186/s12874-021-01383-x
- [20] Pearl, J. (2003). Causality: Models, reasoning, and inference. Econometric theory, 19(675–685), 46. https://doi.org/10+10170S0266466603004109
- [21] Imbens, G. W., & Rubin, D. B. (2015). Causal inference for statistics, social, and biomedical sciences. Cambridge University Press. https://doi.org/10.1017/CBO9781139025751
- [22] Heckman, J. J., & Vytlacil, E. J. (2007). Chapter 70 econometric evaluation of social programs, part I: Causal models, structural models and econometric policy evaluation. In Handbook of econometrics (Vol. 6, pp. 4779–4874). Elsevier. https://doi.org/10.1016/S1573-4412(07)06070-9
- [23] Rahimi, A., & Recht, B. (2007). Random features for large-scale kernel machines. Advances in neural information processing systems (Vol. 20, pp. 1–8). EECS at UC Berkeley. https://doi.org/10.5555/2981562.2981710
- [24] Johansson, F., Shalit, U., & Sontag, D. (2016). Learning representations for counterfactual inference. Proceedings of the 33rd international conference on machine learning (Vol. 48, pp. 3020–3029). New York, New York, USA: PMLR. https://doi.org/10.5555/3045390.3045708
- [25] Crump, R. K., Hotz, V. J., Imbens, G. W., & Mitnik, O. A. (2008). Nonparametric tests for treatment effect heterogeneity. The review of economics and statistics, 90(3), 389–405. https://doi.org/10.1162/rest.90.3.389
- [26] Khoshandam, L., & Nematizadeh, M. (2024). An inverse network DEA model for two-stage processes in the presence of undesirable factors. Journal of applied research on industrial engineering, 11(2), 179–194. https://doi.org/10.22105/jarie.2023.351161.1492
- [27] Jr., F. J. M. (1951). The Kolmogorov-Smirnov test for goodness of fit. Journal of the american statistical association, 46(253), 68–78. https://doi.org/10.1080/01621459.1951.10500769
- [28] Müller, A. (1997). Integral probability metrics and their generating classes of functions. Advances in applied probability, 29(2), 429–443. https://doi.org/https://doi.org/10.2307/1428011

