Chapter Four · failure evidence
What Propensity Score Matching got wrong, from 75 dissertations
Across the empirical records, propensity score matching frequently encounters practical limitations including severe data loss from unmatched observations, persistent covariate imbalance, and vulnerability to unmeasured confounding. In addition, propensity models often suffer from misspecification in high dimensional settings and are repeatedly outperformed by linear regression and alternative causal estimators. These records come from PhD theses at 26 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Caliper enforcement and strict overlap restrictions cause severe sample loss and participant exclusion
Matching protocols and common support trimming frequently discard large proportions of treated and control participants due to caliper mismatches. This substantial sample attrition leads researchers to reject propensity score matching in favor of methods like regression adjustment or subclassification.
Considered and rejected
Considered and rejected: Decided against HDPS propensity score matching in favor of regression with overlap trimming due to severe sample loss (30-80% in matching vs 1-5% in regression).
Use of antiseizure medications during pregnancy and adverse neonatal outcomes · MSpace - University of Manitoba
Lost to a baseline
Nearest neighbor propensity score matching resulted in substantial sample loss due to unmatched cases, leading the author to switch to subclassification matching.
Three Essays On Noncognitive Factors, Friendship Networks And Education Outcomes · Penn
Lost to a baseline
562 ACDF cases excluded due to lack of suitable CDA propensity score matches (caliper mismatch in 2:1 matching)
Lost to a baseline
Chapter 2.4 (Abortion quasi-experiment): Excluded control participants lacking comparable propensity score mix as well as treated women with very high propensity scores outside common support during 1:1 nearest-neighbor matching.
Parenthood and Life Satisfaction: The Consequences of Childbirth, Alternative Pregnancy Outcomes, and Single Parenting on Well-Being in Several Domains of Life · Leibniz Universität Hannover Repository
Lost to a baseline
4 of 24 ONJ cases (and unmatched osteoporosis controls from the pool of 874) could not be matched during propensity score matching, leaving 20 matched pairs.
Lost to a baseline
166 patients were dropped during propensity score trimming due to lacking common support/overlap (overall sample decreased from 8,495 to 8,329 matched patients: 13 in stage I–II HR+/HER2−, 4 in stage III HR+/HER2−, 42 in stage I–II HER2+, 70 in stage III HER2+, 14 in stage I–II TNBC, and 23 in stage III TNBC)
Neoadjuvant versus adjuvant chemotherapy for older adults with stage I–III breast cancer · UT Austin
Lost to a baseline
Study 1 (Manuscript #3): 2 male outlier participants excluded during propensity score matching due to age distribution constraints.
In-Group and Out-Group Perspectives on Content of Social Groups · open_UMR Marburg DSpace 10.0
Considered and rejected
Considered and rejected: Rejected propensity score matching in favor of multivariable regression models, citing selection bias from censoring and sample size loss from unmatched subjects.
Considered and rejected
Considered and rejected: Rejected propensity score matching (PSM) because it discards unmatched control units and relies on inefficient iterative search, opting instead for entropy balancing.
Lost to a baseline
13 exposed individuals and 807 non-exposed individuals from the initial IBD cohort were excluded because they could not be matched using propensity score and disease duration calipers (PS caliper ±0.05, disease duration caliper ±0.5).
Lost to a baseline
278 towns trimmed from main propensity score analysis because their estimated propensity score fell outside the support of the control group (leaving 53 treated and 102 control out of 439)
Essays on the Political Economy of City Status · DukeSpace
Considered and rejected
Considered and rejected: Decided against using propensity score matching in Paper 2 due to small sample size in subgroups and covariate imbalance causing sample attrition, using multivariable logistic regression instead.
Understanding the adverse impact of centralised care on neonatal outcomes · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected full propensity score matching in favor of propensity score covariate adjustment to avoid discarding unmatched cases in small samples.
The Role Of For-Profit Education In Social Stratification And Social Inequality In The United States · Penn
Lost to a baseline
954 treated customers dropped during 1-to-1 propensity score matching with caliper u = 0.05 (sample reduced from 8,730 to 7,776 treated customers)
ESSAYS ON DIGITAL EXPERIENCE · Cornell
Lost to a baseline
171 treated customers dropped during 1-to-1 rolling entry matching without replacement with caliper u = 0.05 (sample reduced from 977 to 806 treated customers)
ESSAYS ON DIGITAL EXPERIENCE · Cornell
Considered and rejected
Considered and rejected: Nearest-neighbor propensity score matching was rejected in favor of kernel matching due to poor common support across treatment and control groups.
Socioeconomic implications of adverse birth outcomes · Leibniz Universität Hannover Repository
High dimensional covariates and model misspecification induce estimation bias and extreme weights
Propensity score models with complex non-linear specifications, high dimensional noise features, or misclassified clustering frequently overfit and produce extreme inverse probability weights. Misspecifying the propensity or outcome model causes persistent estimation bias and degrades the precision of downstream treatment effect estimates.
Tried and failed
generalized additive models applied to propensity score estimation. Outcome: worse than baseline. Reason: did not improve over logistic regression due to high proportion of dummy variables
Disruption analytics in urban metro systems with large-scale automated data · Imperial
Tried and failed
complex non-linear models for propensity score estimation applied to fairness sample re-weighting under MAR. Outcome: worse than baseline. Reason: overfit binary indicators instead of accurately estimating continuous propensity probabilities
ADVANCING ROBUST AND FAIR STATISTICAL AND MACHINE LEARNING MODELS FOR INCOMPLETE DATA · Penn
Tried and failed
covariate matching with high dimensional noise features applied to experimental design treatment assignment. Outcome: worse than baseline. Reason: irrelevant features degrade match quality on important predictive covariates, reducing finite-sample precision
Tried and failed
standard logistic regression for propensity score estimation applied to clustered data with misclassified dropouts. Reason: misclassification of cluster-level missingness induces substantial estimation bias under small cluster sizes
Lost to a baseline
Logistic regression baseline had less noisy propensity score calibration than BICauseTree due to better data efficiency
Towards trustworthy AI: from local explanations to causal understanding · Oxford
Considered and rejected
Considered and rejected: Rejected propensity score matching (PSM) across full marker sets due to the curse of dimensionality, opting for linear probability model regression adjustment
Markers of a real estate agent's value-add · UT Austin
Tried and failed
naive uncalibrated inverse propensity score weighting applied to treatment effect estimation with misspecified imputations. Reason: asymptotic bias persists and does not diminish with sample size under misspecified initial outcome imputations
Tried and failed
high-dimensional propensity score with many empirical covariates applied to observational cohort confounding adjustment. Outcome: unstable. Reason: adding over 200 empirical covariates compromised balance and created extreme inverse probability weights
Real-World Bleeding with Ibrutinib in B-Cell Malignancies · Penn
Lost to a baseline
Horvitz-Thompson and Hajec ratio estimators under network interference suffered from significantly larger standard errors/variance compared to the proposed OR and DR estimators due to sensitivity to extreme estimated propensity scores
Essays on Adaptive Methods for Inference and Prediction under Dependence · DukeSpace
Tried and failed
doubly robust propensity score estimation applied to observational pricing data. Outcome: no signal. Reason: treatment assignment had no correlation with observed features, breaking propensity score estimation
Tried and failed
doubly robust nearest neighbor imputation estimator applied to missing survey data estimation. Outcome: did not generalise. Reason: both the outcome model and the propensity/probability model were misspecified simultaneously
Topics in survey design and nearest neighbor imputation for survey data · Iowa State
Tried and failed
truncating propensity score weights applied to time-varying causal effect estimation. Outcome: worse than baseline. Reason: truncating weights distorted covariate balance over sequential treatment paths, amplifying bias and root mean squared error
Toward Robust and Transparent Estimation of the Effects of Time-dependent Interventions · Harvard
Tried and failed
high-dimensional propensity score with tree-based models applied to causal treatment effect estimation. Reason: variance-governing parameters impacting treatment and outcome relationships caused large estimation bias
Machine Learning Methods for Decision Making Inference in Healthcare · Georgia Tech
Propensity score matching is outperformed by ordinary regression and alternative causal estimators
In settings with linear outcome structures or standard observational cohorts, linear regression and doubly robust alternatives achieve lower bias, lower variance, and higher statistical power than matching. Furthermore, difference in differences and regression adjustments often match or exceed the performance of propensity score methods without adding estimation complexity.
Tried and failed
propensity score matching with difference in differences applied to observational policy impact estimation. Reason: matched estimation yielded results nearly identical to standard difference in differences without added benefit
Tried and failed
propensity score matching for differential expression applied to single-cell perturbation gene expression detection. Outcome: worse than baseline. Reason: did not outperform standard Student's t-test and Kolmogorov-Smirnov statistical tests
Assessing regulatory function of rare and common variants using expression CROP-sequencing · Georgia Tech
Lost to a baseline
Linear regression achieved lower squared error when the outcome function was strictly linear and well-specified, outperforming matching at N=500.
Lost to a baseline
Sequential matching had lower relative sample efficiency (0.717 to 0.898) compared to OLS regression adjustment under perfectly linear models (LI scenario) at small sample sizes (n=50).
Statistical Analysis and Design of Crowdsourcing Applications · Penn
Lost to a baseline
Under the linear outcome regression and linear propensity score simulation setup (OR1RM1), Hainmueller's entropy balancing achieved an SE and RMSE of 3.46 x 10^-2, outperforming the proposed DR method (3.52 x 10^-2).
Lost to a baseline
When the outcome model is strictly linear in observed covariates, TSLS achieves slightly lower absolute bias of the median than the nonparametric full matching IV estimator.
Instrumental Variables and Mendelian Randomization With Invalid Instruments · Penn
Considered and rejected
Considered and rejected: Rejected propensity score matching in Chapter 2.3 in favor of fixed effects linear regressions with within-person centering to examine causal life-course trajectories rather than simply adjusting for pre-event differences.
Parenthood and Life Satisfaction: The Consequences of Childbirth, Alternative Pregnancy Outcomes, and Single Parenting on Well-Being in Several Domains of Life · Leibniz Universität Hannover Repository
Lost to a baseline
Propensity score matching and prognostic score matching showed much higher estimation bias (-42.01% and -201.33%) than BART-CV (-19.50%) on LaLonde PSID-2 observational data
CAUSAL INFERENCE FOR HIGH-STAKES DECISIONS · DukeSpace
Lost to a baseline
Propensity score adjusted conditional log rank test achieved lower statistical power than unadjusted log rank, adjusted Cox score test, and Lin-Wei robust score test across simulation settings, beating only Kong-Slud.
Instrumental Variable and Propensity Score Methods for Bias Adjustment in Non-Linear Models · Penn
Considered and rejected
Considered and rejected: Rejected propensity score matching to generate control groups because DiD handles pre-intervention mean differences and matching introduces regression-to-the-mean bias when selecting extreme initial-year utilizers.
Cost-Effective Management of Diseases: Early Detection and Interventions for Improved Health Outcomes · Georgia Tech
Considered and rejected
Considered and rejected: Rejected propensity score matching in favor of Mahalanobis nearest distance matching due to known methodological pitfalls and bias.
Propensity score matching struggles with continuous treatments and complex data structures
Standard propensity score matching is poorly suited for continuous exposures, multi-group treatments, or data with high rates of missing covariates. Researchers also reject matching when fixed cohort ratios, differing follow-up times, or clustering variables prevent model convergence and valid paired evaluations.
Considered and rejected
Considered and rejected: Rejected using matched propensity score (PSMATCH) or univariate models alone for disease impact, selecting the doubly robust method for lower variance and narrower confidence intervals
Whole-herd drivers of wean-to-finish mortality under field conditions: a data-driven approach · Iowa State
Tried and failed
paired t-tests for post-matching balance evaluation applied to covariate balance assessment after matching. Reason: matching reduces within-pair variance, causing smaller mean differences to yield misleadingly large t-statistics
Contributions To Multivariate Matching In Observational Studies · Penn
Tried and failed
cohort matching at fixed unbalanced ratios applied to Markov health economic modeling. Reason: significantly longer follow-up times and unbalanced mortality between comparator cohorts prevented valid matching
Considered and rejected
Considered and rejected: Standard propensity score matching (PSM) was rejected/avoided in favor of IPWRA and BCM due to inconsistency and bias when matching with multiple continuous covariates.
Nutritional and environmental impacts of livestock production systems in Canada: a food systems perspective · MSpace - University of Manitoba
Considered and rejected
Considered and rejected: Rejected propensity score matching for causal inference due to documented drawbacks in matching compared to inverse propensity weighting (IPW).
Causal Inference with Selection Bias and Complex Observation Data · DSpace at SUNY Buffalo
Considered and rejected
Considered and rejected: Rejected standard propensity score matching/univariate covariate bounds because they fail when covariate distributions and overlap are non-linear and multi-group
Tackling Key Challenges to Guide Clinical Decisions in Cardiovascular Diseases · MIT
Considered and rejected
Considered and rejected: Generalized boosting modeling for propensity score matching was rejected because it does not create a 1:1 control group of equal size compatible with the survival and CBA design
Cost–Benefit Analysis of the Windham School District’s Correctional CTE Program · Texas Tech
Considered and rejected
Considered and rejected: Rejected including stratification and cluster variables directly into the propensity score estimation model due to failed model convergence.
Behavioral health services for co-occurring mental health and substance use disorders under the Affordable Care Act · JScholarship
Considered and rejected
Considered and rejected: Rejected using Inverse Probability of Treatment Weighting (IPTW) propensity score matching because bivariate baseline demographic/clinical characteristics did not significantly differ between cohorts
Considered and rejected
Considered and rejected: Propensity score matching was rejected because it is designed for binary treatments, while maternity leave duration was continuous
Considered and rejected
Considered and rejected: Rejected using baseline LDL, HDL, and continuous HbA1C in the propensity score model due to high EHR missingness (>30-50%), replacing them with binary measurement-presence indicators
Propensity score matching fails to achieve acceptable covariate balance and increases model dependence
Applying nearest neighbor matching, multiple matches, or strict calipers often worsens balance across observed covariates and increases sensitivity to specification choices. As a result, researchers frequently abandon propensity score matching in favor of coarsened exact matching or entropy balancing.
Tried and failed
propensity score weighting on baseline covariate levels applied to correcting differential pre-treatment trends. Reason: conditioning on baseline level covariates failed to balance differential pre-treatment trajectory trends across groups
Essays in Labor and Public Economics · Cornell
Tried and failed
1-to-1 nearest neighbour propensity score matching applied to clustered observational data. Outcome: worse than baseline. Reason: dropping poor pairs or matching with replacement degraded covariate balance and increased model dependence
Considered and rejected
Considered and rejected: Rejected 1:1 nearest neighbour and optimal propensity score matching due to covariate imbalance; used full matching with probit regression instead.
Considered and rejected
Considered and rejected: Rejected propensity score matching and weighting due to inability to achieve adequate covariate balance across treatment and comparison groups, opting instead for baseline covariate regression adjustment in interrupted time series.
Considered and rejected
Considered and rejected: Rejected using a 0.2 standard deviation propensity score matching caliper because the resulting sample (n=2016) was not balanced across all predictor covariates.
Judgment and Decision-Making in the Context of Health and Law · Cornell
Considered and rejected
Considered and rejected: Rejected standard propensity score matching (PSM) due to model dependence, caliper tradeoffs, and subjective bias from uneven density distributions.
Debiased machine learning causal inference for time-varying social variables · Oxford
Considered and rejected
Considered and rejected: Rejected propensity score matching in favor of coarsened exact matching due to concerns over model dependence and covariate imbalance.
Quasi-Experimental Evaluation of Women's Re-Entry in New Jersey - Through a Black Intersectional Lens · Carleton University Institutional Repository
Considered and rejected
Considered and rejected: Rejected using Propensity Score Matching (PSM), specifically nearest neighbor, because it increased covariate imbalance compared to Coarsened Exact Matching (CEM).
The Effects of Financial and Economic Literacy on Individual Policy Preferences · ResearchWorks
Tried and failed
multi-nearest-neighbor matching with simple ATT estimator applied to causal treatment effect estimation. Outcome: worse than baseline. Reason: increasing number of matches compounded bias under poor covariate balance conditions
Modern Econometric Methods for the Analysis of Housing Markets · Virginia Tech
Considered and rejected
Considered and rejected: Propensity score matching rejected in favor of Mahalanobis distance matching to avoid covariate imbalance and model dependence.
Educational Disparities in Chronic Pain and Life Expectancy: Gaps and Pathways · DSpace at SUNY Buffalo
Propensity score methods cannot correct for unmeasured confounding or structural segregation
Matching and weighting techniques fail when key confounding factors like social networks, activity levels, or unobserved site differences remain unmeasured. In addition, propensity score adjustments break down when treatment and control cohorts are structurally separated across underlying productivity or demographic attributes.
Tried and failed
propensity score weighting and test-then-pool synthesis applied to integrating external heterogeneous control datasets. Outcome: did not generalise. Reason: unmeasured confounders and varying study-specific intercepts caused severe type I error inflation and substantial bias
Tried and failed
adding area-level covariates to propensity score models applied to controlling unmeasured confounding in observational studies. Outcome: no signal. Reason: area-level proxies lacked sufficient granularity to adjust for individual-level unmeasured confounding or shift effect estimates
Leveraging Geographic Information for Causal Inference in Pharmacoepidemiology · Harvard
Considered and rejected
Considered and rejected: Rejected Propensity Score Matching (PSM) for comparing experimental and control groups because groups were not overly imbalanced across baseline variables and PSM could not address unmeasured confounders
Treatment compliance of male perpetrators of intimate partner violence · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected using propensity score matching due to inability to match on unobserved social network and information campaign variables.
Analyzing the drivers of agricultural technology adoption among smallholder farmers in Uganda · Oxford
Considered and rejected
Considered and rejected: Rejected propensity score matching because retrospective covariates were limited and omitted-variable bias could not be addressed without an endogenous treatment model.
Unequal starts: the role of different learning environments in the development of inequalities in skills during early childhood · IRIS - UNITN - prod
Considered and rejected
Considered and rejected: Decided against propensity score matching because only summary/aggregate data were provided by registries and unmeasured confounders like activity level remained.
Considered and rejected
Considered and rejected: Rejected using Propensity Score Matching (PSM) for evaluating co-op vs IOF price impacts because PSM fails when participant groups are structurally segregated by productivity.
Left open by the authors
Problems the authors named and did not get to.
Left open
Evaluate machine learning classification methods against logistic regression for propensity score estimation with imbalanced group sizes in marginal structural models. Blocker: None
Left open
Log-transform outcome variables and evaluate childless versus mother comparisons via propensity score matching on merged PSID and O*NET data. Blocker: None
Impact of Occupational Flexibility on Labor Market Outcomes of Women Following Childbirth · MIT
Left open
Apply matching methods or propensity score estimation to the survival model dataset to control for strategic selection in leader tenure and concessions. Blocker: None
The Causes and Consequences of Territorial Nationalism · Harvard
Left open
Develop average treatment effect estimation tools for observational studies using kernel ridge regression-based propensity score weighting. Blocker: None
Left open
Compare standard propensity score weighting with entropy balancing, matching, and stratification methods in added-growth latent growth models via simulation. Blocker: None
Left open
Implement propensity score matching or multivariate clustering on public Census tract data to select control sites for TOD comparative analysis. Blocker: None
Left open
Develop and evaluate methods using large language models for propensity score estimation and causal inference on observational data. Blocker: None
Language Models as Opinion Models: Techniques and Applications · MIT
Left open
Implement propensity score matching, stratification, machine learning estimators, and trimming methods for latent variable outcome models in lavaan simulation studies. Blocker: None
Latent variable outcome analysis models in a propensity score framework · UT Austin
Left open
Develop methodology to quantify residual and unmeasured confounding for continuous dose matching and subclassification. Blocker: Lack of concrete statistical approach or mathematical framework specified in the thesis
Observational Data With A Continuous Exposure: Study Design And Outcome Analysis · Penn
Left open
Apply propensity score matching to balance confounding factors when evaluating hospital nursing resources and surgical patient outcomes. Blocker: Access to restricted hospital discharge databases and nursing survey data
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.