Abstract
Background:
Large-scale propensity score (LSPS) models are increasingly used to control confounding in observational studies, but their reliability in small-sample settings is unclear. Small samples can limit a study's ability to estimate propensity scores accurately and achieve adequate covariate balance, raising concerns about insufficient confounding adjustment. Resulting in small datasets being excluded, despite potentially containing valid information.
Methods:
We evaluated the balancing and bias reduction performance of LSPS models using six target-comparator pairs. To emulate a distributed data network analysis setting, we partitioned two large real-world data sources into smaller subsets. Within each subset, LSPS models were fit locally and evaluated on covariate balance, treatment effect bias and precision after empirical calibration, using real negative controls and synthetic positive controls. Performance was compared to a global LSPS model trained on the full dataset.
Results:
Across most target-comparator pairs, LSPS models substantially reduced bias and, when applying empirical calibration, improved precision. PS matching was more resilient to small sample sizes than PS stratification, achieving adequate balance even at n = 500, while stratification sometimes failed at n = 4000. A slight bias increase was observed at the smallest sample sizes, though not universally. Standard balance diagnostics consistently failed below n = 20 000, while a recently proposed diagnostic accounting for chance imbalance did not.
Conclusions:
LSPS models generally provide reliable bias reduction in small-sample settings, supporting their use in federated analyses. However, standard balance diagnostics may be misleading in small samples, and alternatives should be considered, such as significance checking. When LSPS fails to reduce bias adequately, additional adjustment strategies are required.
Large-scale propensity score (LSPS) models are increasingly used to control confounding in observational studies, but their reliability in small-sample settings is unclear. Small samples can limit a study's ability to estimate propensity scores accurately and achieve adequate covariate balance, raising concerns about insufficient confounding adjustment. Resulting in small datasets being excluded, despite potentially containing valid information.
Methods:
We evaluated the balancing and bias reduction performance of LSPS models using six target-comparator pairs. To emulate a distributed data network analysis setting, we partitioned two large real-world data sources into smaller subsets. Within each subset, LSPS models were fit locally and evaluated on covariate balance, treatment effect bias and precision after empirical calibration, using real negative controls and synthetic positive controls. Performance was compared to a global LSPS model trained on the full dataset.
Results:
Across most target-comparator pairs, LSPS models substantially reduced bias and, when applying empirical calibration, improved precision. PS matching was more resilient to small sample sizes than PS stratification, achieving adequate balance even at n = 500, while stratification sometimes failed at n = 4000. A slight bias increase was observed at the smallest sample sizes, though not universally. Standard balance diagnostics consistently failed below n = 20 000, while a recently proposed diagnostic accounting for chance imbalance did not.
Conclusions:
LSPS models generally provide reliable bias reduction in small-sample settings, supporting their use in federated analyses. However, standard balance diagnostics may be misleading in small samples, and alternatives should be considered, such as significance checking. When LSPS fails to reduce bias adequately, additional adjustment strategies are required.
| Original language | English |
|---|---|
| Number of pages | 10 |
| Journal | Journal of the American Medical Informatics Association |
| DOIs | |
| Publication status | E-pub ahead of print - Aug 2026 |
Fingerprint
Dive into the research topics of 'Evaluating large-scale propensity score adjustment when sample size is small'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver