
Harnessing AI-Generated Information Without Compromising Statistical Validity
Language models can produce useful predictions by drawing on knowledge acquired during pre-training, but these predictions may be biased or unreliable. We develop methods that treat these outputs as auxiliary information and calibrate them against observed data across randomized experiments, adaptive trial designs, and semi-supervised inference. In randomized experiments, our approach CALM uses calibrated language-model predictions to improve power for detecting treatment effects; in the BRIGHTEN depression trial, it identified a significant treatment effect among female Hispanic participants that standard analyses missed. We extend a similar principle to semi-supervised inference, where calibrated language-model predictions helped identify lexical diversity and pronoun ratio as linguistic markers of Alzheimer’s disease.
- CALM: Can language models boost the power of randomized experiments without statistical bias? (2026) ↗
- Integrating digital twins with randomized experiments (NeurIPS 2026) ↗
- Language model augmented semi-supervised statistical inference (ICML 2026) ↗
- Large language models enhanced covariate-adjusted response-adaptive randomization design (NeurIPS 2026) ↗











