Doe Library seen across Memorial Glade on the UC Berkeley campus

School of Public Health, University of California, Berkeley

Jingshen Wang

Jingshen Wang

I am an Associate Professor of Biostatistics in the School of Public Health at the University of California, Berkeley.

My research develops statistical methods for reliable AI and biomedical data science, with recent interests on model evaluation and AI for health, adaptive experimental design, and precision medicine. Across these areas, my group studies how to generate reliable biomedical evidence while making efficient use of data, human expertise, study participants, and computational resources.


Research in Our Group

  • LLM-augmented inference
  • Model Evaluation
  • AI routers
  • LLM Interpretability

Selected projects

CALM calibration weights by covariate and treatment strata, and the resulting variance reduction in each stratum

Harnessing AI-Generated Information Without Compromising Statistical Validity

Language models can produce useful predictions by drawing on knowledge acquired during pre-training, but these predictions may be biased or unreliable. We develop methods that treat these outputs as auxiliary information and calibrate them against observed data across randomized experiments, adaptive trial designs, and semi-supervised inference. In randomized experiments, our approach CALM uses calibrated language-model predictions to improve power for detecting treatment effects; in the BRIGHTEN depression trial, it identified a significant treatment effect among female Hispanic participants that standard analyses missed. We extend a similar principle to semi-supervised inference, where calibrated language-model predictions helped identify lexical diversity and pronoun ratio as linguistic markers of Alzheimer’s disease.

Model outputs labeled by several silver labelers with no gold labels, and the heterogeneous accuracy profiles of those labelers

Allocating and Using Limited Human Expertise for Reliable Model Evaluation

Human expertise remains essential for reliable model evaluation, but expert review is costly and cannot be obtained for every case. We develop methods that make more efficient use of limited human feedback by reusing existing annotations and selectively allocating new expert review where it is most informative. HERO leverages historical expert annotations to improve the reliability and sensitivity of current model evaluation. We also develop adaptive labeling and routing methods that decide when AI-based evaluation is sufficient and when a case should be escalated to a human expert, using information about the cost and quality of different feedback sources and signals that predict likely disagreement with human judgment.

Embedding-level feature selection: masking whole variable groups in the embedding matrix keeps the hidden dimension unchanged, with selection indicators learned by stochastic gradient descent

Reliable and Interpretable Language Models for Health

Reliable use of language models requires more than accurate prediction: we need to understand what information drives their predictions and whether those signals are scientifically meaningful. We develop statistical methods for feature importance and selection in language models with semi-structured data. Similar in spirit to LASSO, our embedding-level framework learns a sparse set of variable groups, but does so while preserving the structure of the original prompt. Applied to multimodal clinical and biomarker data for predicting amyloid and tau PET pathology, our method identified clinically meaningful variables, improved predictive performance, and reduced token use by 57%. This provides a principled way to make language-model predictions more interpretable, efficient, and scientifically useful.


Papers and Manuscripts

This section has three parts: manuscripts and preprints, journal articles, and AI/ML conference papers.

* corresponding author · ** current or former mentee

Manuscripts and preprints

On the theoretical investigation of mediation analysis in Mendelian randomization with summary data. Manuscript. Lyu, R. Q.**; Ma, X.; and Wang, J.*

Journal articles

Agentic AI for scaling diagnosis and care in neurodegenerative disease. Nature Aging, 2026. Breithaupt, A. G.; Weiner, M.; Tang, A.; Possin, K. L.; Sirota, M.; Lah, J.; Levey, A. I.; Van Hentenryck, P.; Zandehshahvar, R.; Gorno-Tempini, M. L.; Giorgio, J.; Wang, J.; Rauschecker, A. M.; Rosen, H. J.; Nosheny, R. L.; Miller, B. L.; and Pinheiro-Chagas, P.
Rapid biphasic decay of intact and defective HIV DNA reservoir during acute treated HIV disease. Nature Communications, 2024. Barbehenn, A.; Shi, L.**; Shao, J.**; Hoh, R.; Hartig, H. M.; Pae, V.; Sarvadhavabhatla, S.; Donaire, S.; Sheikhzadeh, C.; Milush, J.; Laird, G. M.; Mathias, M.; Ritter, K.; Peluso, M. J.; Martin, J.; Hecht, F.; Pilcher, C.; Cohen, S. E.; Buchbinder, S.; Havlir, D.; Gandhi, M.; Henrich, T. J.; Hatano, H.; Wang, J.; Deeks, S. G.; and Lee, S. A. Barbehenn and Shi are co-first authors.

AI/ML conference papers

MedFrameQA: a multi-image medical VQA benchmark for clinical reasoning. EMNLP, 2026. Yu, S.; Wang, H.; Wu, J.; Luo, L.; Wang, J.; Xie, C.; Rajpurkar, P.; Yang, C.; Yang, Y.; Wang, K.; Yu, Y.; and Zhou, Y.