Learning from Medical Images: From Risk Prediction to Patient-Specific World Models
Medical images contain far more information than is typically used for a single diagnostic task. This talk presents our work on learning generalizable, patient-specific representations from routine chest imaging, tracing a progression from task-specific prediction to multimodal foundation models and spatially grounded world models. I will first describe opportunistic cardiovascular risk prediction from low-dose chest computed tomography (CT), illustrating how models can integrate distributed signals across anatomy to predict disease and mortality without additional imaging. I will then present a multimodal and multitask foundation model that jointly learns from 3D CT, clinical variables, and text, for learning transferable chest X-ray representations.
Across these efforts, I will discuss challenges in multimodal learning, label efficiency, distribution shift, and generalization across tasks and datasets. Finally, I will introduce X-WIN, a chest radiograph world model that distills volumetric CT knowledge into 2D X-ray representations through predictive sensing. By learning how observations vary with viewing geometry, X-WIN acquires a latent representation of 3D anatomy that supports transfer and coarse volumetric reconstruction. These results point toward medical AI systems that move beyond task-specific prediction toward richer models of patient state, with applications in risk assessment, spatial reasoning, and image-guided care.