Representation, Structure, and Context: Developing Computational Frameworks for Real-World Health Data
The growing scale and scope of real-world health data offer new opportunities to advance our understanding of human health, yet their availability alone cannot ensure effective use. Data generated through healthcare are noisy, high-dimensional, and heterogeneous, while diverse collection practices often leave only a fragmented view of an individual and their care. In practice, these data are frequently reduced to collections of discrete features, obscuring relationships across medical histories and uncertainty when information is incomplete. This talk explores the development of computational approaches that more fully account for this complexity by capturing comprehensive and contextualized representations of the available information. In doing so, this work moves beyond simply adding more features or increasing model complexity to improve how machine-learning and statistical models learn from messy, real-world data and reflect variability across patient conditions, their care, and downstream outcomes.