Visit my website · Explore the pipelines · Google Scholar · ORCID
I'm Tianyi Yu. I'm drawn to the stories hidden in clinical notes, the subtle signals in images, and the difficult questions behind a prediction.
I want to build AI that earns people's trust. My research explores how models behave when data are incomplete, populations change, and a good benchmark score meets the complexity of care. Here I share the code, questions and lessons along the way.
| 01 / Clinical NLP · 7 projects Biomedical language understanding, sentence classification and cross-corpus evaluation. |
02 / Computer Vision & Biomedical Imaging · 6 projects Skin-lesion modelling, image classification and selective referral. |
| 03 / Trustworthy AI · 9 projects Calibration, uncertainty, robustness and claims matched to evidence. |
04 / Prediction Modelling · 10 projects Clinical risk estimation, temporal validation and transportability. |
| 05 / Health Informatics · 10 projects EHR analysis, clinical data pipelines and population health surveillance. |
06 / AI for Health · 10 projects Reproducible machine learning for health research, with explicit evaluation and data-access boundaries. |
52 research projects across six connected directions. Here are the latest additions—each with its own code, pipeline guide and scientific cover.
![]() Does an augmentation gain remain meaningful when absolute predictive performance falls under observation loss? Explore the code → | ![]() Do missingness stress tests change source-only model selection, or do competing rules choose the same candidate? Explore the code → |
![]() What can objective sleep architecture reveal without turning physiological disruption into a psychiatric label? Explore the code → | ![]() What changes when an AKI progression model moves between clinical databases and requires local updating? Explore the code → |
Biomedical sentence classification examined through calibration and perturbation consistency. CLINICAL NLP · TRUSTWORTHY AI Explore pipeline → | Skin-lesion modelling with lesion-level splits, calibration and risk–coverage evaluation. COMPUTER VISION · BIOMEDICAL IMAGING Explore pipeline → |
ARDS and AHRF prediction across time and ICU databases, with explicit site-shift audits. PREDICTION MODELLING · HEALTH INFORMATICS Explore pipeline → | Next-week escalation modelling with rolling-origin validation and cluster-aware uncertainty. AI FOR HEALTH · POPULATION SURVEILLANCE Explore pipeline → |
Cross-organ gene-set classification under changing normalization and study weights. COMPUTATIONAL BIOLOGY · OPEN-DATA REANALYSIS Explore pipeline → | Word-level eye-tracking analysis with lexical features, grouped evaluation and a reusable CLI. COGNITIVE DATA SCIENCE · PYTHON PACKAGE Explore pipeline → |
Define the question. Make the endpoint, population and information available at prediction time explicit.
Test the transfer. Separate model development from temporal, site and cross-corpus evaluation.
Keep the evidence. Preserve uncertainty, negative findings, provenance and the limits of each experiment.
The full pipeline index includes data-access requirements and scope notes. These repositories host research software and reproduction guides; clinical projects are research-only.
Statistical care. Reproducible code. Claims matched to evidence.






