Skip to content
View terry0809000's full-sized avatar

Block or report terry0809000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
terry0809000/README.md

Tianyi Yu — Curiosity meets care. An AI-assisted portrait with biomedical imagery.

Visit my website   ·   Explore the pipelines   ·   Google Scholar   ·   ORCID

I'm Tianyi Yu. I'm drawn to the stories hidden in clinical notes, the subtle signals in images, and the difficult questions behind a prediction.

I want to build AI that earns people's trust. My research explores how models behave when data are incomplete, populations change, and a good benchmark score meets the complexity of care. Here I share the code, questions and lessons along the way.

Research focus

01 / Clinical NLP · 7 projects
Biomedical language understanding, sentence classification and cross-corpus evaluation.
02 / Computer Vision & Biomedical Imaging · 6 projects
Skin-lesion modelling, image classification and selective referral.
03 / Trustworthy AI · 9 projects
Calibration, uncertainty, robustness and claims matched to evidence.
04 / Prediction Modelling · 10 projects
Clinical risk estimation, temporal validation and transportability.
05 / Health Informatics · 10 projects
EHR analysis, clinical data pipelines and population health surveillance.
06 / AI for Health · 10 projects
Reproducible machine learning for health research, with explicit evaluation and data-access boundaries.

Fresh questions, new code

52 research projects across six connected directions. Here are the latest additions—each with its own code, pipeline guide and scientific cover.

Relative and absolute augmentation gains under observation loss

Relative and absolute augmentation gains under observation loss

Does an augmentation gain remain meaningful when absolute predictive performance falls under observation loss?

Explore the code →
Source-only clinical model selection under missingness stress

Source-only clinical model selection under missingness stress

Do missingness stress tests change source-only model selection, or do competing rules choose the same candidate?

Explore the code →
Objective physiological phenotyping: Sleep-EDF module

Objective physiological phenotyping: Sleep-EDF module

What can objective sleep architecture reveal without turning physiological disruption into a psychiatric label?

Explore the code →
AKI progression: transportability and model updating

AKI progression: transportability and model updating

What changes when an AKI progression model moves between clinical databases and requires local updating?

Explore the code →

Selected pipelines

01 / Beyond Macro-F1

Biomedical sentence classification examined through calibration and perturbation consistency.

CLINICAL NLP · TRUSTWORTHY AI

Explore pipeline →

02 / Selective referral

Skin-lesion modelling with lesion-level splits, calibration and risk–coverage evaluation.

COMPUTER VISION · BIOMEDICAL IMAGING

Explore pipeline →

03 / Temporal transport

ARDS and AHRF prediction across time and ICU databases, with explicit site-shift audits.

PREDICTION MODELLING · HEALTH INFORMATICS

Explore pipeline →

04 / Respiratory surveillance

Next-week escalation modelling with rolling-origin validation and cluster-aware uncertainty.

AI FOR HEALTH · POPULATION SURVEILLANCE

Explore pipeline →

05 / TRACE-Fib

Cross-organ gene-set classification under changing normalization and study weights.

COMPUTATIONAL BIOLOGY · OPEN-DATA REANALYSIS

Explore pipeline →

06 / Reading dynamics

Word-level eye-tracking analysis with lexical features, grouped evaluation and a reusable CLI.

COGNITIVE DATA SCIENCE · PYTHON PACKAGE

Explore pipeline →

How I approach research

Define the question. Make the endpoint, population and information available at prediction time explicit.
Test the transfer. Separate model development from temporal, site and cross-corpus evaluation.
Keep the evidence. Preserve uncertainty, negative findings, provenance and the limits of each experiment.

The full pipeline index includes data-access requirements and scope notes. These repositories host research software and reproduction guides; clinical projects are research-only.

Statistical care. Reproducible code. Claims matched to evidence.

Pinned Loading

  1. ards-ahrf-ehr-temporal-validation ards-ahrf-ehr-temporal-validation Public

    Site-aware temporal EHR validation for ARDS/AHRF using MIMIC-IV and eICU

    Jupyter Notebook

  2. beyond-macro-f1-pubmed-rct-public beyond-macro-f1-pubmed-rct-public Public

    Reproducibility code and sanitized derived evidence for PubMed RCT sentence classification beyond macro-F1

    Jupyter Notebook

  3. bmc-dm-prediction-one-week bmc-dm-prediction-one-week Public

    One-week state-level respiratory illness escalation prediction using public surveillance data

    Python

  4. ham10000-trustworthy-skin-ai ham10000-trustworthy-skin-ai Public

    Jupyter Notebook

  5. trace-fib-pipeline trace-fib-pipeline Public

    Reproducible fibrosis gene-set transfer pipeline: checksum-pinned public inputs, study-aware evaluation, normalization sensitivity, tests and CLI.

    Python

  6. zuco2-scientific-ml-pipeline zuco2-scientific-ml-pipeline Public

    Runnable ZuCo 2.0 eye-tracking machine-learning pipeline with grouped validation and reproducibility outputs.

    Python