Study analysis · Scientific Reports · 2025
Your doctor’s computer can predict if you’ll get diabetes—7 years before you even feel sick.
A computer looked at your medical records and figured out if you’ll get type 2 diabetes years in advance—and even grouped people into three types that respond differently to medicine.
Overview
What the study found
The study in plain English — the bottom line, every takeaway we extracted, and what to do with them.
In simple terms
This study found that a computer program can look at people's medical records and guess who might get diabetes in the future, based on things like weight and medications. But it didn't change anyone's treatment or prove that catching it early helps—it just noticed patterns.
What’s the bottom line?
Scientists taught a computer to read doctors' notes and test results to find people who might get type 2 diabetes years in advance—and to group them into different types based on their health patterns.
How strong is this study?
The study used lots of real patient data from two big hospital systems and tested its computer model carefully, which makes it pretty reliable for spotting patterns. But since it didn't actually help patients or control for things like diet or income, we can't be sure the patterns are truly about diabetes itself.
40 / 100
- COI disclosure+40/40
- Data availabilitydata not shared
- Code availabilitycode not shared
56 / 100
- Randomizationnot randomized
- Blindingblinding unclear
- Control group+15/15
- Sample size (n=10865)+20/20
- Follow-up+10/10
100 / 100
77 / 100
- P-values+15/15
- Effect size+20/20
- Confidence intervals+15/15
- Pre-registrationnot pre-registered
Each component is scored out of 100 and then capped by the study design — a case series cannot reach the ceiling a randomised trial can, however well it is reported.
Where it sits
RCT reviewsReviews of RCTs (Meta-analyses)
Max 100Randomized TrialsRandomized Trials
Max 90Reviews of Cohort StudiesReviews of Cohort Studies
Max 85Cohort StudiesCohort Studies
Max 72Reviews of Case-Control StudiesReviews of Case-Control Studies
Max 63Case-Control StudiesCase-Control Studies
Max 58Cross-Sectional & Case SeriesCross-Sectional & Case Series
Max 50Expert OpinionExpert Opinion
Max 50 / 100
Probability of being correct
Based on clinical experience or non-systematic literature reviews. The lowest level of evidence as they are most susceptible to bias and personal perspective.
This design cannot establish causation — the findings describe an association, not a cause. This is an observational study using retrospective electronic health record data without randomization or intervention. It identifies patterns and associations but cannot rule out confounding factors or establish that the model causes changes in diabetes outcomes.
No Conflicts
No conflicts of interest identified
No conflicts of interest or funding disclosures were reported in the study text.
The study does not include any conflict of interest, funding, or author affiliation disclosures. While the use of large EHR datasets from All of Us and MGB Biobank suggests potential institutional support, no explicit funding sources or industry ties are stated. The absence of disclosure limits full transparency but does not indicate evidence of bias.
Key takeaways
- 01
The computer predicted diabetes 7 years ahead with 75.4% accuracy (AUC 0.754).
- 02
It found 3 types: one (Green) had fewer health problems and lowered blood sugar by 0.64% after metformin; another (Red) had more problems and only lowered it by 0.27%.
- 03
Yes—this means doctors could spot high-risk patients earlier and give the right treatment to the right group, like giving metformin sooner to those who respond best.
Surprising findings
- The Red subtype’s poor response to metformin wasn’t due to higher BMI alone—after adjusting for weight, cardiovascular and mental health differences still persisted.People assume obesity is the main driver of bad diabetes outcomes—but this shows mental health and heart issues independently shape treatment failure.
- The model’s predictive power came from routine EHR data—no genetic tests, no special biomarkers—just standard doctor visits and lab results.We think precision medicine needs DNA tests—but this proves you can get highly accurate predictions from existing medical records.
Practical takeaways
If you have prediabetes, ask your doctor if your EHR data could be analyzed for diabetes subtyping to tailor prevention strategies.
This model isn’t yet available in most clinics—it’s still in research phase and requires access to large EHR datasets.
high confidenceTrack your mental health and cardiovascular symptoms (like sleep apnea or depression)—they may be early signals of a high-risk diabetes subtype.
These patterns are predictive, not diagnostic—don’t self-label based on symptoms alone.
medium confidenceWhy this study matters
Predicts Diabetes 7 Years Ahead
A deep learning model analyzed electronic health records and predicted type 2 diabetes onset up to 7 years in advance with 75.4% accuracy (AUC 0.754)—beating traditional methods like blood sugar tests (AUC 0.632) and risk factor models (AUC 0.693).
Most people only get screened when they’re already at risk—this could let doctors warn you years before symptoms appear, giving you time to prevent it.
Three Diabetes Types—Not One Size Fits All
The model identified three subtypes: Green (mild, few comorbidities), Yellow (moderate), and Red (severe, with obesity, heart disease, and depression). The Red subtype had 2.4x higher rates of sleep apnea and 2.3x higher depression rates than Green.
It’s not just ‘diabetes’—your type determines how bad your complications will be and how well you’ll respond to treatment.
Metformin Works Better for Some
People in the Green subtype saw their HbA1c drop by 0.64% after starting metformin—nearly double the 0.27% drop seen in the Red subtype, despite both groups receiving the same drug.
You might be taking the right medicine—but if your subtype isn’t considered, it might not work well for you.
It’s Not Your Genes—It’s Your Life
The subtypes showed no link to polygenic risk scores (PRS), meaning your genetic predisposition didn’t determine your type—instead, lifestyle, environment, and clinical history did.
This flips the script: your diabetes isn’t just inherited—it’s shaped by your habits, stress, and access to care.
Want the whole report?
Detailed mode opens the full scientific breakdown — every score component, the methodology, conflicts of interest, the evidence analysis behind each claim, and the raw study data.
Overview
What the study found
The study in plain English — the bottom line, every takeaway we extracted, and what to do with them.
Not medical advice. For informational purposes only. Always consult a healthcare professional. Terms
Scientists taught a computer to read doctors' notes and test results to find people who might get type 2 diabetes years in advance—and to group them into different types based on their health patterns.
Research results
The computer predicted diabetes 7 years ahead with 75.4% accuracy (AUC 0.754). It found 3 types: one (Green) had fewer health problems and lowered blood sugar by 0.64% after metformin; another (Red) had more problems and only lowered it by 0.27%.
What this means - more context
Yes—this means doctors could spot high-risk patients earlier and give the right treatment to the right group, like giving metformin sooner to those who respond best.
This study develops a deep metric learning (DML) model to simultaneously predict type 2 diabetes onset and identify subtypes using routine electronic health record (EHR) data, aiming to enable opportunistic screening and precision medicine.
A DML model trained on EHR data from over 10,000 patients predicted T2D onset up to 7 years in advance (AUC 0.754), outperforming traditional models. It also identified three reproducible subtypes (Green, Yellow, Red) based on clinical similarity, with the Red subtype showing higher comorbidities and poorer metformin response (HbA1c reduction: −0.27% vs. −0.64% in Green). Subtypes were not linked to polygenic risk scores, indicating they reflect clinical/environmental factors.
Methods Used
The study used EHR data from 7,567 T2D cases and 3,298 T2D cases in the All of Us and MGB Biobank cohorts. Features included conditions, medications, labs, and demographics. A deep metric learning encoder learned a latent space using triplet loss, followed by logistic regression for onset prediction and K-means clustering (k=3) for subtyping. Models were validated across cohorts and compared to LR, SCARF, TabTransformer, PCA, and UMAP.
Main Finding
The DML model predicted T2D onset 7 years ahead with an AUC of 0.754, surpassing logistic regression (0.706), clinical risk models (0.693), and glycemic measures (0.632). Three subtypes were identified: the Red subtype had significantly higher obesity-related, cardiovascular, and mental health comorbidities, and showed a smaller HbA1c reduction after metformin (−0.27%) compared to the Green subtype (−0.64%).
Confidence Level
High confidence due to large, diverse cohorts (n=10,865), external validation across two independent healthcare systems, rigorous statistical testing with Bonferroni correction, bootstrap confidence intervals, and comparison against multiple established baselines.
Study Flags
Red Flags
- •Retrospective design with potential selection bias from hospital data
- •Lack of family history data may limit predictive power
- •Subtypes form a continuum rather than distinct clusters, requiring further clinical validation
No biological mechanisms were identified in this study. This may be an epidemiological, observational, or survey-based study that reports associations rather than proposing causal biological pathways.
Surprising Findings
The Red subtype’s poor response to metformin wasn’t due to higher BMI alone—after adjusting for weight, cardiovascular and mental health differences still persisted.
People assume obesity is the main driver of bad diabetes outcomes—but this shows mental health and heart issues independently shape treatment failure.
Practical Takeaways
If you have prediabetes, ask your doctor if your EHR data could be analyzed for diabetes subtyping to tailor prevention strategies.
RCT reviewsReviews of RCTs (Meta-analyses)
Max 100Randomized TrialsRandomized Trials
Max 90Reviews of Cohort StudiesReviews of Cohort Studies
Max 85Cohort StudiesCohort Studies
Max 72Reviews of Case-Control StudiesReviews of Case-Control Studies
Max 63Case-Control StudiesCase-Control Studies
Max 58Cross-Sectional & Case SeriesCross-Sectional & Case Series
Max 50Expert OpinionExpert Opinion
Max 50 / 100
Probability of being correct
Based on clinical experience or non-systematic literature reviews. The lowest level of evidence as they are most susceptible to bias and personal perspective.
Non-Scorable
Subject
Lower probability
on the GRADE evidence scale
This study found that a computer program can look at people's medical records and guess who might get diabetes in the future, based on things like weight and medications. But it didn't change anyone's treatment or prove that catching it early helps—it just noticed patterns.
No conflicts of interest were detected in this study. No score impact.
Strengths
- Large sample size across two independent, diverse cohorts (over 10,000 participants)
- Use of validated algorithms (eMERGE, PheCap) for T2D case identification
- Robust internal and external validation with hold-out test sets and cross-cohort transfer testing
Weaknesses
- Retrospective design limits ability to infer temporal causality
- No randomization or intervention to test clinical impact
- Blinding status unknown, raising potential for outcome assessment bias
Methodology
Evidence Keywords
Statistical Reporting
Not medical advice. For informational purposes only. Always consult a healthcare professional. Terms
Scientists taught a computer to read doctors' notes and test results to find people who might get type 2 diabetes years in advance—and to group them into different types based on their health patterns.
Research results
The computer predicted diabetes 7 years ahead with 75.4% accuracy (AUC 0.754). It found 3 types: one (Green) had fewer health problems and lowered blood sugar by 0.64% after metformin; another (Red) had more problems and only lowered it by 0.27%.
What this means - more context
Yes—this means doctors could spot high-risk patients earlier and give the right treatment to the right group, like giving metformin sooner to those who respond best.
This study develops a deep metric learning (DML) model to simultaneously predict type 2 diabetes onset and identify subtypes using routine electronic health record (EHR) data, aiming to enable opportunistic screening and precision medicine.
A DML model trained on EHR data from over 10,000 patients predicted T2D onset up to 7 years in advance (AUC 0.754), outperforming traditional models. It also identified three reproducible subtypes (Green, Yellow, Red) based on clinical similarity, with the Red subtype showing higher comorbidities and poorer metformin response (HbA1c reduction: −0.27% vs. −0.64% in Green). Subtypes were not linked to polygenic risk scores, indicating they reflect clinical/environmental factors.
Methods Used
The study used EHR data from 7,567 T2D cases and 3,298 T2D cases in the All of Us and MGB Biobank cohorts. Features included conditions, medications, labs, and demographics. A deep metric learning encoder learned a latent space using triplet loss, followed by logistic regression for onset prediction and K-means clustering (k=3) for subtyping. Models were validated across cohorts and compared to LR, SCARF, TabTransformer, PCA, and UMAP.
Main Finding
The DML model predicted T2D onset 7 years ahead with an AUC of 0.754, surpassing logistic regression (0.706), clinical risk models (0.693), and glycemic measures (0.632). Three subtypes were identified: the Red subtype had significantly higher obesity-related, cardiovascular, and mental health comorbidities, and showed a smaller HbA1c reduction after metformin (−0.27%) compared to the Green subtype (−0.64%).
Confidence Level
High confidence due to large, diverse cohorts (n=10,865), external validation across two independent healthcare systems, rigorous statistical testing with Bonferroni correction, bootstrap confidence intervals, and comparison against multiple established baselines.
Study Flags
Red Flags
- •Retrospective design with potential selection bias from hospital data
- •Lack of family history data may limit predictive power
- •Subtypes form a continuum rather than distinct clusters, requiring further clinical validation
No biological mechanisms were identified in this study. This may be an epidemiological, observational, or survey-based study that reports associations rather than proposing causal biological pathways.
Surprising Findings
The Red subtype’s poor response to metformin wasn’t due to higher BMI alone—after adjusting for weight, cardiovascular and mental health differences still persisted.
People assume obesity is the main driver of bad diabetes outcomes—but this shows mental health and heart issues independently shape treatment failure.
Practical Takeaways
If you have prediabetes, ask your doctor if your EHR data could be analyzed for diabetes subtyping to tailor prevention strategies.
RCT reviewsReviews of RCTs (Meta-analyses)
Max 100Randomized TrialsRandomized Trials
Max 90Reviews of Cohort StudiesReviews of Cohort Studies
Max 85Cohort StudiesCohort Studies
Max 72Reviews of Case-Control StudiesReviews of Case-Control Studies
Max 63Case-Control StudiesCase-Control Studies
Max 58Cross-Sectional & Case SeriesCross-Sectional & Case Series
Max 50Expert OpinionExpert Opinion
Max 50 / 100
Probability of being correct
Based on clinical experience or non-systematic literature reviews. The lowest level of evidence as they are most susceptible to bias and personal perspective.
Non-Scorable
Subject
Lower probability
on the GRADE evidence scale
This study found that a computer program can look at people's medical records and guess who might get diabetes in the future, based on things like weight and medications. But it didn't change anyone's treatment or prove that catching it early helps—it just noticed patterns.
No conflicts of interest were detected in this study. No score impact.
Strengths
- Large sample size across two independent, diverse cohorts (over 10,000 participants)
- Use of validated algorithms (eMERGE, PheCap) for T2D case identification
- Robust internal and external validation with hold-out test sets and cross-cohort transfer testing
Weaknesses
- Retrospective design limits ability to infer temporal causality
- No randomization or intervention to test clinical impact
- Blinding status unknown, raising potential for outcome assessment bias
Methodology
Evidence Keywords
Statistical Reporting
Scoring
How strong is this study?
The study used lots of real patient data from two big hospital systems and tested its computer model carefully, which makes it pretty reliable for spotting patterns. But since it didn't actually help patients or control for things like diet or income, we can't be sure the patterns are truly about diabetes itself.
40 / 100
- COI disclosure+40/40
- Data availabilitydata not shared
- Code availabilitycode not shared
56 / 100
- Randomizationnot randomized
- Blindingblinding unclear
- Control group+15/15
- Sample size (n=10865)+20/20
- Follow-up+10/10
100 / 100
77 / 100
- P-values+15/15
- Effect size+20/20
- Confidence intervals+15/15
- Pre-registrationnot pre-registered
Each component is scored out of 100 and then capped by the study design — a case series cannot reach the ceiling a randomised trial can, however well it is reported.
Where it sits
RCT reviewsReviews of RCTs (Meta-analyses)
Max 100Randomized TrialsRandomized Trials
Max 90Reviews of Cohort StudiesReviews of Cohort Studies
Max 85Cohort StudiesCohort Studies
Max 72Reviews of Case-Control StudiesReviews of Case-Control Studies
Max 63Case-Control StudiesCase-Control Studies
Max 58Cross-Sectional & Case SeriesCross-Sectional & Case Series
Max 50Expert OpinionExpert Opinion
Max 50 / 100
Probability of being correct
Based on clinical experience or non-systematic literature reviews. The lowest level of evidence as they are most susceptible to bias and personal perspective.
This design cannot establish causation — the findings describe an association, not a cause. This is an observational study using retrospective electronic health record data without randomization or intervention. It identifies patterns and associations but cannot rule out confounding factors or establish that the model causes changes in diabetes outcomes.
No Conflicts
No conflicts of interest identified
No conflicts of interest or funding disclosures were reported in the study text.
The study does not include any conflict of interest, funding, or author affiliation disclosures. While the use of large EHR datasets from All of Us and MGB Biobank suggests potential institutional support, no explicit funding sources or industry ties are stated. The absence of disclosure limits full transparency but does not indicate evidence of bias.