Data-driven prioritization of high-risk individuals for weight loss interventions
All studies were conducted according to all relevant ethical regulations and principles of the Declaration of Helsinki. UKB was granted ethical approval from the North West Centre for Research Ethics Committee (11/NW/0382), in the form of a Research Tissue Bank (RTB) approval. The EPIC-Norfolk study has been approved by the Norfolk Research Ethics Committee, under the reference number 05/Q0101/191. G&H was approved by the London Southeast NRES Committee of the Health Research Authority (14/LO/1240). SURMOUNT-1 was conducted according to the principles of the Declaration of Helsinki and Good Clinical Practice guidelines and was approved by independent ethics committees or institutional review boards at each of the 119 sites across nine countries7.
Statistics and reproducibility
The UKB is a prospective, population-based observational study of ~500,000 individuals residing in the UK, aged 40–69 at enrollment, who were recruited between 2006 and 2010 across >20 centers17. All participants provided informed consent, and study procedures were performed according to the Helsinki declaration. This study was conducted with data accessed under the application numbers 44448 and 30418.
No statistical method was used to predetermine sample size. No randomization or blinding was applied in this observational study. To increase generalizability and reflect a scenario close to clinical reality, we applied exclusion criteria closely matching those of the SURMOUNT-1 trial7 to define the study population (Fig. 1c). However, we did not exclude individuals with prevalent T2D because this population has been investigated separately in SURMOUNT-225, and we did not exclude individuals with weight change >5 kg within the last 90 days because this information was not reliably obtainable7. We included participants with a BMI equal or above 27 kg m−2 with or without prevalent comorbidities, and we excluded participants with contraindications for dual agonist therapies, with self-reported pregnancy or uncertainty about pregnancy or with diagnoses of thyroid cancer7. Additionally, we excluded participants with other potential causes of obesity: that is, owing to Cushing’s syndrome or reported intake of tricyclic antidepressants, atypical antipsychotics, mood stabilizers or paroxetine. We excluded participants with a history of bariatric surgery, estimated glomerular filtration rate < 30 ml min−1 per 1.73 m2, acute or chronic pancreatitis, alcoholism, recent cancer diagnosis (within the last 5 years) or recent MACE within the last 6 months. After applying these criteria, 197,264 participants remained for the final analyses. Analyses are based on distinct samples and did not include technical or biological replicates. All tests were two-sided.
Exposure
To establish a data-driven discovery of predictors of obesity-associated complications, we collated a comprehensive dataset of 2,390 exposure variables by leveraging multimodal data40 (Supplementary Table 9).
We aggregated data on general characteristics of participants, including data on age, sex (self-reported), self-reported overall health, socioeconomic status, deprivation indices and related factors. We included behavior parameters, such as dietary habits and activity levels. To integrate medical history of the participants, we have gathered a comprehensive set of prevalent medical conditions for each participant from different data sources including self-reported conditions, primary care records where available (~45% of the study population), hospital episode statistics and cancer registries, as outlined previously41. We mapped these medical conditions onto ‘phecodes’—that is, medical concept terms—as described41. We use the term disease/diagnosis rather ‘phecode’ for accessibility.
We included information on regular medication and health supplements from self-report, which were categorized based on the Anatomical Therapeutic Chemical classification system as described previously41, yielding a total of 704 individual Anatomical Therapeutic Chemical codes. We have included a wide-ranging set of clinically established blood biomarkers and blood cell counts, the measurement of which have been described before42. We collated basic cardiopulmonary measurements including heart rate, blood pressure and spirometry parameters. In addition to basic body composition parameters (for example, weight, height, BMI, waist circumference, hip circumference and waist-to-hip and waist-to-height ratios), we considered bioelectrical impedance measures. We included metabolites measured with the Nightingale platform using the nuclear magnetic resonance method43,44. We included polygenic scores for a diverse set of 36 different diseases and traits, as provided by the UKB45.
After filtering highly correlated variables (r > 0.9), whereby variables were kept either due to lower missingness or clinical availability, a total of 2,078 features were included in the predictor pool. We applied one-hot encoding to categorical variables, transforming them into binary indicator variables. We standardized continuous variables using Z-score transformation, where each value was centered by subtracting the variable mean and scaled by dividing by its standard deviation. We used multiple imputation to impute missing data using the miceRanger package (version 1.5.0), excluding variables with missingness above 25%, and generated a single dataset with five iterations46.
Outcomes
We computed prediction models for 18 different obesity outcomes3,4,5. We set the recruitment date in UKB as the baseline of the observation period. Follow-up was censored at 10 years based on the time of event for the outcome or date of death as retrieved from the death register, whichever occurred first. Incident complications were defined based on phecode definitions as described in detail previously41, from primary care where available, hospital episode statistics and self-report, except for T2D, for which we adapted a previously described definition to correctly distinguish between the different types, and for myocardial infarction and stroke, for which we adapted algorithmically defined outcomes provided by UKB47 (Supplementary Table 10). For each outcome investigated, we removed prevalent cases (additionally using tests that were available at baseline, such as blood pressure measurements for hypertension). To minimize bias arising from delayed diagnosis, we excluded incident complications occurring within the first 6 months of follow-up. The number of cases based on availability of primary care data are shown in Supplementary Table 11.
Statistical analysis
Feature selection, optimization and internal validation
We employed a two-step machine learning procedure to identify sets of predictors consisting of the 20 most important features for each of the 18 complications, as recently applied to identify proteomic signatures predicting diseases20,21. This procedure involves first applying a least absolute shrinkage and selection operator (LASSO) model to identify the most relevant predictors from a large pool of features. In the second step, these selected predictors are used in a regularized Cox proportional hazards model for optimization and validation, providing a robust framework for feature selection and outcome prediction.
We randomly divided the cohort of eligible participants into three sets for each outcome: (1) feature selection (50%), (2) optimization (25%), and (3) validation (25%). In the feature selection set for each outcome, we applied LASSO with 250 random subsamples, each containing a fraction of 40% of the feature selection set. For each outcome, we compared different pools of possibly selectable features: that is, either by comparing each category (for example blood biomarkers, diseases and drugs and plasma metabolites, among others) of features individually or by increasing the pool of features by adding categories in a stepwise manner. For each outcome and from each set of feature pools, we took forward features with the top 20 highest feature selection scores—that is, those selected most frequently across subsamples—to be used in the optimization and validation sets. Our approach, refined over iterations for similar feature selection use cases20,21,22, is conceptionally equivalent to stability selection strategies48. In the optimization and test sets, we used a Cox proportional hazards model with L2 regularization (ridge). We performed fivefold cross-validation on the optimization set to identify the optimal lambda parameter using a fine grid. For validation, we implemented a bootstrap procedure with 1,000 bootstrap samples to assess model performance on the independent test dataset. In each bootstrap sample, we calculated the C-index to serve as a discriminatory index for predictive performance. We took forward the average C-index and 95% CIs using the 2.5th and 97.5th percentiles of the bootstrap distribution.
Derivation and validation of a shared model (OBSCORE) for all outcomes
In a clinical setting, implementation of a single model with a core set of features that predict multiple complications would be most practical and hence more feasible to implement. Considering that multimorbid obesity—that is, the occurrence of multiple complications in a single individual–is common3, the clinical scenario can be reframed as a multilabel classification problem. We handled this problem as separate binary classification tasks, building separate models for each outcome individually, as explained above. From each model, we extracted the top 20 clinically available features that were consistently selected: that is, with non-zero coefficients in the training set. In the optimization sets, after further pruning for correlated features (r > 0.6), we tested the feature importance of these features using the leave-one-out covariate importance procedure49, and we subsequently ranked these features by mean importance across all 18 outcomes and took forward the top 20. This set of shared features represented a core group of predictors that showed relevance across multiple conditions. We derived 18 separate models to predict each outcome, each retrained using these 20 features in the optimization sets and tested in the validation sets.
External validation
To test for generalizability of the shared OBSCORE model, we additionally performed external validation in the EPIC-Norfolk study20,23. EPIC-Norfolk is a prospective cohort study of more than 25,000 middle-aged individuals from the general population of Norfolk, England.
In the EPIC-Norfolk study, mortality was retrieved from the UK Office of National Statistics, as described before20,21,23. Hospital episode statistics and certificates of death were retrieved from the NHS digital database, with National Health Service numbers. Health records were coded by trained nosologists according to the International Statistical Classification of Diseases and Related Health Problems, 9th (ICD-9) or 10th Revision (ICD-10), and codes were merged owing to the long-term follow-up. Events were determined in case of disease as underlying cause of death or as the reason for hospitalization. To allow for comparisons, we used the same ICD-10/ICD-9 definitions for baseline disease exclusion and incident disease definitions, except for T2D (ICD-9: 250, ICD-10: E10–E14) and stroke (ICD-9: 433–435, ICD-10: I63, I65, I66). For external validation, we included 2,112 individuals from the EPIC-Norfolk cohort with a BMI ≥ 27 and with complete data on predictor variables (18 out of 20 variables were available, except for cystatin C and joint pain). We conducted a complete case analysis. The variables underwent the same standardization processes as in the UKB, including Z-score transformation of continuous variables and logarithmic transformation of skewed variables. We were able to externally validate 14 outcomes with more than 20 incident cases during a 10-year follow-up: angina pectoris (155 cases), arthropathy (122), cholelithiasis (48), coronary atherosclerosis (143), cardiovascular death (94), diaphragmatic hernia (128), GERD (78), hypercholesterolemia (60), hypertension (112), ischemic heart disease (217), MACE (109), myocardial infarction (106), stroke (59) and T2D (114). We used coefficients (model weights) from the OBSCORE model trained with 18 available predictors and tested performance in external validation: that is, without retraining the models. We evaluated performance using C-indices in analogy to UKB analyses, calculated over 1,000 bootstrap samples. We excluded prevalent cases based on diagnoses or early incident cases during follow-up (within the first 6 months).
To further assess generalizability across ancestries, we performed external validation for incident T2D in the G&H study, a population at high risk for T2D and with sufficient incident events to support reliable analyses for this outcome. G&H is a prospective study enrolling individuals from self-reported British Pakistani and British Bangladeshi backgrounds and aged 16 or over since 2015. We curated and harmonized routine UK NHS EHR data from primary care and secondary care and death information from the Office for National Statistics50. We extracted first occurrences of diagnoses and closest quantitative measures from EHRs to study baselines. T2D was ascertained by EHR-based diagnoses as described before51, and we used HbA1c > 48 mmol mol−1 to additionally exclude prevalent undiagnosed cases. A total of 1,740 individuals with complete data on 15 OBSCORE features, with a South Asian BMI cut-off of ≥22 kg m−2, which corresponds to a BMI of ~27 kg m−2 in white populations, as informed by ref. 52, and without prevalent T2D at baseline (study enrollment), were included for external validation. A total of 337 cases occurred during a median follow-up of 6.3 years, excluding incident cases within occurring within 6 months after baseline. We used coefficients from the OBSCORE model trained with the 15 available features (excluding urate, cystatin C, family history of heart disease, joint and non-specific chest pain) in the UKB for external validation.
Performance comparison with established models
We compared this model with the best-performing outcome-specific models—a model containing age, sex and BMI—and to features of other established classifiers of obesity: for example, metabolically healthy obesity (age, sex and binary variable based on systolic blood pressure < 130 mmHg and no antihypertensive medication, waist-to-hip ratio < 0.95 for women or < 1.03 for men and no prevalent T2D) as well as the SCORE2 model (age, sex, smoking status, total cholesterol, high-density lipoprotein cholesterol, systolic blood pressure) and ASCVD risk score (age, sex, smoking status, ancestry, systolic blood pressure, diastolic blood pressure, total cholesterol, high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, T2D, antihypertensive medication, cholesterol-lowering medication, acetylsalicylic acid intake)35,53,54.
Calculation of performance metrics
To test how the shared clinical model performs at risk stratification, we computed the linear predictors for each outcome, resembling individual risk assigned by the prognostic model. We categorized the estimated risk by 20% (quintiles) of individuals for each outcome and calculated and visualized the actual observed 10-year event rate for each 20% group. We calculated risk ratios between the highest 20% and the lowest 20%. Moreover, we calculated rate ratios between the top 50% higher and lower risk groups and compared them to the rate ratios of individuals with obesity versus overweight. Using the estimated risk, we computed additional performance parameters including detection rates across FPRs ranging from 5% to 40%. FPRs were calculated as \(\mathrm{FPR}=\frac{\mathrm{FP}}{\mathrm{TN}+\mathrm{FP}}\), where TN is true negatives; and detection rates (DRs) were calculated as \(\mathrm{DR}=\frac{\mathrm{TP}}{\mathrm{FN}+\mathrm{TP}}\), where TP is true positives and FN is false negatives. Likelihood ratios (LRs) were computed as \(\mathrm{LR}=\frac{\mathrm{DR}}{\mathrm{FPR}}\). To assess the real-world translational potential of the models, we calculated pre-test and post-test probabilities. Pre-test probabilities were derived from cumulative complication incidence in the study, and pre-test odds were calculated as \(\mathrm{pretest}\,\mathrm{odds}=\frac{\mathrm{pretest}\,\mathrm{probability}}{\left(1-\mathrm{pretest}\,\mathrm{probability}\right)}\). Post-test probabilities at 10% FPR were calculated as \(\mathrm{post}-\mathrm{test}\,\mathrm{probability}=\frac{(\mathrm{pretest}\,\mathrm{odds}\,\times \,\mathrm{LR})}{1+(\mathrm{pretest}\,\mathrm{odds}\,\times \,\mathrm{LR})}\). We assessed calibration visually with calibration plots and by calculating ratios and differences between estimated and observed risks (from Kaplan–Meier estimates) in the held-out validation set.
OBSCORE integration in the randomized controlled SURMOUNT-1 trial
SURMOUNT-1 is a randomized controlled trial of participants with BMI ≥30 kg m−2 or ≥27 kg m−2 with at least one weight-related comorbidity treated with tirzepatide (5 mg, 10 mg, 15 mg) versus placebo (NCT04184622)7. We applied the OBSCORE risk model to individuals participating in SURMOUNT-1. A total of 12 features of OBSCORE (age, sex, waist-to-height ratio, hypertension, total cholesterol, high-density lipoprotein cholesterol, HbA1c, alanine aminotransferase, creatinine, cystatin C, urate, smoking status) were available at baseline and at 72 weeks. Coefficients from models retrained with these 12 features were applied to participants from SURMOUNT-1 with available data (nplacebo = 388, nTZP 5mg = 462, nTZP 10mg = 479, nTZP 15mg = 475) to calculate risk scores at baseline and at week 72. We first investigated the heterogeneity of treatment effects (body weight change, waist-to-height ratio change) across baseline predicted risk strata (quartiles) using linear regression models, whereby risk strata were coded numerically to test for linear trends. For each adiposity measure and risk score, we fit a model with treatment contrasts (placebo as reference). The model includes main effects for predicted risk strata and treatment arm, and their interaction term that tests whether treatment effects (versus placebo) change linearly across risk quartiles. EMMs for each treatment arm within each stratum were calculated using the emmeans package, with equal weight given to each stratum regardless of sample size. To test for changes in predicted risk from baseline to week 72 across all 18 outcomes, a one-way analysis of variance across treatment arms was employed, followed by unadjusted pairwise contrasts of the group means.
Subgroup and sensitivity analyses
To investigate performance of OBSCORE in different subgroups, we evaluated the performance of OBSCORE separately in individuals of genetically inferred non-European (n = 9,668) and European (sample size matched) ancestries as well as in individuals with the lower 50% and higher 50% of the Townsend deprivation index (median of the study population −2.05), without retraining in the subgroups. To test whether discriminative performance may be influenced by early incident cases that represent delayed diagnosis, we investigated C-indices of OBSCORE when excluding early incident cases additionally at 1 year, 1.5 years and 2 years after baseline for each outcome.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.



