Chinese General Practice ›› 2026, Vol. 29 ›› Issue (26): 3829-3837.DOI: 10.12114/j.issn.1007-9572.2025.0290

• Article·Proactive Health • Previous Articles     Next Articles

Revealing the Relationship between Lifestyle and Sub-health Based on Machine Learning: a Cross-sectional Study in Guangdong Province

  

  1. 1. The Fifth Clinical College of Guangzhou University of Chinese Medicine, Guangzhou 510095, China
    2. Department of Endocrinology, Guangdong Provincial Second Hospital of Traditional Chinese Medicine, Guangzhou 510095, China
    3. Guangdong Provincial Key Laboratory of Research and Development in Traditional Chinese Medicine, Guangzhou 510095, China
    4. Department of Traditional Chinese Medicine, Nanfang Hospital, Southern Medical University, Guangzhou 510515, China
    5. Department of Respiratory, Guangdong Provincial Second Hospital of Traditional Chinese Medicine, Guangzhou 510095, China
  • Received:2025-09-25 Revised:2026-07-05 Published:2026-09-15 Online:2026-08-11
  • Contact: BI Jianlu

基于机器学习揭示生活方式与亚健康之间的关系:一项在中国广东省的横断面研究

  

  1. 1.510095 广东省广州市,广州中医药大学附属第五临床医院
    2.510095 广东省广州市,广东省第二中医院内分泌科
    3.510095 广东省广州市,广东省中医药研究开发重点实验室
    4.510515 广东省广州市,南方医科大学南方医院中医科
    5.510095 广东省广州市,广东省第二中医院呼吸科
  • 通讯作者: 毕建璐
  • 作者简介:

    作者贡献:

    刘阳负责资料收集、数据整理分析及论文初稿撰写;冯小洁负责文献资料筛选;陈洁瑜、罗仁负责论文修订;赵晓山、毕建璐负责论文选题、构思设计、团队组建、全程质量控制及论文审校。

  • 基金资助:
    国家自然科学基金联合基金(U22A20365); 国家自然科学基金资助项目(T2341019); 国家自然科学基金重点项目(81830117); 广州市科技计划项目(2024B03J1343); 广东省中医药局科研项目(20241208); 广东省科技计划项目(2020A1414050029); 广东省医学科学基金资助项目(B2025015)

Abstract:

Background

Sub-health status, as a subclinical and reversible stage between health and chronic disease, has received widespread attention. Previous studies have indicated that various lifestyle factors significantly influence the dynamic changes of sub-health.

Objective

This study aims to explore the association between diverse lifestyle factors and sub-health status using machine learning methods. In addition, it seeks to develop a predictive model for sub-health to enable early identification, diagnosis, and intervention in sub-health populations.

Methods

This study was a cross-sectional survey conducted from September 2017 to January 2024. A multistage cluster sampling method was employed to randomly select survey sites and units in stages from seven representative cities in Guangdong Province, with a total of 20 375 participants ultimately included. Data were collected on-site during routine health examinations by uniformly trained investigators using standardized questionnaires, including the Sub-health Assessment Scale and the Health-promoting Lifestyle ProfileⅡ (HPLP-Ⅱ). For statistical analysis, SPSS 25.0 and Python 3.9 were used for data processing and analysis. Feature selection was first performed using recursive feature elimination, and sub-health prediction models were then developed based on Logistic regression, support vector machine, k-nearest neighbors, and XGBoost. Subsequently, five-fold cross-validation and data balancing techniques were employed to evaluate and optimize model performance, aiming to identify associated factors and predict sub-health status.

Results

A total of 20 375 participants were included. Among them, 2 586 (12.7%) participants were classified as healthy, and 17 789 (87.3%) were classified as having sub-health status. In the sub-health group, there were 8 581 (48.2%) males and 9 208 (51.8%) females. Univariate analysis showed that there were statistically significant differences between the healthy and sub-health groups in sex, age, educational level, marital status, BMI, alcohol consumption, and the scores of all 52 HPLP-Ⅱ items (P<0.05), whereas no statistically significant difference was observed in smoking status (P>0.05). Feature selection results indicated that BMI and sleep quality were the key factors most closely associated with sub-health status. The model constructed based on the XGBoost algorithm demonstrated good performance, with an accuracy of 0.860 and an F1 score of 0.672. SHAP analysis further confirmed that these variables had high contributions to the model and played an important role in predicting sub-health status.

Conclusion

This study indicates that BMI and sleep quality are key factors associated with suboptimal health status, providing a scientific basis for the prevention and management of chronic diseases. Against the backdrop of rapid advances in predictive, preventive, and personalized medicine, early prediction and precise intervention for suboptimal health may facilitate disease state reversal and ultimately improve population health outcomes.

Key words: Sub-health, Lifestyle, Health promotion, Predictive learning models, Machine learning, SHapley Additive exPlanations

摘要:

背景

亚健康是介于健康与慢性病之间、具有可逆性的亚临床状态,现已受到广泛关注。现有研究证实,各类生活方式因素会显著影响亚健康状态的动态变化。

目的

本研究旨在采用机器学习方法探讨多种生活方式与亚健康的关联,同时构建亚健康预测模型,以期实现亚健康人群的早识别、早诊断及早干预。

方法

本研究为横断面调查,开展时间为2017年9月—2024年1月。采用多阶段整群抽样方法,在广东省7个具有代表性的城市中分阶段随机抽取调查点及单位,最终共纳入20 375名单位体检人群。由经过统一培训的调查人员在体检现场使用标准化问卷开展数据收集,包括亚健康评估量表、健康促进生活方式量表Ⅱ(HPLP-Ⅱ)等。采用SPSS 25.0和Python 3.9进行数据处理与分析,首先利用递归特征消除法进行特征筛选,基于逻辑回归、支持向量机、k近邻算法和XGBoost构建亚健康预测模型;随后结合五折交叉验证及数据平衡技术对模型性能进行评估与优化,以分析影响因素并预测亚健康状态。

结果

20 375名研究对象中,健康人群2 586名(12.7%)、亚健康人群17 789名(87.3%)。亚健康人群中,男8 581名(48.2%),女9 208名(51.8%)。单因素分析结果显示,健康人群与亚健康人群的性别、年龄、受教育程度、婚姻状况、BMI、饮酒情况及HPLP-Ⅱ的52个条目得分比较,差异有统计学意义(P<0.05);两者吸烟情况比较,差异无统计学意义(P>0.05)。特征选择结果显示,BMI和睡眠质量是与亚健康状态密切相关的关键因素。基于XGBoost算法构建的模型表现良好,准确率为0.860,F1分数为0.672。夏普利可加性解释(SHAP)分析结果进一步证实,上述变量在模型中具有较高贡献度,对亚健康状态的预测具有重要影响。

结论

本研究表明,BMI和睡眠质量是影响亚健康状态的关键因素,这可以为慢性病的有效预防与管理提供科学依据。在全球预测医学、预防医学及个性化医疗快速发展的背景下,亚健康风险的早期预测与精准干预有望实现疾病状态逆转,从而提升公共健康水平。

关键词: 亚健康, 生活方式, 健康促进, 预测学习模型, 机器学习, 夏普利可加性解释