Chinese General Practice

    Next Articles

Development and Analysis of a Hypertension Risk Prediction Model for Tibetan Older Adults in the Tibet Region: Based on Data from the National Health Services Survey in Tibet

  

  1. 1.Medical College, Xizang University, Lhasa 850000, China 2.Lhasa Key Laboratory of Public Health Safety and Health Policy Research (Xizang University), Lhasa 850000, China 3.Plateau Health Science Research Center, Xizang University, Lhasa 850000, China
  • Received:2025-12-15 Revised:2026-03-10 Accepted:2026-03-27
  • Contact: Zhaxi Dawa; E-mail: zhaxi0891@qq.com

西藏地区藏族老年人高血压患病风险预测模型的建立与研究:基于国家卫生服务调查西藏数据

  

  1. 1.850000 西藏自治区拉萨市,西藏大学医学院 2.850000 西藏自治区拉萨市公共卫生安全与卫生政策研究重点实验室(西藏大学) 3.850000 西藏自治区拉萨市,西藏大学高原健康科学研究中心
  • 通讯作者: 扎西达娃;E-amil:zhaxi0891@qq.com
  • 基金资助:
    第七次西藏自治区卫生服务调查扩点部分项目 (18080278)

Abstract: Background Hypertension is a well-recognized major risk factor for cardiovascular and cerebrovascular diseases. In plateau regions, unique environmental and sociocultural factors may lead to differences in the epidemiology, prevention, and control priorities of hypertension compared with plain areas. However, there is currently a lack of a specific hypertension risk identification tool for Tibetan older adults in the Tibet Autonomous Region. Objective Based on data from the National Health Services Survey, to develop a hypertension risk prediction model for this population, and to provide a potential auxiliary tool for community-based preliminary screening and targeted health management of hypertension risk. Methods A cross-sectional study design was used. A total of 2 140 Tibetan older adults aged ≥ 60 years from the Tibet region were included. Variables were first screened by univariate analysis, and then a multistage feature selection strategy (LASSO regression combined with multivariate Logistic regression) was used to determine the final predictive variables. Based on the selected variables, five machine learning models were constructed and compared for their diagnostic value in predicting hypertension among Tibetan older adults: LASSO regression, RandomForest, XGBoost, LightGBM, and NaiveBayes. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, and specificity. Clinical utility was further assessed using calibration curves, decision curve analysis, and clinical impact curves. Results Among the 2 140 Tibetan older adults included, 926 had hypertension, yielding a prevalence of 43.27%. The participants were randomly divided into a training set (n=1 499) and a test set (n=641) at a 7:3 ratio. In the training set, 649 participants had hypertension (prevalence 43.30%), and in the test set, 277 participants had hypertension (prevalence 43.21%); the difference in prevalence between the two sets was not statistically significant (χ2 =0.001, P=0.972). Multivariate Logistic regression identified 8 predictive factors significantly associated with hypertension status: different cities, BMI category, age, insomnia symptoms within the past 2 weeks, number of other chronic diseases, daily tooth brushing frequency, social medical insurance, and employment status. Model comparison showed that NaiveBayes performed best in the test set, with an AUC of 0.640, accuracy of 59.8%, sensitivity of 54.9%, and specificity of 63.5%. A nomogram was constructed based on the final 8 factors. Conclusion Based on cross-sectional data, this study developed a hypertension risk prediction model for Tibetan older adults incorporating 8 factors. The NaiveBayes model showed limited discriminative ability in the test set, with an AUC of 0.640, accuracy of 59.8%, sensitivity of 54.9%, and specificity of 63.5%. The model and nomogram may provide a reference for community-based hypertension screening in this population, but their diagnostic performance and practical utility require further validation in prospective studies.

Key words: Hypertension, Machine learning, Predictive learning models, Tibetan nationality, Aged, Tibet

摘要: 背景 高血压是已知的心脑血管疾病主要危险因素之一。在高原地区,独特的自然环境和社会文化因素可能使得高血压的流行状况与防治重点与平原地区有所不同。然而,目前缺乏针对西藏自治区藏族老年人群的特异性高血压风险识别工具。目的 基于国家卫生服务调查数据,构建适用于该人群的高血压风险预测模型,旨在为西藏地区社区开展高血压风险初筛和针对性健康管理提供一种可能的辅助工具。方法 采用横断面研究设计,纳入西藏地区≥60岁藏族老年人共2 140人。首先通过单因素分析筛选变量,随后采用多阶段特征选择(LASSO回归结合多因素Logistic回归)确定最终预测变量。基于筛选出的变量,构建并比较了LASSO回归、RandomForest、XGBoost、LightGBM及NaiveBayes共5种机器学习模型对藏族老年人高血压患病的诊断价值。使用受试者工作特征曲线下面积(AUC)、准确率、灵敏度、特异度等指标评价模型性能,并通过校准曲线、决策曲线分析及临床影响曲线评估其临床实用性。结果 纳入的2 140名藏族老年人中高血压患者926例,高血压患病率为43.27%。按7∶3比例随机分为训练集(n=1 499)和测试集(n=641)。训练集中高血压患者649例(高血压患病率为43.30%),测试集中高血压患者277例(高血压患病率为43.21%),两组患病率比较,差异无统计学意义(χ2=0.001,P=0.972)。通过多因素Logistic回归分析,最终确定了8个与高血压状态显著相关的预测因素:不同地市、BMI分级、年龄、两周内失眠症状、其他慢性疾病种数、每天刷牙次数、社会医疗保险、就业状况。模型比较结果显示,NaiveBayes模型在测试集上表现最佳,其AUC为0.640,准确率为59.8%,灵敏度为54.9%,特异度为63.5%。基于最终确定的8个因素构建了列线图。结论 本研究基于横断面数据,构建了一个包含8个因素的藏族老年人高血压风险预测模型。NaiveBayes模型在测试集中AUC为0.640、准确率59.8%、灵敏度54.9%、特异度63.5%,展现出有限的区分能力。该模型及诺莫图工具或可为该人群的高血压社区筛查提供参考,但其诊断效能与实用性仍需在前瞻性研究中进一步验证。

关键词: 高血压, 机器学习, 预测学习模型, 藏族, 老年人, 西藏

CLC Number: