一覧 Case Studies Cluster analysis and factor analysis based on symptom profiles and physiological indicators

Cluster analysis and factor analysis based on symptom profiles and physiological indicators

This case applied the multivariate methods of cluster analysis and factor analysis to deeply understand the diversity of disease within a group of patients with a particular cardiovascular disease. In chronic diseases—especially cardiovascular disease—even with the same diagnostic name, symptom profiles, physiological indicators, and responses to treatment often differ greatly from patient to patient.

Categorizing this patient diversity and identifying latent subtypes is essential for advancing individualized medicine and formulating more effective treatment strategies. Dr.DataScience analyzed the patients’ complex clinical data and made it possible to classify patients from a new perspective.

In accordance with our confidentiality agreement, no specific figures or detailed medical background are disclosed.

Background and objective

The client believed that, even among patients with a particular cardiovascular disease who present with seemingly similar symptoms, there is latent diversity in the combinations of symptoms and the patterns of laboratory values. What was required was to classify this diversity into statistically distinct subtypes and to understand the characteristics of each subtype, with the aim of obtaining important insights leading to predicting disease progression, stratifying treatment responsiveness, and even optimizing the design of future clinical trials.

Through this analysis, what was required was to clarify what distinct subtypes the patient group falls into based on symptom profiles and physiological indicators, what latent axes of symptoms or physiological characteristics characterize these subtypes, how the presence or absence of a particular medical history affects patient classification and characteristics, and how the resulting subtypes can help in formulating individualized treatment strategies.

Data and variables

The data used in the analysis of this case were the clinical records and laboratory data of patients with a particular cardiovascular disease.

    • Symptom score: a quantitative score (5-point scale) that comprehensively evaluates multiple principal symptoms reported by the patient (e.g., degree of shortness of breath, presence of edema, frequency of chest pain, fatigue, etc.).
    • Biomarker score: a quantitative score (5-point scale) that comprehensively evaluates each patient’s multiple physiological laboratory values and vital signs (e.g., BNP level, heart rate, ejection fraction, blood pressure, etc.).
    • Presence or absence of medical history: a categorical variable indicating the presence or absence of the patient’s principal medical history or comorbidities (e.g., “no medical history,” “has medical history,” “unknown”).

Analytical methods

In this case, I performed the following analyses regarding the subtype classification of patients with a particular cardiovascular disease.

  1. Preliminary analysis
    • I tabulated the distribution of the “presence or absence of medical history” attribute of the target patients to grasp the background of the patient group.
    • To check whether the score distributions of the main evaluation items, “symptom score” and “biomarker score,” followed a normal distribution, I performed the Shapiro–Wilk test (significance level 0.05) and judged the appropriateness of the analytical methods.
  2. Cluster analysis
    • For 2,423 patient records, to classify patients in terms of “symptom score” and “biomarker score,” I adopted cluster analysis using the non-hierarchical k-means method.
    • To determine the number of clusters, I adopted the elbow method, setting the optimal number of clusters at five, at the point where the trend of the within-cluster sum of squares converged.
    • I interpreted the characteristics of each cluster in detail by the combination of “symptom score” and “biomarker score,” identifying patient groups with particular symptom patterns or physiological characteristics.
  3. Factor analysis
    • To further quantitatively evaluate the classification obtained from the cluster analysis and to extract latent evaluation axes, I performed factor analysis on the individual items that make up the symptom score and the biomarker score (e.g., shortness of breath, chest pain, BNP level, heart rate, etc.).
    • The number of factors was determined by considering the Kaiser–Guttman criterion (eigenvalue of 1 or more) and the rate of increase in the cumulative proportion of variance explained.
    • I interpreted the identified factors in clinical terms and, on a factor-score plot, visualized the position of each patient on latent axes such as “symptom severity” and “degree of cardiac dysfunction.” This made it possible to more clearly identify patient subtypes with composite characteristics.
    • Considering the differences in characteristics by the patient’s “presence or absence of medical history,” I performed cluster analysis and factor analysis separately for each group: “no medical history,” “has medical history,” and “unknown.”

Overview of the main results and clinical considerations

As a result of the analysis, patients with a particular cardiovascular disease were classified, based on the combination of “symptom score” and “biomarker score,” into five distinct subtypes (clusters) with different characteristics. For example, a severe type with heavy symptoms and worsened biomarkers, a latent-risk type with mild symptoms but abnormalities in biomarkers, and a stable type with both symptoms and biomarkers stable were identified. By analyzing these clusters in detail, the clinical characteristics of each subtype and the need for treatment intervention became clear.

Furthermore, the factor analysis extracted multiple latent factors behind the patients’ symptoms and physiological indicators, such as “disease severity” and “compensatory mechanism.” From the factor-score plot, it became possible to quantitatively identify patient groups that had tended to be overlooked—for example, those whose “symptoms appear minor but whose physiological indicators are in fact worsening.” This highlights a patient diversity that cannot be fully addressed by uniform diagnosis or treatment, providing important grounds for formulating more individualized treatment strategies.

In addition, the subgroup analysis by “presence or absence of medical history” suggested that patients with and without a medical history have different characteristics in how symptoms appear and in their biomarker patterns, highlighting the need for clinical guidelines and treatment approaches specialized to each stratum.

These findings deepen the understanding of the pathology of cardiovascular disease and not only enable appropriate risk assessment and treatment selection for each patient stratum, but also contribute to enhancing the power to detect treatment effects in future clinical trials by targeting more homogeneous patient groups.

Dr.DataScience’s contribution

In this case, Dr.DataScience provided advanced data analysis to reveal the diversity of disease from the complex clinical data of patients with cardiovascular disease. Its main contributions were as follows.

  1. Patient stratification through multivariate analysis
    • Going beyond simple descriptive statistics or simple comparisons, by introducing the advanced multivariate methods of cluster analysis and factor analysis, I categorized patients from multiple aspects—”symptoms” and “biomarkers”—and elucidated the latent structure behind them.
    • This provided an objective, data-driven approach to patient classification, which had tended to rely on rules of thumb.
  2. Contribution to individualized medicine
    • Through factor analysis, I clarified how each patient is positioned on the latent factor axes and quantitatively identified patient subtypes with particular composite characteristics.
    • These results provide the client with clear grounds for formulating individualized treatment strategies tailored to each patient’s characteristics, greatly contributing to the advancement of individualized medicine.
  3. Deepening clinical implications
    • The subgroup analysis considering an important background factor such as “presence or absence of medical history” highlighted differences in the patterns of symptom onset and physiological responses according to patient background, providing more detailed and practical clinical implications.

By fusing statistical expertise with a deep understanding of the medical field, Dr.DataScience contributed to data-based patient management and the optimization of treatment strategies, helping to improve the quality of care in the field of cardiovascular disease.

© Dr.データサイエンス. All Rights Reserved.