一覧 Case Studies Identifying influencing factors from multiple angles with canonical correlation analysis and analysis of variance

Identifying influencing factors from multiple angles with canonical correlation analysis and analysis of variance

In this case study, I describe how Dr.DataScience used advanced statistical methods to help a client who wished to clarify, within the diverse medical data they held, the latent associations among multiple symptom groups or laboratory-value groups, or the influence that particular patient attributes or treatment interventions have on clinical outcomes.

In the medical setting, a wide variety of information is collected daily—patient attributes, medical history, subjective symptoms, test results, treatment content, and more. Deeply understanding the relationships behind these complex data groups, and the influence individual factors have on outcomes, is essential to elucidating pathology, improving diagnostic accuracy, and optimizing treatments.

By taking the characteristics of medical data into account and deriving statistically valid and clinically meaningful insights, Dr.DataScience helped the client make data-driven decisions with greater confidence. In accordance with our confidentiality agreement, no specific disease names, treatments, or patient information are disclosed; however, the methods used, the process, and the type of findings obtained are the same as in the actual analysis.

Background and objective

The client wished to deeply understand, for a certain disease’s patient data, how multiple symptom patterns and laboratory-value groups are related to one another, and to what degree particular patient background factors or treatment choices influence particular clinical outcomes.

In particular, they aimed to analyze, in a statistically meaningful way, a dataset containing much qualitative information (e.g., types of subjective symptoms, treatment options, etc.), and to obtain objective evidence useful for formulating diagnostic and treatment strategies. Revealing the complex relationships among many variables—relationships that tend to be overlooked by conventional simple analyses—was an important objective of this analysis.

Data and variables

This analysis used anonymized, large-scale medical survey data. The subjects were data from a patient cohort collected for a particular purpose. Specifically, the following main types of variables were included.

  1. Questionnaire items (variables to be analyzed)
    • Items consisting of text-based categorical responses concerning patients’ subjective symptoms, lifestyle, medical history, treatment options, and the like (e.g., “yes,” “no,” “frequently,” “occasionally,” etc.). These qualitative data were preprocessed by converting them into numbers for the analysis.
  2. Particular clinical-outcome variables
    • Categorical indicators whose influence we wished to evaluate, such as the presence or absence of a particular symptom, the stage of treatment response, or the presence of complications (e.g., “improved / not improved,” “occurred / did not occur,” etc.).

Analytical methods

  1. Data cleaning and preprocessing
    • Because the questionnaire responses were entered as text-based categorical values, I created dummy variables (categorical values converted into numbers) to put them into a form suitable for statistical analysis.
    • I also ensured data consistency, taking into account cases where the number of dummy variables linked to each questionnaire item was multiple and uneven.
  2. Canonical correlation analysis
    • ・Pre-analysis check
      • Prior to the canonical correlation analysis, I checked for multicollinearity within and between each set of variables, taking care that redundant variables would not affect the interpretation of the results.
      • In the data after dummy-variable conversion, I checked whether the sample size needed for the analysis was sufficiently secured, and whether responses were extremely concentrated in particular categories (data bias). I paid particular attention because correlation coefficients tend to come out strong when an original questionnaire item has many kinds of values.
    • ・Analysis of the correlations
      • I grouped the questionnaire items concerning patients’ various symptoms, conditions, and treatments (the converted dummy variables) into “variable sets (packages),” each composed of multiple variables.
      • To analyze the correlations among these variable sets, I used a method called canonical correlation analysis. This analysis makes it possible to group the created dummy variables into a single variable set and to evaluate the correlations between such sets.
      • The canonical correlation analysis yields canonical correlation coefficients. As with ordinary correlation coefficients, these can be interpreted such that the value approaches 1 when the correlation between variable sets is highest and approaches 0 when the correlation is low.
    • ・Interpreting the results
      • In addition to the obtained canonical correlation coefficients, I examined the standardized canonical loadings (weights) and cross-loadings of the original questionnaire items on each canonical variate. This allowed me to interpret in detail which questionnaire items contributed most to forming each canonical variate, and between which specific items the association across variable sets was strong.
      • I also computed redundancy coefficients to evaluate how much of the variance of one variable set was explained by the canonical variates of the other, deepening the practical significance of the analysis results.
  3. Multivariate analysis of variance (MANOVA) and analysis of variance (ANOVA)
    • ・Pre-analysis check
      • I confirmed the assumptions of analysis of variance—namely, the normality and homogeneity of variance of the response variable in each group—both visually (histograms, box plots) and statistically (e.g., the Shapiro–Wilk test, Levene’s test, etc.). For multivariate analysis of variance in particular, I also confirmed multivariate normality and the homogeneity of the variance-covariance matrices (e.g., Box’s M test).
    • ・Influence analysis
      • To analyze how much each questionnaire item (explanatory variable) indicating a patient’s background or experience influenced the particular clinical outcome (response variable) of special interest to the client, I performed multivariate analysis of variance.
      • This analysis makes it possible to evaluate, for multiple response categories of the particular clinical outcome, how much statistical influence is exerted by the presence or absence of response data within each questionnaire item.
    • ・Post-hoc tests and effect sizes
      • Where the analysis of variance found an overall statistically significant difference, I performed post-hoc tests to identify specifically which categories differed. When making multiple comparisons, I applied multiple-comparison adjustments (e.g., Tukey’s HSD, Bonferroni correction) to prevent the detection of false significant differences.
      • In addition to p-values, I calculated effect sizes. This allowed me to evaluate not only statistical significance but also the clinical importance and substantive magnitude of the detected influence, deepening the interpretation of the results.

Overview of the main results and clinical considerations

This analysis revealed, through canonical correlation analysis, that statistically significant associations exist among multiple groups of medical-related questionnaire items. This suggested that multiple symptoms or lifestyle factors that appear unrelated at first glance may in fact be linked by a common latent structure or patient characteristic.

In addition, the analysis of variance revealed that several patient background factors and experiences had a statistically significant influence on particular clinical outcomes. For example, the direction and strength of specific influences were quantitatively shown—such as that a patient group that had previously undergone a particular treatment showed a different tendency in the improvement of a certain symptom compared with a group that had not.

These findings provide the client with solid, data-based evidence for refining diagnostic criteria, individualizing treatment strategies, improving patient-education programs, or optimizing clinical-research design. By identifying clinically meaningful patterns and influencing factors within complex medical data, it became possible to indicate a path toward more effective medical interventions.

Dr.DataScience’s contribution

This case clearly demonstrates the ability of Dr.DataScience’s statistical-analysis specialist to delve deeply into the essence of complex medical survey data—including text-based responses—held by a client, and to draw out clinically valuable insights.

We first thoroughly carried out the seemingly complex preprocessing of converting qualitative data into dummy variables and building multiple questionnaire items into meaningful variable sets. On this foundation, by combining two advanced statistical methods—canonical correlation analysis to explore the associations among multiple variable sets, and multivariate analysis of variance to evaluate the influence of each factor on a particular clinical outcome—we made it possible to capture both the overall picture of the data and the individual influences from multiple angles.

Through this approach, we drew out the maximum latent value of the client’s data and provided reliable scientific evidence leading to the solution of clinical problems. Through detailed analysis of complex survey data in the medical field, Dr.DataScience powerfully supports the client’s clinical problem-solving and solid, data-driven decision-making.

© Dr.データサイエンス. All Rights Reserved.