* This article is based on an actual analysis project; however, in light of the confidentiality agreement (NDA) with the client, while maintaining the framework of the medical and healthcare field, specific details such as disease names and variables have been substantially altered from the actual case. We ask for your understanding in advance.
This case evaluated the association between “daily training implementation status” and “indicators of motor-function recovery” over several months in functional-recovery training after cerebrovascular disease. In long-term observational data in actual clinical practice, it is common for the duration of attendance and the number of valid months of records to vary from patient to patient.
Through this analysis, rather than simply comparing each patient’s mean values, I aimed to weight the data by “the length of the period over which data were recorded (the amount of information)” as the reliability of the data, and to extract more accurate correlations closer to the reality of actual clinical practice. By combining weighted correlation analysis with detailed subgroup analysis, Dr.DataScience derived robust objective grounds from highly variable real clinical data and contributed to formulating effective exercise-therapy plans.
Background and objective
In this case, what was required was to clarify, in a patient group discharged from a recovery-phase ward, “the implementation frequency of which training content (gait training, upper-limb function training, etc.) is strongly correlated with the maintenance and improvement of post-discharge quality of life (QOL) and motor-function evaluation.” The main objective of the analysis was to evaluate, in an integrated manner, patient data with differing valid observation periods, and to identify effective intervention factors overall and in particular patient strata (subgroups such as disease severity and age group).
Data and variables
This analysis used per-patient monthly training records and periodic evaluation data collected from multiple medical facilities. The main variables analyzed were as follows.
- Evaluation variables 1 (intervention indicators): monthly mean values such as “the number of times various training was implemented” and “the time spent on self-directed training.”
- Evaluation variables 2 (clinical evaluation indicators): monthly mean values such as “motor-function evaluation score (FIM, etc.)” and “quality of life (QOL) evaluation indicators.”
- Weighting variable: each patient’s “number of months for which data recording was validly carried out (observation period).”
- Stratification variables (subgroups): disease severity, age group, presence of comorbidities, etc.
Analytical methods
In this case, to perform a precise correlation evaluation that accounts for the differences in each patient’s observation period (differences in data reliability), I selected and applied the following statistical methods.
- Evaluation of correlations
- Adopted method: weighted correlation analysis
When calculating correlations by averaging several months of data per patient, it is inappropriate to treat the mean value of a patient with only one month of data the same as the mean value of a patient with 12 months of data. In this analysis, I performed weighted correlation analysis adopting “the number of valid recorded months” as the weight. This made it possible to emphasize the trends of patients for whom data were obtained stably over the long term, and to suppress spurious correlations due to short-term, chance variation. - Methods not adopted: ordinary Pearson correlation analysis or Spearman’s rank correlation coefficient
Ordinary correlation analysis treats all subjects (patients) as a single, equal data element. Because the information about “data reliability”—the length of the observation period—is completely lost, and there is a risk that the analysis results are strongly pulled by the data of a patient who happened to produce extreme values in just one month (overestimating the influence of outliers), I judged it inappropriate as an aggregation process for the kind of uneven longitudinal data in this case and refrained from adopting it.
- Conducting subgroup analysis
- Adopted method: comprehensive subgroup analysis by stratification
Not only the overall correlation, but for all-pattern subgroups by the combination of “severity,” “age group,” and “comorbidities,” I iteratively performed weighted correlation analysis.
- Ensuring significance and reliability
- I calculated not only the strength of the correlation coefficient but also the p-value by a t-test accounting for the degrees of freedom (sample size) and the upper and lower limits of the 95% confidence interval, and extracted only strong associations satisfying an absolute correlation coefficient of 0.45 or more, p < 0.050, and a certain sample size.
Overview of the main results and clinical considerations
As a result of the weighted correlation analysis, significant correlations were confirmed between particular interventions and clinical indicators. For example, in the overall analysis, a moderate or stronger positive weighted correlation was found between “the time spent on self-directed training” and “the motor-function evaluation score” (p = 0.012).
Furthermore, the subgroup analysis revealed that this correlation was not pronounced in the “high-severity patient group,” while it showed a very strong positive correlation (R > 0.65, p = 0.003) in the “patient group of moderate severity and without particular comorbidities.”
Items that, in ordinary correlation analysis (without weighting), did not reach significance because they were buried in the error of patients with little data (p = 0.085, etc.) were brought to the surface as significant associations (p = 0.024) by applying weighting, successfully capturing the true trends of patients who had continued training over a long period.
Dr.DataScience’s contribution
In this case, Dr.DataScience created highly precise objective grounds by adding statistically valid correction to the challenge peculiar to real clinical data—”missingness and variability in observation periods.”
- Optimal method selection accounting for data reliability
- Rather than relying on simple Pearson correlation of mean values, by selecting “weighted correlation analysis,” which incorporates the amount of information (number of recorded months) as a weight, I prevented distortion of the analysis results by a small amount of unstable data and derived reliable results close to clinical intuition.
- Providing implications through comprehensive subgroup analysis and conditional extraction
- I built an analytical procedure that comprehensively and automatically processes correlation analysis for dozens of subgroup patterns combining multiple background factors. By rigorously extracting conditions on three criteria—the strength of the correlation coefficient, the p-value, and the sample size—I provided valuable insights directly linked to individualized medicine: “which intervention should be recommended for which patient stratum.”