一覧 Case Studies Survival analysis using propensity score matching and RMST

Survival analysis using propensity score matching and RMST

* This article is based on an actual analysis project; however, in light of the confidentiality agreement (NDA) with the client, while maintaining the framework of the medical and healthcare field, specific details such as disease names and variables have been substantially altered from the actual case. We ask for your understanding in advance.

This case evaluated the efficacy, in terms of “overall survival (OS)” and “progression-free survival (PFS),” of a new molecular-targeted drug compared with standard treatment for patients with advanced solid cancer, after adjusting for patients’ baseline background factors. In clinical observational data, confounding—whereby bias in patient background arises between treatment groups—is unavoidable.

Through this analysis, I aimed to perform confounding adjustment by propensity score matching (PSM) and, further, to provide a robust evaluation method for cases where the “proportional-hazards assumption” in survival analysis is not satisfied. By making full use of advanced statistical methods such as RMST (restricted mean survival time), Dr.DataScience provided a precise efficacy evaluation with little bias and contributed to building evidence for treatment selection in actual clinical practice.

Background and objective

In this case, what was required was to clarify whether there was a statistically significant difference in survival time between the new molecular-targeted drug (intervention group) and standard treatment (control group). The main objective of the analysis was to evaluate the influence of the difference in treatment on survival and to quantitatively clarify differences in treatment effect by subgroup, such as “the expression level of a particular biomarker (high/low)” and “the degree of tumor progression.”

Data and variables

This analysis used observational-study data provided by a particular medical institution. The main variables analyzed were as follows.

    • Response variables: overall survival (OS) and progression-free survival (PFS), and the presence or absence of each respective event. The observation period was set at a maximum of 36 months.
    • Main explanatory variable (intervention): use of the new molecular-targeted drug vs. use of standard treatment.
    • Adjustment and stratification variables: biomarker expression level, tumor progression, age, sex (in the order male, female), ECOG performance status, BMI, etc.

Analytical methods

In this case, to adjust for confounders between groups and to address the constraints of the mathematical assumptions in survival analysis, I selected and applied the following statistical methods.

  1. Adjustment for confounders (PSM)
    • Adopted method: propensity score matching (PSM)
      I calculated propensity scores using logistic regression and performed 1:1 matching on samples whose score difference was 0.2 or less. To evaluate covariate balance after matching, I adopted not only p-values but also the standardized mean difference (SMD), assessing imbalance against the guideline that the absolute value of the SMD be less than 0.1.
    • Method not adopted: reliance on p-values alone
      In evaluating the balance of background factors, I refrained from using test p-values alone. When the sample size is small, p-values tend to become significant with slight variations, and conversely, when the sample size is large, even a clinically meaningless difference is judged significant; I therefore judged that the SMD, which shows substantive balance, should be emphasized.
  2. Group comparison of background factors
    • Adopted method: Fisher’s exact test
      I adopted it for comparing categorical variables when creating the post-matching background-factor table.
    • Method not adopted: chi-squared test
      Initially I intended to prioritize the chi-squared test, but because cells with an expected frequency of less than 5 exceeded 20% of the total in the dataset, the conditions for its application were not met, so I judged it inappropriate and did not adopt it.
  3. Survival analysis and verification of proportional hazards
    • Adopted methods: log-rank test and Cox proportional hazards model (general model), Schoenfeld residuals
      I adopted the log-rank test and the Cox proportional hazards model for the group comparison of survival time. At the same time, to confirm whether the Cox model’s premise—that “the hazard ratio is constant over time (proportional hazards)”—was satisfied, I performed a test using Schoenfeld residuals.
  4. Response when proportional hazards break down
    • Adopted method: RMST (restricted mean survival time)
      In some subgroup analyses, it was found that the proportional-hazards assumption was problematic (e.g., the survival curves crossed). Therefore, I adopted the RMST model, which calculates and compares the area under the survival curve up to a specific time point (36 months in this case), and tested the difference and ratio of the RMST.
    • Method not adopted: interpretation by the Cox model’s hazard ratio alone
      For data where proportional hazards are not satisfied, determining treatment effect by the Cox model’s hazard ratio (HR) alone risks overestimating or underestimating the effect, so I refrained from adopting it.

Overview of the main results and clinical considerations

In the analysis after adjusting for background factors by PSM, the intervention group (new molecular-targeted drug) showed a significant prolongation of survival compared with the control group for both overall survival (OS) and progression-free survival (PFS). The significance probability of OS by the log-rank test was p = 0.003, and that of PFS was p = 0.001, confirming the effect with extremely high reliability.

Notably, although the proportional-hazards assumption was not satisfied in some datasets, the analysis at the 36-month time point using RMST proved that the intervention group’s RMST was significantly longer (with a significant difference in both the difference and the ratio of RMST), demonstrating that the treatment’s efficacy was robust.

In addition, the subgroup analysis found a particularly strong treatment effect in the patient group with high expression of a particular biomarker. On the other hand, in the biomarker-low-expression group, the disappearance of a significant difference after PSM (p = 0.056) and an attenuation of the hazard ratio were observed; from the standpoint of stratified medicine, I was able to present clear risk factors as to which patient profiles the treatment should be prioritized for.

Dr.DataScience’s contribution

In this case, Dr.DataScience avoided the traps of the assumptions lurking in survival analysis and created robust evidence directly linked to decision-making in the clinical setting.

  1. Rigorous verification of assumptions and optimal method selection
    • Rather than mechanically applying the Cox model in survival analysis, I thoroughly verified proportional hazards using Schoenfeld residuals. By not overlooking cases where the assumption broke down and switching to the alternative, robust evaluation method of RMST, I derived unbiased, objective evidence.
  2. Ensuring the quality of PSM and multifaceted balance evaluation
    • Avoiding evaluation that relies on p-values alone, by using the SMD (standardized mean difference) to objectively evaluate covariate balance, I achieved true confounding adjustment free from the influence of sample size. For the bias in expected frequencies in small data, I switched to Fisher’s exact test—pursuing statistical validity down to the finest detail.
  3. Clear interpretation of the results and provision of clinical implications
    • By tracking in detail the transition of the hazard ratio from univariate to multivariate analysis (the influence of confounding), I identified the true independent risk factors. This provided valuable insights directly linked to narrowing down the patients indicated for treatment in actual clinical practice—going beyond a mere “there was/was not an effect.”

© Dr.データサイエンス. All Rights Reserved.