一覧 Case Studies Clarifying the treatment effects of new anticancer drugs with the Cox model and forest plots

Clarifying the treatment effects of new anticancer drugs with the Cox model and forest plots

In this case study, I describe how Dr.DataScience used advanced survival analysis and effective data-visualization methods to help a client who wished to objectively evaluate the treatment effects of two anticancer drugs and, further, to clarify how those effects differ according to patient characteristics (subgroup effects). In cancer treatment, accurately evaluating a drug’s efficacy and identifying which drug is optimal for which patients are directly linked to improving patients’ prognoses.

Through a multifaceted approach including univariate analysis, multivariate analysis, and subgroup analysis, Dr.DataScience derived statistically robust insights and expressed them as visually excellent forest plots, helping the client make data-driven decisions with greater confidence. In accordance with our confidentiality agreement, no specific disease names or figures are disclosed; however, the analysis and figures produced, and the type of findings obtained, are the same as in the actual case.

Background and objective

Using data collected in a clinical trial, the client wished to comparatively evaluate the treatment effects of two anticancer drugs (Abcdefghmab and Jklmnopmab) and to clarify their influence on a particular clinical outcome (e.g., overall survival, progression-free survival, etc.).

In addition to an overall evaluation of treatment effects, an important objective was to verify in detail whether the treatment effects of the two drugs differed according to various background factors such as patient age, disease stage, and the presence of comorbidities (the presence of subgroup effects). The aim was to establish guidance for more effective drug selection and to obtain objective information useful for advancing individualized medicine based on patient stratification.

Data and variables

This analysis used an anonymized clinical-trial dataset. The subjects were data from a patient cohort with a particular cancer type. Specifically, the following main types of variables were included.

  1. Treatment groups
    • Abcdefghmab group vs. Jklmnopmab group
  2. Response variables
    • Time until an event occurs (time variable): the number of days or the period until a particular clinical outcome (e.g., death, disease progression, etc.) occurs.
    • Whether the event occurred (event variable): a binary variable indicating whether the event of interest occurred during follow-up (e.g., 0 = no event/censored, 1 = event).
  3. Covariates (explanatory variables)
    • Multiple continuous and categorical variables thought to potentially influence the outcome or treatment effect, such as patient attributes (age, sex, etc.), disease stage, particular biomarker values, medical history, presence of comorbidities, and prior treatment history.

Analytical methods

  1. Data cleaning and preprocessing
    • I checked the data quality and preprocessed the data into the form optimal for analysis, including handling missing values unsuited to analysis and converting categorical variables into dummy variables.
  2. Cox proportional hazards model
    • For the survival data whose response variable was time-to-event, I performed a Cox proportional hazards model to evaluate how the principal comparison factor (treatment group) and multiple covariates influence it.
    • I examined in detail the model’s assumption of proportional hazards (that the influence of each factor is constant over time) using methods such as the Schoenfeld residuals test, and verified the validity of the model.
    • Through the analysis, I calculated the degree to which each explanatory variable influences the risk of the target event as a hazard ratio (HR), and confirmed statistical significance (p-values).
    • ・Univariate analysis
      • First, to evaluate the influence that each factor (the treatment group and individual covariates) has on the risk of the clinical outcome on its own, I performed individual Cox proportional hazards models.
    • ・Multivariate analysis
      • Next, I entered multiple factors (the treatment group and the covariates to be adjusted for) into the model simultaneously and performed a multivariate Cox proportional hazards model that evaluates each one’s independent influence after adjusting for the others. This aimed to eliminate the influence of confounders and obtain a pure estimate of the treatment effect.
  3. Subgroup analysis and creation of forest plots
    • To verify the possibility that the treatment effect differs by patient characteristics, I performed subgroup analyses evaluating the treatment-group hazard ratio for each predefined subgroup, such as age group, disease stage, and the status of particular biomarkers.
    • I visually expressed the hazard ratios and 95% confidence intervals of each factor (mainly the treatment group) obtained from these subgroup analyses as a forest plot. This figure makes it possible to intuitively understand the direction of the treatment effect in each subgroup and the certainty of that effect (the width of the confidence interval).

Forest plot

As requested, to clearly express the results of the subgroup analysis by the Cox proportional hazards model, I created a forest plot. This forest plot can intuitively show differences in treatment effect according to patient characteristics—for example, “in the ISS Stage III subgroup, Jklmnopmab reduces the risk of the event statistically significantly more than Abcdefghmab.”

(Figure: Example of a forest plot)
  1. How to read a forest plot
    • ・The central vertical line (the null-hypothesis line): This line is drawn at the position where the hazard ratio is 1.0. A hazard ratio of 1.0 means that the factor (here, the drug Jklmnopmab compared with Abcdefghmab) has no influence at all on the risk of the event.
    • ・The point for each subgroup (the black dot): This represents the estimate of the hazard ratio (effect size) in each subgroup. If this point is to the right of the central vertical line, the Jklmnopmab group has a higher risk of the event than the Abcdefghmab group; if to the left, the risk is lower.
    • ・The horizontal line for each subgroup (the 95% confidence interval): The horizontal line through the point indicates the 95% confidence interval of that hazard ratio. The wider this interval, the greater the variability of the estimate and the lower its reliability.
    • ・The relationship between the confidence interval and the central vertical line: If the horizontal line crosses the central vertical line (hazard ratio 1.0), the treatment effect in that subgroup (the difference between the two drugs) is interpreted as not statistically significant (p > 0.05). If the horizontal line does not cross the central vertical line, the treatment effect in that subgroup is interpreted as statistically significant (p < 0.05).

Overview of the main results and clinical considerations

Through this analysis—and the forest-plot visualization in particular—the client was able to clearly grasp not only the overall treatment effects of the two anticancer drugs (Abcdefghmab and Jklmnopmab) but also the differences in treatment effect within particular patient subgroups. The univariate and multivariate analyses showed the overall tendencies, and the forest plot delved further, visually showing which drug is more effective in particular patient groups, or whether there is no difference in effect.

These visual insights are extremely useful information, as they allow the principal factors and the associated risk or protective effects to be grasped intuitively without having to decipher complex statistical tables.

Dr.DataScience’s contribution

This case demonstrated that Dr.DataScience can present the statistical insights obtained from a client’s medical data not only as support for decision-making in the clinical setting but also as a strength in publication as an academic paper. Through rigorous, highly reproducible statistical methods (univariate and multivariate Cox proportional hazards models, including rigorous verification of the assumptions) and the presentation of results via intuitive, visually persuasive forest plots, Dr.DataScience ensures high reliability in the peer-review process and contributes to maximizing the clinical and academic impact of the research. In this way, I powerfully support the client’s valuable data being shared with the world in a form backed by solid scientific evidence.

To evaluate the treatment effects of the two anticancer drugs, I first performed appropriate univariate and multivariate Cox proportional hazards model analyses. I carefully confirmed the model’s assumption of proportional hazards, and then expressed the hazard ratios and confidence intervals of each factor obtained from the analysis as an intuitive, easy-to-understand forest plot, helping the client quickly understand complex statistical data and accelerate data-driven decision-making. Dr.DataScience demonstrates expertise not only in conducting analyses but also in presenting the results in a practical form, powerfully advancing the client’s clinical problem-solving and the use of data in the medical setting.

© Dr.データサイエンス. All Rights Reserved.