一覧 Case Studies Nonparametric (Mann–Whitney U test) comparative analysis for non-normal distributions

Nonparametric (Mann–Whitney U test) comparative analysis for non-normal distributions

This case carried out a detailed comparative examination of operation times (decision time and activity time) between cases of differing urgency, in order to evaluate efficiency in the emergency-medical-response setting. In a medical setting that demands prompt and appropriate response, making optimal decisions and taking optimal action within limited time is extremely important, and establishing efficient response protocols suited to the characteristics of each case is an urgent challenge.

Through this analysis, I aimed to quantitatively grasp the characteristics of response times according to case urgency and, from the results, to obtain concrete implications for operational improvement and greater efficiency in the field. By analyzing real operational data in detail and providing a statistically rigorous evaluation, Dr.DataScience contributed to solving challenges in the medical setting.

Background and objective

In this case, what was required was to evaluate in detail the operational efficiency between cases of differing urgency in the emergency-medical-response setting. Specifically, the objective was to clarify whether there was a statistically significant difference in the initial “decision time” and the subsequent “activity time” between a “high-urgency case group” and a “low-urgency case group.” Through this analysis, I aimed to obtain insights leading to the formulation of optimal response protocols, staffing, and improvements in education and training suited to the characteristics of each case group.

Data and variables

This analysis used large-scale real operational data spanning multiple years, provided by a particular emergency-medical-response institution. The main variables analyzed were as follows.

    • Comparison groups: the high-urgency case group and the low-urgency case group
    • Main evaluation items: the “decision time” from the occurrence of the event until a decision is made, and the actual “activity time” after the decision. These times are recorded in minutes.

Analytical methods

In this case, I performed the statistical analysis for the comparative examination of response times in the following steps.

  1. Preliminary evaluation of the data distribution
    • First, for the main evaluation variables “decision time” and “activity time,” I examined in detail the distributions in the all-year data and in each year’s data.
    • I evaluated the distributional characteristics in multiple patterns, including the full sample, the data with outliers excluded, and data with transformations such as logarithmic and square-root transformation applied.
  2. Confirmation of normality
    • In addition to visual inspection, using statistical methods such as the Shapiro–Wilk test, I confirmed that the data distribution did not follow a normal distribution in any of the analysis patterns.
    • This suggested that, because the time data tend to be extremely skewed in particular groups, parametric statistical methods assuming a normal distribution were inappropriate.
  3. Selecting the appropriate test method
    • Based on the evaluation that the data did not follow a normal distribution, I judged and adopted the Mann–Whitney U test—a nonparametric method that evaluates the difference in the overall distribution rather than the difference in medians—as the most appropriate statistical method for testing the difference between two independent groups.
  4. Multiple-comparison correction
    • To address the problem of multiple comparisons arising from performing multiple statistical tests (for both “decision time” and “activity time,” and comparisons spanning multiple years), I corrected the p-values using the Holm method. This appropriately controlled the overall Type I error rate and enhanced the reliability of the results.
  5. Conducting the difference tests
    • Combining the selected Mann–Whitney U test with multiple-comparison correction, I tested whether there was a statistically significant difference in “decision time” and “activity time” between the “high-urgency case group” and the “low-urgency case group.” This was done for both the all-year dataset and the dataset for each individual year.

Overview of the main results and clinical considerations

As a result of the analysis, for the “dataset: all years,” it became clear—and continued to be clear even after multiple-comparison correction—that the “decision time” and “activity time” of the low-urgency case group were statistically significantly longer than those of the high-urgency case group (significance level p < 0.05).

Notably, for some variables, even though the medians were equivalent, the Mann–Whitney U test detected a significant difference. This is because this test can capture not merely the difference in medians but differences in the shape and location of the entire distribution, suggesting that there is an essential difference in the patterns of response time between the high-urgency and low-urgency case groups.

This result is very important for quantitatively grasping the characteristics of response according to urgency. For example, it was shown that for high-urgency cases, because more rapid decisions and actions are required, response times may tend to be shortened.

On the other hand, for low-urgency cases, there may be different operational characteristics—such as requiring more time for detailed situation assessment and gathering peripheral information, or having more time to spare for response because priorities differ.

Dr.DataScience’s contribution

In this case, Dr.DataScience made a substantial contribution to drawing practical implications from a large volume of operational data.

  1. Ensuring statistical rigor
    • For the realistic challenge that the actual data did not follow a normal distribution, I thoroughly evaluated the distributional characteristics of the data and, based on the results, selected and applied the appropriate nonparametric method—the Mann–Whitney U test.
    • Furthermore, to appropriately control the increased risk of Type I error from conducting multiple tests, I applied multiple-comparison correction by the Holm method, enhancing the reliability and validity of the results.
  2. Extracting practical insights
    • By using a method that evaluates the difference in the entire distribution rather than merely comparing means, I clearly captured the subtle yet important differences in “decision time” and “activity time” between cases of differing urgency.
    • Through this detailed analysis, it became possible to more deeply understand the operational characteristics of the field and to obtain solid grounds for formulating concrete action plans for greater efficiency and improved response quality.
  3. Clear interpretation of the results and provision of implications
    • I presented the results of the statistical analysis not merely as figures or p-values, but clearly as clinical and operational implications—what they mean for field operations and what improvements they could lead to.

In this way, the statistically rigorous and practical insights provided by Dr.DataScience made a direct contribution to the client’s decision-making process.

By clarifying, based on objective data rather than mere rules of thumb or intuition, the reality of the optimal decision times and activity times for cases of differing urgency, it became possible to formulate concrete improvements such as reviewing field operation protocols, more efficient staffing of limited medical resources, and developing effective education and training programs tailored to the characteristics of each case group.

In this way, data-science expertise contributed to solving challenges in a medical setting demanding both complexity and high precision, ultimately helping to dramatically improve the quality and efficiency of emergency medical response.

© Dr.データサイエンス. All Rights Reserved.