* This article is based on an actual analysis project; however, in light of the confidentiality agreement (NDA) with the client, while maintaining the framework of the medical and healthcare field, specific details such as disease names and variables have been substantially altered from the actual case. We ask for your understanding in advance.
This case compared and verified, in critically ill sepsis patients admitted to the intensive care unit (ICU), the predictive accuracy of multiple novel blood and urinary biochemical indicators (biomarkers) and a machine-learning-based “integrated prediction score” calculated from the patients’ basic information, for predicting the “early onset of acute kidney injury (AKI).”
In the clinical setting, when there are multiple test indicators for predicting the onset of a disease, it is insufficient merely to evaluate individually whether each indicator “is useful for prediction”; it is necessary to rigorously compare whether “the new indicator is statistically significantly superior to the existing one.” However, when comparing multiple indicators measured simultaneously from the same patient population, a strong correlation arises among the data (a paired data structure), so using the wrong statistical method carries the risk of misjudging the superiority of the tests. Furthermore, repeating round-robin comparisons among multiple indicators markedly increases the probability of a significant difference appearing by chance (a Type I error).
Through this analysis, I aimed to apply an advanced statistical test that can correctly handle the paired data structure obtained from the same patient, and further to control the increase in error accompanying multiple tests using the concept of the false discovery rate (FDR). By combining the precise comparison of the area under the ROC curve (AUC) with multiple-comparison correction that prevents false positives while maintaining power, Dr.DataScience scientifically identified the optimal predictive indicator that can be truly relied upon in the clinical setting, making a substantial contribution to establishing an early-treatment-intervention protocol for critically ill patients.
In this case, what was required was to determine the optimal indicator for accurately predicting, at an extremely early stage before onset, “acute kidney injury (AKI)”—a complication with high lethality that develops following septic shock and the like. At the target medical institution, in addition to the prediction indicator currently in mainstream use, there were a total of five evaluation axes: three trace biochemical indicators newly under consideration for introduction, and a “novel integrated prediction score” automatically calculated from patients’ electronic medical-record information.
The greatest clinical challenge was that, because introducing a new test method or prediction score requires corresponding cost and effort, one had to prove as an objective figure “whether the new indicator truly has discriminative ability surpassing the existing indicators.” Merely lining up and looking at the sensitivity and specificity of each indicator cannot determine whether the difference is statistically meaningful or merely chance variability in sampling.
The main objective of the analysis was to calculate, for all five prediction indicators, the area under the ROC curve (AUC)—an index of the ability to discriminate the presence or absence of AKI onset—and to clarify whether a statistically significant difference exists among the AUCs. Another goal was to perform appropriate multiple-comparison correction based on two different premises (case settings)—”determining superiority among the biochemical indicators,” which is the focus of the study, and “determining superiority among all indicators including the integrated prediction score”—and to derive conclusions based on rigorous evaluation criteria.
This analysis used prospectively collected clinical observational data of several hundred critically ill patients at a particular advanced acute-care medical institution. For all patients, the test values at admission were comprehensively obtained. The main variables analyzed were as follows.
In this case, to compare the discriminative ability of multiple prediction variables obtained from the same patient and to avoid the statistical fallacies accompanying multiple tests, I selected and applied the following extremely rigorous statistical methods.
In the analysis of “Case 1,” which compared only the biochemical indicators, I evaluated by the DeLong test whether there was a difference between indicator A, which had the lowest predictive ability (AUC = 0.315), and indicator D, which showed moderate predictive ability (AUC = 0.887); the raw p-value was p = 0.016, and a significant difference was found. Furthermore, when I applied multiple-comparison correction by the BH method to this result, the calculated q-value was q = 0.047, which fell below the predetermined false-discovery-rate threshold (FDR = 0.05), so the robust conclusion was drawn—even after correction—that “indicator D has statistically significantly superior predictive ability to indicator A.”
On the other hand, in the DeLong tests among the top three indicators—indicator D (AUC = 0.887), indicator C (AUC = 0.940), and indicator B (AUC = 0.946)—the raw p-values all exceeded 0.05 (for example, p = 0.057 for indicator D vs. indicator C), and naturally no significant difference was found after BH correction (q > 0.05).
Furthermore, in the analysis of “Case 2,” which added the novel integrated prediction score (AUC = 0.149), the criterion for multiple-comparison correction became stricter due to the increased total number of tests. As a result, in the comparison of indicator A and indicator D, although the raw p-value was p = 0.016, the q-value after BH correction was q = 0.062. This presented a result on a delicate borderline left to clinical judgment: under the extremely strict criterion of FDR = 0.05, “it cannot be definitively asserted that there is a significant difference (reserved),” whereas if the FDR = 0.10 criterion permissible for exploratory indicators is adopted, it is judged “significant.”
From these multifaceted analysis results, an extremely clear overall clinical conclusion was obtained: “Indicators B, C, and D have clearly superior predictive ability compared with indicator A and the novel integrated prediction score, but among these top three indicators there is no significant superiority or inferiority that can be statistically proven.” This enabled the medical institution to decide on a cost-effective introduction plan: it need not introduce all three expensive top indicators, but may select just one that is easiest to operate from the standpoint of ease of measurement and cost.
In this case, Dr.DataScience completely plugged the statistical pitfalls lurking in a simple comparison of indicators and created unshakable objective evidence to support an important decision in the medical setting.