What One Fragile Subgroup Is Worth: Enrichment, Comparator Uncertainty, and Probability of Success for a Phase II Trial of a Cardiac Extracellular Matrix Hydrogel
Abstract
Live Interactive Clinical Interface: https://ventrigelcds.streamlit.app/ Correction notice: This version replaces v1 (Computational Patient Selection for Extracellular Matrix Therapies: A Machine Learning Framework for Phase II Trial Optimization). Its headline result, a Random Forest classifier reporting ROC-AUC 1.00, was invalid. The label the classifier predicted was computed from the same features it was given, so the accuracy measured how well a tree can recover a rule its author had written, and said nothing about predicting patients. The reported feature importances reproduced hand-chosen penalty weights, the hyperparameter search described in the text does not appear in the code, and one reported threshold corresponds to nothing in the data generator. The accompanying correction notice gives the full account. Background: The VentriGel first-in-man trial (NCT02305602) treated 15 post-infarction patients with an injectable decellularized porcine myocardial matrix. It reported, after the fact, that left ventricular remodeling improved mainly in patients treated more than twelve months after their infarction. This study asks what that would mean for a Phase II design, and how much of it holds up. Methods: No simulated patients were used. The trial's Supplemental Appendix gives every endpoint as a mean with its standard error and n, split by stratum, which fixes the standard deviation exactly. The transcription was checked by recomputing all 38 published p-values that could be checked, and by rebuilding each pooled mean and variance from its strata. The trial never tested the early stratum against the late one; that test was run here and corrected for the number of endpoints examined. The missing control arm was anchored to eight control-arm estimates from six published sources, and the standard error on each anchor was carried through. Results are given as assurance, which is power averaged over the uncertainty in the effect and in the comparator, then multiplied by the probability that the subgroup effect is real. Results: All 38 p-values reproduced, with a median standard-deviation reconstruction error of 3.3%. One endpoint of nine shows a nominally significant interaction: end-systolic volume, -16.9 mL, 95% CI -32.2 to -1.6, p = 0.034. It survives neither Bonferroni nor Benjamini-Hochberg correction, although the strata are balanced on all eight baseline measures and the pattern runs opposite to regression to the mean. Natural history after infarction turned out to depend on the era of the trial and to disagree in sign among acute cohorts, while chronic cohorts are stable. Against anchored comparators an unselected trial needs 1,046 patients where an enriched one needs 92. The chronic comparator, however, carries a standard error of 4.4 mL, about the size of the 7.6 mL effect measured against it; carrying that through raises the enrollment for an 80% chance of success from 174 to 406 and caps assurance at 90.9%. At even odds on the subgroup effect being real, the chance the programme succeeds is 40%. Conclusions: Enrichment turns a trial that cannot be powered into one that can, but the conclusion rests on a fragile interaction and on a comparator estimated from 28 patients. A 2x2 design powered on the interaction costs 440 patients, close to the 406 an 80%-assurance enriched trial needs, and it answers the question a sponsor actually has. The next money is better spent measuring the six-month volume change in untreated chronic post-infarction patients than on more enrollment.
// Source
Authors: Palash Rakshit