Is There a Best Hypergraph Neural Network? A Significance-Aware Recomputation and Statistical Audit of DHG-Bench
Abstract
Deep hypergraph learning is evaluated almost entirely through leaderboards that rank methods by mean accuracy over a few random seeds, usually without significance testing. Is there a best hypergraph neural network, or does the apparent ordering reflect seed noise? We independently recomputed the node-classification track of DHG-Bench on a single GPU with twenty random seeds (against five upstream) and a different software stack, and applied a four-layer statistical audit to the per-seed accuracies: a reproducibility check, per-dataset paired Wilcoxon tests with Holm correction, an across-datasets Friedman/Iman–Davenport omnibus with Nemenyi and Holm-corrected pairwise tests, and a variance decomposition. Within a single dataset, twenty seeds distinguish most method pairs (74–98%), so the protocol is not underpowered. Across the nine datasets where all 17 methods complete, the omnibus rejects global equality (Kendall’s W=0.45), yet no pair survives Holm correction, and the top methods fall within one critical-difference band. One dataset carries more seed noise than between-method signal and cannot rank methods. The recompute also documents a non-reproducible method, a label-range data fault, and missing per-dataset configurations in the public release. No single method is statistically best across these datasets, so single-leader claims are not supported; we release a reusable significance-aware evaluation protocol.
// Source
Authors: Valeriya V. Tynchenko, Sergei Kurashkin, А. С. Бородулин, Tee Connie, Ahmad Hammoud, Vadim S. Tynchenko
Institutions: Bauman Moscow State Technical University, Siberian Federal University, Multimedia University