Engineering & Technologypreprint2026-08-17

Hidden-FFN on Apple GPUs: Compact Inference and Training from Homogeneous Sparse FFNs to WANDA-Pruned Language Models

Open access0 citations

Abstract

This paper presents a first GPU implementation of Hidden-FFN (HFFN), whose companion CPU worksestablished an exact compact runtime for fixed sparse FFNs. Four Metal paths are evaluated: homogeneous FFNinference and training, then Pythia-70M, 160M, and 410M inference and training with fixed WANDA supports.Only FFNs are replaced; active weights, gradients, and AdamW states remain compact, with no dense fallback orpost-attention fusion.On an Apple M1 GPU, homogeneous inference reaches 2.52× and 1.61× speedup over iso-support PyTorch/MPSat 5% and 10% density; training reaches 3.38× at 5% and 1.69× at 20%. Pythia-410M decode reaches 1.72×at 10%, while training at 20% and N = 32 scales from 1.15× on 70M to 1.57× on 410M. Full WikiText-2inference preserves NLL within 1.3 × 10−6. Across five ten-epoch training attempts per model, all 15 HFFNtrajectories remain finite and 14 of 15 pass the paired final-quality gate. The resulting first-version baselinetherefore establishes both the performance opportunity and a functional compact training runtime, withouttreating its measured operating range as a GPU ceiling of HFFN.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-17

Authors: Maxime Parker