TopK Sparse Autoencoders Across Three Model Architectures: Dictionary Collapse, Dense-Feature Degeneracy, and the Limits of Activation-Pattern Feature Matching
Open access0 citations
Abstract
TopK SAEs trained at matched sparsity and dictionary size on Llama-3.2-3B, Mistral-7B-v0.3 and Qwen2.5-3B. Reports abrupt dictionary collapse with partial recovery, dense-feature degeneracy invalidating frequency-based feature selection, and a null cross-architecture matching result whose detection floor is measured by planted-signal power analysis (MDES 0.95-1.00).
// Source
View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-14
Authors: J. Melton
Institutions: American Standard (United States)