AI & Computingpreprint2026-08-14

TopK Sparse Autoencoders Across Three Model Architectures: Dictionary Collapse, Dense-Feature Degeneracy, and the Limits of Activation-Pattern Feature Matching

Open access0 citations

Abstract

TopK SAEs trained at matched sparsity and dictionary size on Llama-3.2-3B, Mistral-7B-v0.3 and Qwen2.5-3B. Reports abrupt dictionary collapse with partial recovery, dense-feature degeneracy invalidating frequency-based feature selection, and a null cross-architecture matching result whose detection floor is measured by planted-signal power analysis (MDES 0.95-1.00).

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-14

Authors: J. Melton

Institutions: American Standard (United States)