When the Proxy Inverts: Distributional Metrics Rank Post-Training Quantization Methods Backwards
Abstract
Standard post-training quantization (PTQ) evaluation often relies on distributional proxy metrics such as KL divergence, perplexity (PPL), and tensor-level reconstruction errors (e.g., Frobenius error). This work systematically investigates the breakdown between these proxy metrics and actual downstream functional performance—specifically focusing on structured emission (tool calling and JSON generation), short answer, classification, and open generation. Through empirical evaluations across multiple task families and multi-stage statistical testing (including independent replications up to pooled $N=240$), we demonstrate key rank inversions where lower distributional error proxies fail to predict actual functional task fidelity. Furthermore, we analyze the distinction between strict syntax and semantic fidelity and outline low-cost evaluation protocols for practitioners. Contents of this deposit: Raw LaTeX source files (main.tex, body.tex contained within ax.tar). Compiled research preprint (PDF).
// Source
Authors: Asadbek Xayitov