AI & Computingpreprint2026-08-15

Information-Preserving Learned Compression for Compute-Efficient Language Model Training: A Proof of Concept and a Decisive Experimental Programme

Open access0 citations

Abstract

This work investigates whether the learning-relevant information in a corpus can be represented in substantially fewer computational units through a learned, content-adaptive, information-preserving compact representation, and whether a language model can be trained directly on that representation. The study presents a staged proof-of-concept programme comprising conventional lossless compression baselines, a hand-designed semantic shorthand, shorthand followed by Zlib-9, a learned vector-quantized discrete codebook, direct autoregressive training on a fixed four-bytes-to-one-code representation, and an information-capacity analysis of the fixed-rate constraint. The experiments establish that learned compact codes can be constructed and consumed directly by a model without reconstructing the source byte stream. They also identify codebook collapse and consequent information loss as a principal engineering obstacle and show analytically why aggressive fixed-rate coding cannot represent arbitrary high-entropy blocks within a small fixed code width. The work therefore presents a proof of concept rather than a demonstrated training-compute reduction. It specifies a decisive follow-up experiment using 2×, 4×, 8× and 16× effective representation compression, matched parameter counts, matched-FLOP evaluation, encoder and decoder cost accounting, information-preservation measurements, downstream capability benchmarks, multiple random seeds and statistical uncertainty. The central research question is whether information-preserving learned compression can produce a favourable trade-off between compression ratio, information preservation, model capability and total system computation.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-15

Authors: Rabinarayan Dash