Dynamic Low-Power Inference in Quaternary Tensor Accelerators
Abstract
We propose Dynamic Low-Power (DLP) operational modes for modern GPU and 3D-stacked architectures based on topological pruning and bus asymmetry. Rather than operating at full symmetric connectivity and power consumption, we demonstrate that hardware designs can activate a secondary operational mode in which the system intentionally gates off portions of the interconnect—including vertical Through-Silicon Vias (TSVs)—while asymmetrically biasing the remaining routes. Physical Mechanism The underlying physical mechanism is the Non-Hermitian Skin Effect (NHSE). Under controlled asymmetry and link pruning, computation naturally concentrates toward topological boundaries, such as a 3D chip's heatsink layer. This allows interior or thermally stressed silicon regions to remain powered off without loss of computational performance. Validation We validate the approach using quaternary neural-network inference mapped onto planar and 3D spatial meshes under open boundary conditions. Exact symbolic analysis establishes the underlying mathematical structure of continuous-time information drift. Hardware simulations on mobile processors confirm the operational behavior. Extreme topological strangulation reveals zero measured latency penalties under the tested conditions. Key Enabler The central enabler is quaternary quantization. The discrete four-state weight space maintains stable inference under asymmetric routing and across topological phase transitions. This enables a hardware–software co-design paradigm in which operational mode switching and active thermal isolation become first-class architectural features. Results and Applications Result: More than 50% reduction in energy cost while maintaining robust inference accuracy of approximately 80%, effectively mapping the experimentally accessible percolation limits of the architecture. The proposed approach is directly applicable to: • Consumer GPUs• Enterprise tensor processors• Edge AI accelerators• 3D-stacked computing architectures Attention: The paper may be more up-to-date. Licensing Manuscript: CC BY-NC 4.0Algorithmic implementation (software/hardware agnostic): PolyForm Noncommercial License 1.0.0
// Source
Authors: Andres Sebaatian Pirolo