Fovea and Peripheral based Vision for Vision-Language Models.
Open access0 citations
Abstract
AbstractIn this paper it is proposed a novel method that aims to achieve human-like vision (interms of efficiency) for VLMs. In doing so, an almost x10 token input reduction is achieved,and the training time is only 2-5 minutes on low end gpu hardware. Though, additionally,limitations of this method are presented.
// Source
View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-02
Authors: L. Monaco