AI & Computingpreprint2026-08-02

Fovea and Peripheral based Vision for Vision-Language Models.

Open access0 citations

Abstract

AbstractIn this paper it is proposed a novel method that aims to achieve human-like vision (interms of efficiency) for VLMs. In doing so, an almost x10 token input reduction is achieved,and the training time is only 2-5 minutes on low end gpu hardware. Though, additionally,limitations of this method are presented.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-02

Authors: L. Monaco