AI & Computingarticle2026-09-07

DualGLEAN: Dual Allocation for VLM-Guided Generalized Category Discovery in Remote Sensing Images

Open access0 citations

Abstract

Generalized category discovery (GCD) aims to classify known categories while discovering novel ones in unlabeled data, yet existing methods lack mechanisms to correct boundary-ambiguous samples that receive noisy pseudo-labels, as they primarily rely on visual feature learning without external semantic guidance. Vision-language models (VLMs) offer a natural source of cross-modal semantic correction. However, applying VLM-guided contrastive signals directly within the GCD training loop proves counterproductive because the locally-oriented InfoNCE loss conflicts geometrically with the globally oriented K-means objective in the shared backbone space. We identify the root cause as a dual resource allocation problem: the VLM-derived signal must be allocated to the correct feature subspace to avoid geometric conflict with K-means clustering (space allocation), and the limited VLM inference budget must be allocated to the correct samples to maximize discriminative return (budget allocation). These two decisions are coupled; failure on either renders the other ineffective. To resolve this, we propose DualGLEAN, a framework that addresses the dual allocation challenge through two coupled mechanisms: decoupled contrastive alignment (DCA), which routes the VLM-guided neighbor contrastive loss to a dedicated projector space while preserving the backbone space for global clustering, and compound uncertainty querying (CUQ), a three-stage filtering metric that jointly evaluates predictive entropy, boundary proximity, and local label inconsistency to direct VLM queries exclusively to truly boundary-critical samples. Extensive experiments on the AID and RSSDIVCS datasets demonstrate that DualGLEAN achieves strong performance, improves four diverse GCD baselines as a plug-in module, generalizes across seven VLM backbones, introduces zero additional trainable parameters to the base GCD network, and incurs a total VLM API cost of only CNY 2.45 per full training run on the AID dataset under the default search-scope configuration, with the cost scaling linearly with the query budget.

// Source

View paper (DOI)Open access versionOpenAlexRemote SensingPublished 2026-09-07

Authors: Hongfu Li, Yuxiang Xie, Jing Zhang, Yanming Guo, Xin Zhang

Institutions: National University of Defense Technology, National Defence University, University of Defence