EVEREST:Endogenous Vision-Language Reinforcement Reasoning Exploration for Urban Socio-Semantic Segmentation
Qixiu Li, Zhongzhi He, Xiang Zhu, Xiaoyong Li, Jiarun Lin, Weifeng Xu
Abstract
Urban socio-semantic segmentation leverages digital and satellite imagery to provide critical spatial semantic information for downstream applications such as urban resource allocation. Although existing methods achieve high segmentation accuracy, they still suffer from inaccurate delineation of target boundaries. The underlying issue is that current models primarily rely on passively aggregated global cross-modal cues, lacking active exploration of the environment. To address this limitation, we propose the EVEREST model, which adopts an egocentric exploration strategy that enables the model to actively investigate boundary cues and perform self-correction. In addition, we formulate discrete natural-language prompts as pseudocode to regularize the execution logic. Reinforcement learning is further employed to implement this irreducible process and elicit the model's structured reasoning capability. Our EVEREST achieves optimal performance on all metrics in the real world urban socio-semantic dataset, demonstrating the superiority of our model. Codes are available at https://anonymous.4open.science/r/EVEREST-9D21/.
Create a lesson
Related papers
How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
Corey D. C. Heath
Self-Reflective Multi-modal Reasoning for Short-Video Fake News Detection
Pinjie Xu, Yuzhou Yang, Zhikai Tan et al.
Emotion Understanding in Streaming Video with Trajectory-Aware Reliability
Qingsong Wang, Qigong Lei, Zitong Wang et al.
WaveOp-LiteFM: Lightweight Neural-Operator Flow Matching for Satellite-to-Radar Precipitation Retrieval
Chunlei Shi, Yecheng Zhang, Yufeng Zhu et al.
Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion
Zilong Huang, Junyi Peng, Junjie Li et al.
Task-disentangled Low-Rank Adaptation for Versatile Audio-visual Multi-modal Learning Tasks within a Unified Framework
Hanyu Xuan, Mengqi Zhang, Junjun Mao et al.