REVIEW 3 major objections 5 minor 71 references
WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding
T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Ultra-high-resolution remote sensing improves when models organize a small, topology-preserving evidence set under global context, not when they simply see more pixels.
desk verdict Solid training-free UHR systems paper with real multi-benchmark gains; the “organize better” slogan is only partly isolated by the ablations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Minimal Support Evidence Set (MSES) under Global Context Constraint, plus Structured Evidence Reasoning via Structured Evidence Metadata (SEM) and a Topology-Preserving Evidence Board (TPEB). Together they select a small complementary patch set and re-present it so the model retains spatial grounding and relative layout for joint global-local reasoning.
What would settle it
On the same UHR benchmarks, fix the visual budget and either (a) replace the greedy MSES selection with random or purely question-only top-k patches of equal size, or (b) strip SEM and the topology-preserving board while keeping the same patches: if accuracy does not fall and the rise-then-fall budget curve disappears, the claim that organization—not access—drives the gains fails.
Extended reading notes
Core claim
The paper claims that for ultra-high-resolution remote sensing, effective understanding is a problem of constructing and organizing the right evidence under global context constraints, not of expanding visual access. Its training-free framework, WeaveEarth, first builds a compact Minimal Support Evidence Set that is relevant, low-redundancy, and spatially complementary, then feeds a frozen vision-language model a unified interface of global thumbnail, structured evidence metadata, and a topology-preserving evidence board, yielding consistent accuracy gains over passive whole-image adaptation and active multi-round search.
Load-bearing premise
A training-free greedy score of question similarity plus global-thumbnail consistency, coverage, and redundancy, with a fixed small evidence budget and a hand-arranged board layout, is enough for frozen vision-language models to treat the selected patches and metadata as truly minimal yet sufficient support for spatial answers.
Editorial extensions
If this is right
- UHR remote-sensing VQA can be improved without fine-tuning backbone VLMs by redesigning only the inference input interface.
- Passive resolution scaling and multi-round zoom search are not the only viable routes; a single-pass structured evidence interface can match or beat them at lower latency.
- Spatial-relation and complex-reasoning subtasks benefit most when local patches retain explicit coordinates, roles, neighbors, and relative layout.
- Evidence quantity has a sweet spot: too few patches miss clues, too many dilute them, so minimal-yet-sufficient sets are preferable to always adding more crops.
- The same construction-plus-organization pattern can be dropped onto different frozen open-source VLMs with stable relative gains.
Reading between the lines
- If encoder similarity under a global thumbnail is the bottleneck, swapping in a stronger or task-adapted cross-modal encoder could raise the ceiling without changing the rest of the pipeline.
- The counting limitation the authors note suggests multi-scale or instance-aware evidence units as a natural next module when targets are dense and tiny.
- The same ‘organize better, not access more’ principle may transfer to other large-image domains (pathology slides, satellite video frames) where answers depend on sparse regions and long-range layout.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. WeaveEarth is a training-free, plug-and-play framework for ultra-high-resolution remote sensing VQA that reformulates the task as structured evidence construction and reasoning under global context constraints. Stage 1 (Global-Aware Evidence Construction) scores overlapping patches with question and global-thumbnail similarity (Eq. 1), expands neighborhoods, and greedily builds a compact Minimal Support Evidence Set via Rel+αCov−βRed (Eq. 2). Stage 2 (Structured Evidence Reasoning) attaches Structured Evidence Metadata and arranges patches into a Topology-Preserving Evidence Board, which, with the global thumbnail, is fed to a frozen VLM. On LRS-VQA, MME-RealWorld, and XLRS-Bench, and across Qwen3-VL-8B, LLaVA-v1.6-7B, and IXC-2.5-7B, the method reports consistent gains over passive whole-image and active multi-round UHR baselines, with ablations (Table 3), efficiency comparisons (Fig. 3), and budget curves peaking at |S|=6 (Fig. 4) offered as evidence that gains come from better organization rather than expanded visual access.
Significance. If the organization-over-access thesis holds, the paper offers a practical and conceptually clean alternative to both costly whole-image adaptation and high-latency multi-round search for UHR RS understanding. Strengths include: (i) training-free transfer across three frozen backbones (Table 4); (ii) multi-benchmark evaluation with sample-size-weighted averages; (iii) component ablations and an accuracy–latency comparison against representative passive and active methods; (iv) a budget-sensitivity curve that rises then falls, consistent with a “minimal yet sufficient” evidence regime; and (v) public code. These make the work useful for deployment-oriented RS-VLM pipelines even if some mechanistic claims need tighter controls.
major comments (3)
- [§4.3–4.5, Tables 3–4, Figs. 3–4] Central claim isolation (Tables 1–4, Figs. 3–4; §4.3–4.5): The paper’s load-bearing thesis is that gains come from organizing evidence (MSES+SEM+TPEB under GCC), not from expanded visual access. Fig. 4 only varies WeaveEarth’s own |S|; Table 3 removes modules from the full system, so every ablated row still uses encoder-selected multi-patch input and residual structure. There is no matched visual-budget control that feeds the same number of high-resolution, question-relevant crops plus the global thumbnail without SEM and without topology-preserving layout (plain top-k multi-crop). Without that condition, it remains unclear whether the frozen VLM uses metadata/topology or simply benefits from several relevant high-res patches. Please add this control on at least one backbone and two benchmarks; if the unstructured multi-crop closes most of the gap, the slogan and interpretation of Figs.
- [§3.2, Eqs. (1)–(2)] Eq. (2) and greedy MSES construction (§3.2): Rel(S,q), Cov(S), and Red(S) are named but not defined operationally (feature aggregation for Rel; coverage metric for Cov; overlap/redundancy measure for Red). The claim that a training-free greedy maximizer yields a “minimal yet sufficient” support set therefore cannot be audited or reproduced from the main text alone. Please specify exact formulas, the encoder used for sim(·,·), values of λ/α/β, overlap/grid settings, and either a short justification that greedy is adequate or a small comparison against a stronger combinatorial baseline on a subset.
- [§4.1, Fig. 4] Hyperparameter and selection sensitivity (§4.1, Fig. 4): Free parameters include λ, α, β, evidence budget (default 6), patch grid/overlap, and TPEB layout. Only |S| is swept. Given that the weakest assumption is that encoder similarity under a global thumbnail plus hand-chosen budget yields answer-critical evidence, at least a limited sensitivity study for λ and (α,β)—or a clear statement that defaults transfer without retuning across backbones/benchmarks—is needed to support the training-free, plug-and-play claim.
minor comments (5)
- [Table 1] Table 1: WeaveEarth is listed as 8B while some compared UHR methods use 3B/7B; a short note on parameter fairness (or reporting the same backbone for all UHR methods where possible) would help readers.
- [Fig. 5] Figure 1 / case study (Fig. 5): “ZoomSearth” appears to be a typo for ZoomSearch; please correct consistently.
- [§4.6] §4.6 Limitation: Counting remains weak because small dense objects are hard in the thumbnail; the planned multi-scale extension is reasonable—consider quantifying residual counting error rates by object size if space allows.
- [§3.3, Fig. 2] SEM format in Fig. 2 is informative; ensure the exact prompt template that injects SEM+TPEB into each backbone is in the appendix or code for full reproducibility.
- [§4.2] No error bars or multi-seed variance are reported; for a systems paper this is common, but a brief note on run-to-run stability (or deterministic decoding settings) would strengthen Tables 1–4.
Circularity Check
No significant circularity: empirical training-free systems paper whose claims rest on external benchmarks and ablations, not on a derivation that reduces to its own inputs by construction.
full rationale
WeaveEarth is a training-free engineering framework (Global-Aware Evidence Construction via encoder similarities + greedy MSES under Eq. 1–2, then SEM + TPEB) evaluated by accuracy gains on external UHR RS benchmarks (LRS-VQA, MME-RealWorld, XLRS-Bench) against frozen VLM baselines and prior UHR methods (Tables 1–4). Ablations (Table 3) remove modules and report drops; budget sensitivity (Fig. 4) is an empirical curve peaking at the chosen budget of 6 rather than a tautological identity; efficiency comparisons (Fig. 3) and multi-backbone transfer are likewise external measurements. There is no self-definitional equation equating a claimed prediction to a fitted input, no uniqueness theorem imported from overlapping authors that forces the result, no ansatz smuggled via self-citation as the sole load-bearing step, and no renaming of a known identity presented as a first-principles derivation. Design choices (scoring form, budget, layout) are ordinary hyper-parameters of an empirical method; they do not make the reported accuracy numbers true by construction. Per the analyzer rules this is the expected non-finding for a self-contained systems paper.
Assumptions & free parameters
free parameters (4)
- λ (global-context weight in s_i)
- α, β (coverage and redundancy weights)
- evidence budget |S| (default 6)
- patch grid / overlap / TPEB layout scale
assumptions (4)
- domain assumption Frozen general VLMs can perform joint global–local spatial reasoning when given a thumbnail, a small set of patches, and explicit spatial metadata/topology layout.
- domain assumption Cross-modal encoder similarity to question and global thumbnail is a useful proxy for answer-critical local regions in UHR RS images.
- domain assumption A compact, low-redundancy, spatially complementary support set is preferable to more patches under fixed VLM budgets.
- ad hoc to paper Greedy selection adequately approximates the combinatorial MSES objective in Eq. (2).
invented entities (4)
-
Minimal Support Evidence Set (MSES)
-
Structured Evidence Metadata (SEM)
-
Topology-Preserving Evidence Board (TPEB)
-
Global Context Constraint (GCC)
Cite this review
Pith. "Pith review of WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding." pith.science (2026). https://pith.science/paper/WFM45NYL
@misc{pith2026260710120,
author = {Pith},
title = {Pith review of: WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/WFM45NYL}},
note = {Machine review of arXiv:2607.10120}
}
read the original abstract
Ultra-High-Resolution (UHR) remote sensing image understanding requires Vision-Language Models (VLMs) to capture both the global scene layout and sparse yet task-critical local details under limited computational budgets. Existing methods mainly follow two paradigms. One is passive perception, which relies on resolution expansion or token compression and may therefore discard fine-grained details. The other is active perception, which depends on multi-round zooming and search, but suffers from high latency, contextual fragmentation, and error accumulation. We argue that a more effective path toward UHR understanding lies not in accessing more, but in organizing better. To this end, we propose WeaveEarth, a training-free framework that reformulates UHR understanding as a problem of structured evidence construction and reasoning under global context constraints. Specifically, WeaveEarth first employs Global-Aware Evidence Construction to select a compact, low-redundancy, and spatially complementary Minimal Support Evidence Set. It then introduces Structured Evidence Reasoning, which weaves local evidence, spatial metadata, and relative topology into a unified reasoning interface, thereby enhancing the VLM's ability to perform global-local joint reasoning. Extensive experiments show that WeaveEarth consistently outperforms strong baselines and existing UHR methods across multiple UHR remote sensing benchmarks and multiple frozen VLM backbones. Code is available at https://github.com/XianZhi-Ma/WeaveEarth.
Figures
Reference graph
Works this paper leans on
-
[1]
2025.Claude 3.7 Sonnet and Claude Code
Anthropic. 2025.Claude 3.7 Sonnet and Claude Code. Retrieved February 24, 2025 from https://www.anthropic.com/news/claude-3-7-sonnet
2025
-
[2]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
arXiv 2025
-
[3]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
arXiv 2025
-
[4]
Yuxiang Cai, Yongheng Shang, and Jianwei Yin. 2024. MultiDAN: Unsupervised, Multistage, Multisource and Multitarget Domain Adaptation for Semantic Seg- mentation of Remote Sensing Images. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 1168–1177
2024
-
[5]
Antoine Carreaud, Elias Naha, Arthur Chansel, Nina Lahellec, Jan Skaloud, and Adrien Gressin. 2026. Context-Aware Semantic Segmentation via Stage-Wise Attention. (2026). arXiv preprint arXiv:2601.11310
arXiv 2026
-
[6]
Hengzhi Chen, Liqian Feng, Wenhua Wu, Xiaogang Zhu, Shawn Leo, and Kun Hu. 2025. F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing Segmentation. (2025). arXiv preprint arXiv:2506.07847
arXiv 2025
-
[7]
Yunkai Dang, Meiyi Zhu, Donghao Wang, Yizhuo Zhang, Jiacheng Yang, Qi Fan, Yuekun Yang, Wenbin Li, Feng Miao, and Yang Gao. 2025. A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs. (2025). arXiv preprint arXiv:2512.17319
arXiv 2025
-
[8]
Lamei Di, Bin Zhang, Yiming Wang, and Wenxia Zhang. 2025. Frequency Meets Semantics: Text-Visual Fusion with Directional Spectral Enhancement for Salient Object Detection in Optical Remote Sensing Images. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 1987–1996
2025
Show all 71 references
-
[9]
Renxiang Guan, Junhong Li, Siwei Wang, Wenxuan Tu, Miaomiao Li, En Zhu, Xinwang Liu, and Ping Chen. 2025. Multi-view Graph Clustering with Dual Relation Optimization for Remote Sensing Data. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 7346–7355
2025
-
[10]
Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, Lei Zhao, Zhuoyi Yang, Xiaotao Gu, Xiaohan Zhang, Guanyu Feng, Da Yin, Zihan Wang, Ji Qi, Xixuan Song, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Yuxiao Dong, a...
2024 arXiv
-
[11]
Yuan Hu, Jianlong Yuan, Congcong Wen, Xiaonan Lu, and Xiang Li. 2023. RSGPT: A Remote Sensing Vision Language Model and Benchmark. (2023). arXiv preprint arXiv:2307.15266
2023 arXiv
-
[12]
Ling Huang, Wenqian Dong, Song Xiao, Jiahui Qu, Yuanbo Yang, and Yunsong Li
-
[13]
InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM)
Language-Guided Visual Prompt Compensation for Multi-Modal Remote Sensing Image Classification with Modality Absence. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 5161–5170
-
[14]
Zhong Ji, Changxu Meng, Yan Zhang, Haoran Wang, Yanwei Pang, and Jungong Han. 2024. Eliminate Before Align: A Remote Sensing Image-Text Retrieval Framework with Keyword Explicit Reasoning. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 1662–1671
2024
-
[15]
Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Jing Li, Jiayang Li, Yue Zhou, Hongjie He, and Jonathan Li. 2025. GRASP: Geospatial pixel Reasoning viA Structured Policy learning. (2025). arXiv preprint arXiv:2508.17102
2025
-
[16]
Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan. 2024. GeoChat:Grounded Large Vision- Language Model for Remote Sensing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 27831–27840
2024
-
[17]
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024. LLaVA- OneVision: Easy Visual Task Transfer. (2024). arXiv preprint arXiv:2408.03326
2024 arXiv
-
[18]
Jianhui Li, Chao Wu, Yingchao Piao, Yuchu Qin, Xiaoping Du, Lili Zhang, and Huadong Guo. 2023. How can we support the UN Sustainable Development Goals when open data is stagnant?Science Bulletin68, 12 (2023), 1216–1218
2023
-
[19]
Ke Li, Di Wang, Ting Wang, Fuyu Dong, Yiming Zhang, Luyao Zhang, Xiangyu Wang, Shaofeng Li, and Quan Wang. 2026. RSVG-ZeroOV: Exploring a Training- Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images. InProceedings of the Fortieth AAAI Confer...
2026
-
[20]
Ke Li, Di Wang, Haojie Xu, Haodi Zhong, and Cong Wang. 2024. Language- Guided Progressive Attention for Visual Grounding in Remote Sensing Images. IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–13
2024
-
[21]
Ke Li, Ting Wang, Di Wang, Yongshan Zhu, Yiming Zhang, Tao Lei, and Quan Wang. 2026. ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery. (2026). arXiv preprint arXiv:2604.01893
2026
-
[22]
Qingyun Li, Shuran Ma, Junwei Luo, Yi Yu, Yue Zhou, Fengxiang Wang, Xudong Lu, Xiaoxing Wang, Xin He, Yushi Chen, and Xue Yang. 2026. Co-Training Vision Language Models for Remote Sensing Multi-task Learning. (2026). arXiv preprint arXiv:2511.21272
2026
-
[23]
Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou. 2024. RemoteCLIP: A Vision Language Foundation Model for Remote Sensing.IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1–16
2024
-
[24]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 26286–26296
2024
-
[25]
2024.LLaV A-NeXT: Improved reasoning, OCR, and world knowledge
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024.LLaV A-NeXT: Improved reasoning, OCR, and world knowledge. Retrieved January 30, 2024 from https://llava-vl.github.io/blog/2024-01-30-llava- next/
2024
-
[26]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS). 34892–34916
2023
-
[27]
Jiaqi Liu, Lang Sun, Ronghao Fu, and Bo Yang. 2026. Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models. (2026). arXiv preprint arXiv:2509.22221
2026
-
[28]
Ruixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li, Wanchen Li, Lanxuan Xue, Hui Qiao, Weizhan Zhang, Deyu Meng, and Xiangyong Cao. 2025. ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks. (2025). arXiv preprint arXiv:2511.12267
2025
-
[29]
Wang Liu, Puhong Duan, Xudong Kang, and Shutao Li. 2025. Squeezing Context into Patches: Towards Memory-Efficient Ultra-High Resolution Semantic Seg- mentation. InProceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI). 1603–1611
2025
-
[30]
Weiqi Liu, Yongshan Zhang, Xinxin Wang, and Lefei Zhang. 2025. Deep Multi- Level Contrastive Clustering for Multi-Modal Remote Sensing Images. InPro- ceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 1239–1247
2025
-
[31]
Xu Liu and Zhouhui Lian. 2024. RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts. (2024). arXiv preprint arXiv:2412.05679
2024 arXiv
-
[32]
Ye Liu, Shitao Song, Miaohui Wang, Hao Gao, and Jun Liu. 2025. DE-Unet: Dual- Encoder U-Net for Ultra-High Resolution Remote Sensing Image Segmentation. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 18 (2025), 12290–12302
2025
-
[33]
Junwei Luo, Yingying Zhang, Xue Yang, Kang Wu, Qi Zhu, Lei Liang, Jingdong Chen, and Yansheng Li. 2025. When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning. (2025). arXiv preprint arXiv:2503.07588
2025 arXiv
-
[34]
Xianzhi Ma, Jianhui Li, Changhua Pei, and Hao Liu. 2025. GeoMag: A Vision- Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing. In Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 5441–5450. ACM MM, 2026, Rio de Janeiro, Brazil ...
2025
-
[35]
Li Mi, Manon Béchaz, Zeming Chen, Antoine Bosselut, and Devis Tuia. 2025. GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 6122–6131
2025
-
[36]
2024.GPT-4o mini: advancing cost-efficient intelligence
OpenAI. 2024.GPT-4o mini: advancing cost-efficient intelligence. Retrieved July 19, 2024 from https://openai.com/index/gpt-4o-mini-advancing-cost-efficient- intelligence/
2024
-
[37]
2024.Hello GPT -4o
OpenAI. 2024.Hello GPT -4o. Retrieved May 13, 2024 from https://openai.com/ index/hello-gpt-4o/
2024
-
[38]
Ruizhe Ou, Yuan Hu, Fan Zhang, Jiaxin Chen, and Yu Liu. 2025. GeoPix: Multi- Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing. (2025). arXiv preprint arXiv:2501.06828
2025 arXiv
-
[39]
Yuwen Pan, Rui Sun, Yuan Wang, Tianzhu Zhang, and Yongdong Zhang. 2024. Rethinking the Implicit Optimization Paradigm with Dual Alignments for Re- ferring Remote Sensing Image Segmentation. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 2031–2040
2024
-
[40]
Chao Pang, Xingxing Weng, Jiang Wu, Jiayu Li, Yi Liu, Jiaxing Sun, Weijia Li, Shuai Wang, Litong Feng, Gui-Song Xia, and Conghui He. 2025. VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis. In Proceedings of the Thirty-Ninth AAAI Conference on A...
2025
-
[41]
Khan, and Salman Khan
Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad S. Khan, and Salman Khan. 2025. GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing. (2025). arXiv preprint arXiv:2501.13925
2025 arXiv
-
[42]
Run Shao, Ziyu Li, Zhaoyang Zhang, Linrui Xu, Xinran He, Hongyuan Yuan, Bolei He, Yongxing Dai, Yiming Yan, Yijun Chen, Wang Guo, and Haifeng Li
-
[43]
Asking like Socrates: Socrates helps VLMs understand remote sensing images. (2025). arXiv preprint arXiv:2511.22396
2025 arXiv
-
[44]
Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Ming- wei Zhu, and Jianwei Yin. 2025. ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration. (2025). arXiv preprint arXiv:2411.16044
2025 arXiv
-
[45]
João Daniel Silva, João Magalhães, Devis Tuia, and Bruno Martins. 2024. Large Language Models for Captioning and Retrieving Remote Sensing Images. (2024). arXiv preprint arXiv:2402.06475
2024 arXiv
-
[46]
Di Wang, Shunyu Liu, Wentao Jiang, Fengxiang Wang, Yi Liu, Xiaolei Qin, Zhim- ing Luo, Chaoyang Zhou, Haonan Guo, Jing Zhang, Bo Du, Dacheng Tao, and Liangpei Zhang. 2025. GeoZero: Incentivizing Reasoning from Scratch on Geospa- tial Scenes. (2025). arXiv preprint arXiv:2511.22645
2025
-
[47]
Fengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang, Haotian Wang, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang, Hongzhen Wang, Wenjing Yang, Bo Du, and Jing Zhang. 2025. GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution. (2025). ...
2025
-
[48]
Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yifan Zhang, Long Lan, Xue Yang, Hongda Sun, Yulin Wang, Di Wang, Jun Song, Jing Zhang, and Bo Du. 2026. GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imager...
2026
-
[49]
Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yuhao Zhou, Di Wang, Yifan Zhang, Haoyu Wang, Haiyan Zhao, Hongda Sun, Long Lan, Jun Song, Yulin Wang, Jing Zhang, Wenlong Zhang, and Bo Du. 2026. Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in ...
2026
-
[50]
Fengxiang Wang, Hongzhen Wang, Zonghao Guo, Di Wang, Yulin Wang, Ming- shuo Chen, Qiang Ma, Long Lan, Wenjing Yang, Jing Zhang, Zhiyuan Liu, and Maosong Sun. 2025. XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?. ...
2025
-
[51]
Kang Wu, Yingying Zhang, Lixiang Ru, Bo Dang, Jiangwei Lao, Lei Yu, Junwei Luo, Zifan Zhu, Yue Sun, Jiahao Zhang, Qi Zhu, Jian Wang, Ming Yang, Jingdong Chen, Yongjun Zhang, and Yansheng Li. 2025. A semantic-enhanced multi- modal remote sensing foundation model for Earth obser...
2025
-
[52]
Kelu Yao, Nuo Xu, Rong Yang, Yingying Xu, Zhuoyan Gao, Titinunt Kitrungrot- sakul, Yi Ren, Pu Zhang, Jin Wang, Ning Wei, and Chao Li. 2025. Falcon: A Remote Sensing Vision-Language Foundation Model (Technical Report). (2025). arXiv preprint arXiv:2503.11070
2025
-
[53]
Liang Yao, Fan Liu, Delong Chen, Chuanyi Zhang, Yijun Wang, Ziyun Chen, Wei Xu, Shimin Di, and Yuhui Zheng. 2025. RemoteSAM: Towards Segment Anything for Earth Observation. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 3027–3036
2025
-
[54]
Liang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang, Rui Min, Shengxiang Xu, Shimin Di, and Pai Peng. 2026. RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow. InProceedings of the Fortieth AAAI Conference on Artificial Intelligence. 11883–11891
2026
-
[55]
Liang Yao, Fan Liu, Shengxiang Xu, Chuanyi Zhang, Rui Min, Shimin Di, and Yuhui Zheng. 2026. RemoteZero: Geospatial Reasoning with Zero Human Anno- tations. (2026). arXiv preprint arXiv:2605.04451
2026 arXiv
-
[56]
Liang Yao, Shengxiang Xu, Fan Liu, Chuanyi Zhang, Bishun Yao, Rui Min, Yongjun Li, Chaoqian Ouyang, Shimin Di, and Min-Ling Zhang. 2026. RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs. (2026). arXiv preprint arXiv:2604.07765
2026 arXiv
-
[57]
Bo Yuan, Danpei Zhao, Zhuoran Liu, Wentao Li, and Tian Li. 2024. Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 2117–2126
2024
-
[58]
Yang Zhan, Zhitong Xiong, and Yuan Yuan. 2025. SkyEyeGPT: Unifying remote sensing vision-language tasks via instruction tuning with large language model. ISPRS Journal of Photogrammetry and Remote Sensing221 (2025), 64–77
2025
-
[59]
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, J...
2024 arXiv
-
[60]
Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin, Zonghao Guo, Fengxi- ang Wang, Xue Yang, Kaiwen Wei, and Lei Wang. 2025. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding. (2025). arXiv preprint arXiv:2512.02715
2025
-
[61]
Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, Jun Li, and Xuerui Mao. 2025. EarthMarker: A Visual Prompting Multimodal Large Language Model for Remote Sensing.IEEE Transactions on Geoscience and Remote Sensing63 (2025), 1–19
2025
-
[62]
Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, and Xuerui Mao. 2024. Earth- GPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain.IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–20
2024
-
[63]
Xu Zhang, Junyao Ge, Yang Zheng, Kaitai Guo, and Jimin Liang. 2025. Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Remote Sensing. (2025). arXiv preprint arXiv:2512.19302
2025 arXiv
-
[64]
Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin, and Tieniu Tan. 2025. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficu...
2025 arXiv
-
[65]
Zilun Zhang, Zian Guan, Tiancheng Zhao, Haozhan Shen, Tianyu Li, Yuxiang Cai, Zhonggen Su, Zhaojun Liu, Jianwei Yin, and Xiang Li. 2025. Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning. (2025). arXiv preprint arXiv:2509.21976
2025 arXiv
-
[66]
Zilun Zhang, Haozhan Shen, Tiancheng Zhao, Zian Guan, Bin Chen, Yuhao Wang, Xu Jia, Yuxiang Cai, Yongheng Shang, and Jianwei Yin. 2025. Enhancing Ultrahigh Resolution Remote Sensing Imagery Analysis With ImageRAG: A new framework.IEEE Geoscience and Remote Sensing Magazine13, ...
2025
-
[67]
Yang Zhao, Shusheng Li, and Xueshang Feng. 2025. Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 11862–11870
2025
-
[68]
Siru Zhong, Xixuan Hao, Yibo Yan, Ying Zhang, Yangqiu Song, and Yuxuan Liang. 2024. UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross- Domain Adaptation. InProceedings of the 32nd ACM International Conference on Multimedia. 6307–6315
2024
-
[69]
Yue Zhou, Jue Chen, Zilun Zhang, Penghui Huang, Ran Ding, Zhentao Zou, PengFei Gao, Yuchen Wei, Ke Li, Xue Yang, Xue Jiang, Hongxin Yang, and Jonathan Li. 2026. DVGBench: Implicit-to-explicit visual grounding benchmark in UAV imagery with large vision–language models.ISPRS Jou...
2026
-
[70]
Yunqi Zhou, Chengjie Jiang, Chun Yuan, and Jing Li. 2025. Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search. (2025). arXiv preprint arXiv:2511.20460
2025
-
[71]
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...
2025 arXiv
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.