REVIEW 7 cited by
TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than several specialist models trained to perform specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single image instruction-following model on scene classification, visual question answering, and captioning. We publicly release our data, model, and code at https://github.com/ermongroup/TEOChat .
Forward citations
Cited by 7 Pith papers
-
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
MONITRS provides roughly 10,000 FEMA disaster events with temporal Sentinel-2 imagery, news-derived captions, and QA pairs; fine-tuning TEOChat on it raises event classification accuracy from about 50% to about 89%.
-
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing
SkyVLaM introduces a temporal basis perceiver and adaptive dense selection to improve language-conditioned video segmentation in UAV scenes, and contributes the SkyVid dataset.
-
GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing
ChronoBench decomposes long-term remote sensing understanding into four cognitive levels, and the GeoChrono model, using per-location temporal trajectories, achieves 78.34% accuracy—over 20 points above prior MLLMs—bu...
-
VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.
-
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...
-
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
A remote sensing LVLM that augments visual features with retrieved captions and routes them through level-specific experts improves performance on several RS vision-language benchmarks.
-
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.
Discussion (0). Continue with ORCID to comment.