REVIEW 8 cited by
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Referring expression comprehension (REC) involves localizing a target instance based on a textual description. Recent advancements in REC have been driven by large multimodal models (LMMs) like CogVLM, which achieved 92.44% accuracy on RefCOCO. However, this study questions whether existing benchmarks such as RefCOCO, RefCOCO+, and RefCOCOg, capture LMMs' comprehensive capabilities. We begin with a manual examination of these benchmarks, revealing high labeling error rates: 14% in RefCOCO, 24% in RefCOCO+, and 5% in RefCOCOg, which undermines the authenticity of evaluations. We address this by excluding problematic instances and reevaluating several LMMs capable of handling the REC task, showing significant accuracy improvements, thus highlighting the impact of benchmark noise. In response, we introduce Ref-L4, a comprehensive REC benchmark, specifically designed to evaluate modern REC models. Ref-L4 is distinguished by four key features: 1) a substantial sample size with 45,341 annotations; 2) a diverse range of object categories with 365 distinct types and varying instance scales from 30 to 3,767; 3) lengthy referring expressions averaging 24.2 words; and 4) an extensive vocabulary comprising 22,813 unique words. We evaluate a total of 24 large models on Ref-L4 and provide valuable insights. The cleaned versions of RefCOCO, RefCOCO+, and RefCOCOg, as well as our Ref-L4 benchmark and evaluation code, are available at https://github.com/JierunChen/Ref-L4.
Forward citations
Cited by 8 Pith papers
-
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
PerceptionLM releases 2.8M human-labeled fine-grained video QA pairs and spatio-temporal captions, plus models and a new benchmark, arguing that human data, not just synthetic data, is needed for detailed video understanding.
-
HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding
Under frozen perception, an explicit query-region alignment hook plus perception-grounded abstention reduces conditional binding errors and hallucinations, with an exact Acc=(1-SeeErr)(1-SayErr) decomposition.
-
The Mechanistic Emergence of Symbol Grounding in Language Models
Symbol grounding emerges in Transformers and state-space models through middle-layer 'aggregate' attention heads that connect environmental cues to words, but not in unidirectional LSTMs.
-
Describe Anything: Detailed Localized Image and Video Captioning
DAM achieves state-of-the-art detailed localized captioning on images and videos via focal prompting and a localized vision backbone trained on VLM-expanded segmentation data and self-labeled web images.
-
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
A hierarchical window transformer that injects image details into upsampled CLIP features and compresses them with cross-scale window attention improves MLLM fine-grained perception by 3.7% on average over LLaVA-UHD.
-
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
RoboOS is a hierarchical cloud-and-edge framework that coordinates heterogeneous robots, but its headline model gains are weakened by benchmark designs that overlap with training data.
-
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.
-
CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models
A LoRA-fine-tuned MiniGPT-v2 with USDA retrieval answers food and calorie questions from one image, but calorie accuracy is never measured.
Discussion (0). Continue with ORCID to comment.