REVIEW 7 cited by
GREC: Generalized Referring Expression Comprehension
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The objective of Classic Referring Expression Comprehension (REC) is to produce a bounding box corresponding to the object mentioned in a given textual description. Commonly, existing datasets and techniques in classic REC are tailored for expressions that pertain to a single target, meaning a sole expression is linked to one specific object. Expressions that refer to multiple targets or involve no specific target have not been taken into account. This constraint hinders the practical applicability of REC. This study introduces a new benchmark termed as Generalized Referring Expression Comprehension (GREC). This benchmark extends the classic REC by permitting expressions to describe any number of target objects. To achieve this goal, we have built the first large-scale GREC dataset named gRefCOCO. This dataset encompasses a range of expressions: those referring to multiple targets, expressions with no specific target, and the single-target expressions. The design of GREC and gRefCOCO ensures smooth compatibility with classic REC. The proposed gRefCOCO dataset, a GREC method implementation code, and GREC evaluation code are available at https://github.com/henghuiding/gRefCOCO.
Forward citations
Cited by 7 Pith papers
-
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.
-
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
RefBench-PRO organizes REC into attribute, position, interaction, relation, commonsense, and reject tasks; no tested MLLM exceeds 72%, and Ref-R1 raises Qwen2.5-VL-7B from 57.6 to 69.4 on it.
-
Generalised Medical Phrase Grounding
MedGrounder grounds radiology sentences to zero, one, or multiple scored image regions, outperforming single-box and grounded-report baselines on multi-box and non-groundable phrases.
-
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
ReMeREC introduces a relation-aware multi-entity referring expression comprehension framework and the ReMeX dataset, reporting state-of-the-art grounding and relation prediction, with some evaluation caveats.
-
MDC-R: The Minecraft Dialogue Corpus with Reference
MDC-R adds expert annotations of anaphoric and deictic reference, with block-level IDs and bounding boxes, to 101 Minecraft building dialogues, and shows that current referring-expression models struggle on this dynam...
-
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
Griffon-R generates its own grounding hints and rationale before answering, achieving state-of-the-art visual reasoning on VSR and CLEVR while improving MMBench, ScienceQA, and TextVQA.
-
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
TACO couples thinking with final answers, filters unstable training samples, reweights easy or hard samples, and adds multi-scale test inference, improving LVLM visual reasoning accuracy over VLM-R1.
Discussion (0). Sign in to comment.