Pith. sign in

Paper Citation Record · LEDGER

Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.16866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.16866 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.748237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:30.418964Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c015d103-93b1-4f2f-8f8d-eeb9445e64a7 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.035027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.035027Z digest=sha256:dffcd495ca3efb6e19a56d09ac3810df3798ba0e673bb83b9637de0b7ff16213

Observation 8847315a-6971-4b00-a272-49be1b9d1b92 · inbound

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models cites this paper.

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:36:34.710222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:36:34.710222Z digest=sha256:5af1ef1dfd894fec6335befb7994ea389fceff60e38f85fe7f14322f56c4590d

Observation 0b2efd09-6a65-4065-bb2f-bb74d5075251 · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.433530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.433530Z digest=sha256:181b05032eabd071134f74ae89e2311c73e47ef73d6390b04592fce6bcc3346e

Observation 05ff9b75-378d-4771-9df1-bd40edd1ee85 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.476383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:28bf146e8e4d106c137a76fb0c1346edba04d272564157eb34778a099e273697

Observation 750f8264-b226-4935-b8ae-ca15a37c60bd · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.748237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.748237Z digest=sha256:48edc55d1a884e115c96baba940a90cb507475998fad15cd5280fe915fa919d9

Observation 53d18302-1b89-4e28-9c2f-f93f8e69bb9f · inbound

Describe Anything: Detailed Localized Image and Video Captioning cites this paper.

Describe Anything: Detailed Localized Image and Video Captioning Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:15.679702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:15.679702Z digest=sha256:f0232869cc152b93f05962df63f16d3289df1446282c67b43bcdab395900d52a

Observation 15668e53-1611-4e51-8800-4b57851c1a04 · inbound

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration cites this paper.

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:35.289793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:35.289793Z digest=sha256:a087d492b7ea4726895561a265d9c1820959bf9e2552aac78ce585075b2c53da

Observation b62095e3-85b8-401e-96ee-90db625f3260 · inbound

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training cites this paper.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.615927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:72a82ea985abd2af109f3c25d9ffcfc0aea06138ed59bdbb5b2a598be3e0a6e5

Observation 0d56b3be-ae1c-4d34-9335-5e815e5e7573 · inbound

The Mechanistic Emergence of Symbol Grounding in Language Models cites this paper.

The Mechanistic Emergence of Symbol Grounding in Language Models Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T09:45:01.649392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:45:01.649392Z digest=sha256:fa73e39f197396d76be9d5b27aa3d729c134a72e986254431e5cde0a95e80843

Observation 7f446895-3ac8-403a-9351-5e3c71dd2064 · inbound

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs cites this paper.

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:30.421473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:16:53.915716Z digest=sha256:d0546046a839708d43dd9e5051c9af3f887d32b095cefb886abf6cd634910844

Observation a2cf8f59-6714-45bc-8395-4bdbb072b083 · inbound

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding cites this paper.

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T11:14:52.015620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:14:52.015620Z digest=sha256:6bf1f0fa9de03d62204d07de2caf1d9ead8d7c54276b74e49e733ed2fb52d884