Pith. sign in

Paper Citation Record · LEDGER

Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2404.07973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.07973 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:26:28.879917Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.769941Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52c46c55-579c-46fe-97ec-d0cf92d06fcc · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.316983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:a1f9f4d696cf3a3898a61a2f3dccff51154f41b08aa78f2de796bcdf6b1f8e79

Observation bf729760-af01-45a6-893f-d69d7bddf096 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 298

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.322854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:893f510bafc61256d5cffd66510b0b2cc10b6a49f00e1ed15ea3517a74f4d714

Observation eb2197b6-7ad7-43e4-bd0f-417a2b0b5481 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:26.618781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:c3d903753a69015054bb6239bccaba3667081ed18f86143125bdf2deea01e9b9

Observation 30ff817f-1f3f-4dfc-b3db-e497b9323d08 · inbound

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images cites this paper.

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T17:26:28.879917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:26:28.879917Z digest=sha256:fb38b7552d7607129f2270e1b785fb0a0f827e93dbf84de62409fe15f673626f

Observation 6d4e3094-5d35-4e77-99a9-2c7866e3cecf · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.039848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:eb8871f9610d65a563567bbbf32fa43cd750f7b9f8417032675ce605e3d17db3

Observation f3329cb4-26ea-4ff8-b71b-a584ec2106a5 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.173791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b3a213224cfa1161d99385fd8cc6d49f0aaedfd1d58617a1f55f37b6b807ccc0

Observation 19323831-a40a-49de-b004-6b7c50a0a822 · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.903410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.903410Z digest=sha256:f03a1fea6227fd70897332dda4701f897097548b0fc4059d2eeb046c834328e5

Observation ae642888-c3de-4cbe-aa24-3161505b0301 · inbound

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos cites this paper.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.279027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.279027Z digest=sha256:36a712c5fb4412a18a60d38b65361412f2cbec43d6d2e03fd41dde668021f3c0

Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.768322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.768322Z digest=sha256:db722540f0431f7111dcb402f1f7fa23255acb5544c9ae127b6eb23aec95fbb0

Observation cf03d328-6614-4812-9942-6988fa4739e2 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:03.713136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:03.713136Z digest=sha256:1aca2cf5c48c6b85b12d7ffca7acb7c690c177bb078c4d6417e0b9e1dcf0b13b

Observation 4e50f82b-d859-4e4a-b256-471307846976 · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.130365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.130365Z digest=sha256:035b515e38d261c9486d526b81e5ba3ae22e79cb9ec077b5adfe9bff9fc2d1e8

Observation 7c4736ca-60ba-4593-9f5a-14e5318b593a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.946719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:a3301274d78cd4f3ce17a4c0c570009322d3a1b4898841f3670df74853eb9f2b

Observation 65a71161-201a-4948-826e-c30a6127c282 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.257522Z digest=sha256:508eaa0f958cc0ef7dc772b6e6d9f6b9821e35d6dfc7d4df5ed4d4803c0281e3

Observation 34366149-934d-4980-ac63-399f83fe1d0c · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.900012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:37279261ec07e424053818eb6ac543e5e0fa388642a85049e671e9c58dfa7760

Observation 8e0166ae-07b0-46e7-b378-14b89929b444 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 224

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.196499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:e9bd22a429c87c58a4e140cc0e82ddb800edffb208755fc2d9b9c68d05a67169

Observation f6556e68-116d-4302-a7a0-ae0bb9f0cddf · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.572846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:55dc856fca72f3fae4cdc273197ec8018f03316605605ae3d90f3c7b491144d9

Observation 79c338b0-a657-400a-aeeb-d4f246691794 · inbound

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method cites this paper.

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:33:41.834379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:14:25.423302Z digest=sha256:1d93981d6d618d3cc289ce364a177ce547412ca9eb4f53faadde2fc10853c3ab

Observation 7edf8782-c164-4990-aea0-70f39a55e4cd · inbound

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images cites this paper.

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:01.649239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:26:47.052025Z digest=sha256:3fea266f9319c7da52000af6ad6dc4ecf76a25a52f63b865c5c3753c0c7a9fb2

Observation 3c192de7-c3cf-4579-bbb6-8210ba4bad56 · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.145207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:b168b4baa6b271808e14f49fa767e8e2c5e858de082cd47bf1ebf2a8dfed6b89

Observation 51c62cd9-08e9-43a4-8844-3cd356ccf0fa · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.197690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:6e54fa85103c8e46cd4117d149ab0d56fd1e87c0f2ea164f8becea9ace88eb6a

Observation 19cf84d4-0dc3-42cc-891a-1538e68afda9 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.771463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:b0668cce209627141c3b536eea841d37f89366786d749ac83bc1db878fb5e993

Observation 50f680ef-28ec-4ccd-8347-75f771161114 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:972e8380cab42b087d08ca6f21e65bca9a47cc8ccb161b05ad8e1073cbade893

Observation 8378c805-3cb9-4fde-afd2-956cd2c9cfa9 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 194

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:1cc7f676acb3157a1bd47ad076d92ec51a297c7480d1e02657c7f112875e547d

Observation 81988803-75fe-4b14-a681-bdc1f7221ccd · inbound

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement cites this paper.

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T10:23:58.883843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:23:58.883843Z digest=sha256:fae7e78c81677003dc8c99fa370e52c96673b29623bf137d75ccd1bba841fc26

Observation 577a0dee-2ba4-4ba8-b0ae-991bf551e466 · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:45.544126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:45.544126Z digest=sha256:f5dd6c0e379992becd7bb2af78b4b51f0e5aeb602bba47d49e3676e2bfc27eab