Pith. sign in

Paper Citation Record · LEDGER

ChatterBox: Multi-round Multimodal Referring and Grounding

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.13307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.13307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:52:30.302929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.009759Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b459bffb-b2d8-49b4-b2f8-581389ed9697 · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.302929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.302929Z digest=sha256:cadf39e350481804bcad2d7c698dd6d75b96261b5ec0f56cdbbced8367e7398c

Observation 0416a874-b557-45af-b665-202975af1755 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.272318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.272318Z digest=sha256:0be288f2776e0abbf71721b256b28ddf582159015c0b0cea92f61a01c729857f

Observation 5f725db5-be10-455e-9c8c-8492c98d1363 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.891664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:37db6d7268fc8b9d678ff348341c75a352959011a7a05c1e61e79688841548da

Observation 01f06659-eee1-4f52-a87b-5cf22220bed6 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.380638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:f3da3d4478edc03f335cf9722aa36b2baf48b9cf15ba3a750f0ede25c1f6fa64

Observation 665d09d2-0384-475b-81e5-6c6e3077e78b · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.194382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:35581df18d49f2a865de3e61d20ca2f3561d3b2b0cbf0402aa3b4ed3a9ab8f36

Observation b6219926-56a0-4a5a-95f8-3b4ec1a0777a · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.011788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d6783318120a0533a1fd038b6d4701e8805d0da2034163e60928b7dc83d657f3