Pith. sign in

Paper Citation Record · LEDGER

ChatterBox: Multi-round Multimodal Referring and Grounding

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.13307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.13307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:52:30.302929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.009759Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b459bffb-b2d8-49b4-b2f8-581389ed9697 · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.302929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.302929Z digest=sha256:856f040ed5f60b39bd1fed5aae02259f495cf263bcc0ff4576e3a7a3eba1322b

Observation 0416a874-b557-45af-b665-202975af1755 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.272318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.272318Z digest=sha256:8a868e690b6589cdfda842ff28fd7c411131a246fc74720c49a203525f83f3d6

Observation 5f725db5-be10-455e-9c8c-8492c98d1363 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.891664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:3edb60c3e97efe31cb0c8f02915b536f125416a76c71f49a65bf7095270a907e

Observation 01f06659-eee1-4f52-a87b-5cf22220bed6 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.380638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:983f38d79120435e077c4e249d548cbbe379e5aff8675bca1f46b4936d5ac57f

Observation 665d09d2-0384-475b-81e5-6c6e3077e78b · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.194382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:33006ef21bb1317111a1b43710d160b4c015419ebcac75118f946a9574907456

Observation b6219926-56a0-4a5a-95f8-3b4ec1a0777a · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.011788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:1a6af17790fce788541fdd02d881c28e80d3094d02bbb774fe07ebdb77494eb1