Pith. sign in

Paper Citation Record · LEDGER

ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2304.14407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.14407 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.185621Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T22:53:33.136794Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19efc1b1-4ac3-49c7-892a-bae99dec71b5 · inbound

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction cites this paper.

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:53:33.140158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T22:51:08.753650Z digest=sha256:058584828f1bf206de3b06360bac4652cfad5258a780c11a53210d28faf32f10

Observation ec2a373d-f810-4828-a02f-1d37a7f2e28d · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.185621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.185621Z digest=sha256:39018be5f2efab7f7c21c0ca422b1e5beb92e5dc5d57cb18f37b1e62214f9fce

Observation cd22ef6a-5f86-4416-ad29-b5ba31469af1 · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.967715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.967715Z digest=sha256:a8970f76d2569c761563631b14bae2d3c52ae9471f201dba7df0f0cf9b3b5c19

Observation 06d262f9-5160-4354-bb14-4ce73c8dbb54 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:51.278905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:51.278905Z digest=sha256:f8cc7cb3677ac264268c91ad64898fcfcc4d22c70ecfc13e502d9e038dc18b47

Observation 0dd44e4c-e53b-44f3-be17-484d8fd9ff9b · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.124471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.124471Z digest=sha256:fbf360f1e4d7234e4545273e5efc3bd1b1fc78e4c16fcf4226c76f910a248090

Observation ac6a474e-5933-453f-9972-1f4d49aeb7e2 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:09.748930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:09.748930Z digest=sha256:d286b865145e69c8f82f5ef71b5317a527a7a4b4c1bc6ab4c556a2415ff676c7

Observation c9e2d0ce-b931-4d3f-84e6-2435b6d1b4d1 · inbound

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding cites this paper.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:03.282561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:03.282561Z digest=sha256:6231c6c8a5f5ae89c774487aef90bbdc64f68e88cd4d353d73a96aa0c50477b8

Observation 5763ab29-2d37-442c-86c7-bfe7cc13504d · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:27.035766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:c0ec008d17c6f3066a6afc37eb9e06fb839d408975a7cf6663c0147e6bf611a5

Observation 1f8f433f-4812-4bc4-8563-c4b538c5550d · inbound

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents cites this paper.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.132631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:a66aae6106fc5bdfd776cbb64f6645cf05b6c65105edb99f3e696040b2b9b156

Observation 46f6a05a-b353-4bc2-8862-e30fd90d321a · inbound

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding cites this paper.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.473686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:340e0665511e37626dff20da6f8a65199b49f4f17a143ca74c0bc90eb6bae1f6