Pith. sign in

Paper Citation Record · LEDGER

Understanding Long Videos with Multimodal Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2403.16998.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.16998 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:44:50.317635Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.913681Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b4cf9ec3-f3d7-4196-b36c-00a50127b88e · inbound

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering cites this paper.

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Understanding Long Videos with Multimodal Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:50.317635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:44:50.317635Z digest=sha256:d534f55f8bf0d17f0a297427fca558c7ec8c620d0e026b64949791e104e1b407

Observation d6ee249b-7f99-4517-b17d-92e41ccbddc6 · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering Understanding Long Videos with Multimodal Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:32.306216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:32.306216Z digest=sha256:b67f3895fda1903b1248759a3d0dc08f00c2a4af8c67b83ba2b8d9d4ad4429b4

Observation c930cb75-2276-44ae-adab-2cf010a1b749 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding Understanding Long Videos with Multimodal Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.871880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:f08e7511a4a79aea81f05cfd13b985d5a92cee0bb158782ed1a42bddc16eb7eb

Observation aa8772c6-d631-472d-b460-9edf5730b0a3 · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding Understanding Long Videos with Multimodal Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:22.177953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T12:03:09.408019Z digest=sha256:0ce8176a03de3b209b1c281dab5f73e00dd7e5135c4c06908d836c65ca7ecd80

Observation ac4c87ab-087c-4c2d-8458-7591a1866ac2 · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding Understanding Long Videos with Multimodal Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:37:41.750865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T17:34:10.344111Z digest=sha256:6b816aa9ea4165127dae379fa9fcab04c1cb81c1da27f74278b1d4d12b119162

Observation 9f09208b-d686-492b-96d3-8d3ef94bd99a · inbound

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models cites this paper.

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models Understanding Long Videos with Multimodal Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:31.182712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:31:40.463891Z digest=sha256:cdf4e4cbd083cae9a11466b09f5e22db24b0434ecff97d5bd34922238af1dafa

Observation 0a58456e-931f-4ce1-a837-dae08856644d · inbound

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture cites this paper.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Understanding Long Videos with Multimodal Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.915306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:e65c1463f1c5171a2ca40bb06637ac85db5ad7d4287f95befdc3ac8c71d0ca9c