Pith. sign in

Paper Citation Record · LEDGER

Foundation Models for Video Understanding: A Survey

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.03770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.03770 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:02.474751Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:25:09.414000Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f82032cd-ba4f-40fb-82a5-c206f06c27c6 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Foundation Models for Video Understanding: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.474751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:02.474751Z digest=sha256:0a73db066ac1a964940ae980edf42781de5d6ab26b40a5f897d1d66b4f954b0d

Observation a0f8d93c-3dd1-4fe1-99ee-bb4195847a37 · inbound

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review cites this paper.

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review Foundation Models for Video Understanding: A Survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:51.313475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:51.313475Z digest=sha256:6e4f6d2f847d53e50131338d6c8a05911101a54fe8933f2c3eb21fd552a51824

Observation d9c5ec05-8374-4726-822d-161b0b5cbc8d · inbound

How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction? cites this paper.

How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction? Foundation Models for Video Understanding: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:47:34.195792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:47:34.195792Z digest=sha256:147d9b75c37ebf40d148765c106a30e1a4f9e57b477a88d7b4ec5a5a17b843de

Observation d21fed4b-9c62-4799-aaf2-71ddaa2232ec · inbound

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications cites this paper.

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications Foundation Models for Video Understanding: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:08.167795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:08.167795Z digest=sha256:2869104b1c8a47da5db65ca03e7dff6cb3f478b702358f16085c533db330d684

Observation 95461f5d-2fd6-4227-a5dc-1a5feab35128 · inbound

Towards channel foundation models (CFMs): Motivations, methodologies and opportunities cites this paper.

Towards channel foundation models (CFMs): Motivations, methodologies and opportunities Foundation Models for Video Understanding: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:44.013977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:44.013977Z digest=sha256:6f431be600adbe8bb7100ebddfe970d672413dd94ca2dd84ca05d9817cc1cf5b

Observation 085a55b7-2b21-43a7-8a36-80c70466fbb1 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Foundation Models for Video Understanding: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.068720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.068720Z digest=sha256:03c3b9c68dcee9eba8d40f64fcd6c070e1f71e3aabce7bc5948cc2fa89313218

Observation 0ee60542-d7c1-4097-af91-440d060e6b87 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Foundation Models for Video Understanding: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.695590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.695590Z digest=sha256:d088b602ed81d62c4857b3cf2324a7a37a81ea07cbbaf539a5238daf8a0cdd90

Observation 217712fc-9fa7-468a-8a99-833d6bb663f6 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models Foundation Models for Video Understanding: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.600749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.600749Z digest=sha256:f1d20e113418196a4c62778dd1b0ae05348fed3efaba76ad61f3b214959bb251

Observation d40b607d-74ce-4e26-acbf-e826b0c92f83 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Foundation Models for Video Understanding: A Survey

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.854630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:b993ae2cbc2558849f6487f42919484936d5e2d753c7d11f002a52abe8992f39

Observation 2f579b52-7fdf-4d0e-a219-75c2105d3562 · inbound

IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models cites this paper.

IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models Foundation Models for Video Understanding: A Survey

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:46:11.501822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T02:32:05.523343Z digest=sha256:5bc288f5de239c3ca3b3966536d6149f388cc094f02131ca4ae9d465bc863640

Observation 5fff32ad-f7d2-408f-86d0-413567330b0e · inbound

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding cites this paper.

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding Foundation Models for Video Understanding: A Survey

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.610362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T11:52:26.828633Z digest=sha256:9408fa6fb66a8e4fd7c4486f50f0b8413c78c56d9339d0cfb3c38a9d9cdf7d7a

Observation 0bcab266-27a7-4fab-8856-88bbeeff037e · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis Foundation Models for Video Understanding: A Survey

Reference 133

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.612224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:0edca48a603fca58e97e7cae69e36aae9ae943bd0b94035bce7f61bd902bb643

Observation af3879f6-b6b7-4e58-a6a4-0367716f0e62 · inbound

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark cites this paper.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Foundation Models for Video Understanding: A Survey

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:23.343071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:18:00.436343Z digest=sha256:bd4024a3552641be10bc3ceb5975b06891ea25196642b0cfcf82a3401c0b93cc

Observation 5205c761-7247-497a-a491-565ce0053d5f · inbound

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark cites this paper.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Foundation Models for Video Understanding: A Survey

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.415571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T00:24:25.203292Z digest=sha256:98a3983674e5d82644e0ec4f543d9675cf429a66993e5264f0bc70f77ba02819

Observation cdcf1409-f6e6-4da9-baab-aed54170f6a1 · inbound

Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC cites this paper.

Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC Foundation Models for Video Understanding: A Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:33:25.410861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T15:31:45.565641Z digest=sha256:bc211ba41751a42bf356b76b0afd99ce4d7318b32c407f5d1f91f728d81aaed0

Observation bc24fa89-09ca-4088-b910-5e1e172959bc · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Foundation Models for Video Understanding: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:2d0c81905c24eb14196704959759d6a8fb117379e87246f8da55a9dded4d6012