Pith. sign in

Paper Citation Record · LEDGER

HourVideo: 1-Hour Video-Language Understanding

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2411.04998.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04998 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:32:55.211096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T16:03:08.061136Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cbe8f332-5b99-4c76-8ab7-33c1155b89e5 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling HourVideo: 1-Hour Video-Language Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.666934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:35930b326ff3433131d95246a36f47cb4a0bace0995aadd9a5e8c73f3250ec30

Observation 9fd797b2-7f4d-4da5-84de-bb69088dd631 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models HourVideo: 1-Hour Video-Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.211096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.211096Z digest=sha256:7e2ece7395fded73443c5963159452c4b85b448b366de5cae85de9e8e7aaf8fe

Observation fa46b6a8-778e-4b90-8dcb-ef52e97c5367 · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.393069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.393069Z digest=sha256:2f3af6d42216dc8af4a7a770086fb7afde3579e9152c064ba77d252f652c3b1e

Observation c04387ad-e5a2-4f71-b3a1-cb412a049b00 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HourVideo: 1-Hour Video-Language Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.694510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:66a7aab11470697171b8f59826b786d6c0d8284ba3f797d5479d15d0e91991eb

Observation b5c2abd0-7de7-4cfb-a7dd-de5ee9ae5132 · inbound

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos cites this paper.

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos HourVideo: 1-Hour Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:36.983483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:36.983483Z digest=sha256:ade6c9bdeb703e7b582d8facc447d04b3868d6ecfc69d5fc61028eceddcbfa09

Observation 414f0402-51a4-494a-af35-088dcae167d2 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.359538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:1992352feae74ed02e371daec193ad352ec4a6ea223b41eb12d8a4fb860d42f0

Observation ed00242b-b453-4eac-8022-3632c09b33c5 · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models HourVideo: 1-Hour Video-Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:58.871782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:58.871782Z digest=sha256:92ec423fda2cf2662deb5a5f56a999a4c8162c8294fe581be52ae6a86d405ab9

Observation 1be4414a-8ae4-4d38-8c76-e76aaa204ed4 · inbound

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations cites this paper.

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations HourVideo: 1-Hour Video-Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:26.874122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:26.874122Z digest=sha256:ab11fdcd81ef28472f1ef57c40daf2fe5115955db82db13187b7535e413cbd08

Observation 01bcc5ab-e0c7-4408-a60e-48774759a8fe · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.211447Z digest=sha256:bcb559e8839abcfeb85dcc7de0d63464edf4f3221c3ffe6dc00f360e2b58df9d

Observation 4f25828d-e529-40a3-a226-d630579dc041 · inbound

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment cites this paper.

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment HourVideo: 1-Hour Video-Language Understanding

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:36:00.359960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:35:34.760659Z digest=sha256:a02798460da498cfb9140541ff0049e3eaa29c789521e54dee4322110018d4db

Observation 7d037126-b102-4049-b78a-c9a19bdd683b · inbound

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment cites this paper.

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment HourVideo: 1-Hour Video-Language Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:16.657829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:16.657829Z digest=sha256:dded1bbfda017b1cd26ae04c3ae223ac046d3f4764f460c71c9156bb20d287e6

Observation 864588dc-4e06-40b1-b981-5da984992834 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.063406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:494d199ad869888ea2eaeb838d1da11b5a493a54045c8d781c4a24fb325db447