Pith. sign in

Paper Citation Record · LEDGER

Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2409.14485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.14485 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:35:30.244486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.682244Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 60d92558-0dc5-40b6-b968-b9685ac8c8b2 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.433404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:fe950cbc399b58553f88a806c17fb1f9e4a1cf753488c05a55a217bd60649487

Observation 1cf91e4a-058b-46e0-825a-d4ae68ca3742 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.987016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a29def8c1a3639e0d2e00e41851d25b5162a20160b3ad1fd8023802d1c9d338a

Observation 731ff6fb-eb7b-48c7-b925-bcf2ff8d0414 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.503594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b8df03eda47578414ddd66eda0e5890e4e241eac03bd9ea21576fd62d016a24e

Observation 8467a30b-0658-47f3-8e17-081184e9bdde · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.770114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:f4c1ca932bbd9bdf20d0c04f07fb546ada66862b760f4fc6fef95665b1d265aa

Observation 7b71ba94-2146-45e8-8a0f-07d06ad2a63e · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.244486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.244486Z digest=sha256:e7359e3193041764f0e3eb115f0e6fe9dbc3ab2e59bf1af76bb57e1917cead4f

Observation beec19e8-68f9-4737-ac36-e866b13ef6d7 · inbound

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos cites this paper.

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:36.988659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:36.988659Z digest=sha256:db40574bcb7b27e0963f522cf7707171778c09cc3d97c887854f9dd11fc2fcc9

Observation 70421979-ef9e-42f6-a482-5d1d7dcc3c9e · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.956964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.956964Z digest=sha256:842131f2a74dfbe16c24788ee42cebfeb372641d375476e0695371938bc978de

Observation 5523b6de-9a63-4e83-8221-c14dd59cc6bc · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.251039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.251039Z digest=sha256:3250f24f333e7cc12eda8d904e82b7a7162575e8103086871aba54c87c2aa045

Observation 8d9c289e-2d5a-4082-a8ec-fcc72a4a1192 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.584507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.584507Z digest=sha256:77b2d08639f735c4094a27fa4b926209bb1ffa03122f23b5dc9a5009b907dd3c

Observation f19b5cbc-5429-425a-8010-6d84e9d4ed61 · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.806587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.806587Z digest=sha256:d59bdc39a51abbb004d392286157b2f022271150a1b9335d0811df2f6081b712

Observation 358355fd-9998-444c-aae7-b87c1418de3a · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.637236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.637236Z digest=sha256:1d7dae2a227bd1e40c8713a88a9437b2991c35e1cbf8da3279adb68fb2fe7d90

Observation c8017646-7848-455d-855d-be4096fd1aca · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:04.246155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:04.246155Z digest=sha256:a7c0dab2573b3a011fcbc13f3a12e9c2faa63ec5a9655fbd76abf01979d55441

Observation e951b058-651e-41cb-82f6-f9ca968c09dd · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.226160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.226160Z digest=sha256:5179423f23297a925b02d3a89a0f675f8665a44239e6787543f46b2898b973e8

Observation b3b21a56-e1ad-4778-a6ce-c322c3a42042 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.507321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.507321Z digest=sha256:3a7eada83fb10fda68df15a2b9662f2cfd46ee19254194dce2a9be03131374a5

Observation 51b1b8f9-be18-4d35-bc41-a649916768d0 · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.755141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.755141Z digest=sha256:60ae6515121c4e00c5173f5e1841330ec64afe1430e4b912f22cd83bfeb621a3

Observation bee69751-5848-4f41-9434-e3ee6d7c25b8 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:07.069669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:07.069669Z digest=sha256:a6c1c45a22e79884f8953d380ec8c71d7aaa11531f75d81e2c31eee8e8599586

Observation 402060c6-11ee-452e-afb6-8d77ad4cb01e · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:02.472353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:02.472353Z digest=sha256:17de67afefa1e3793f51e17d816abd9c48db3ca99dcbb340b6393c0c1966d221

Observation 11789708-fe6c-41c0-b5a3-1c112e649775 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.223612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.223612Z digest=sha256:cd265972976dfe01648970ed98ce0ed272d89bafcd78080280e3ef3127b9e9df

Observation a03ecc67-4006-4690-9bf6-bdae02f2ad6f · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.899236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.899236Z digest=sha256:5ea25c802648043546afbf01c0119750289498a80b691f977cfd3083aa497f01

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:8fd82b6ac9ac73c85c10e5cf1d91278ead83186856eb01c8f46162b9257e413e

Observation e6347225-408e-4c8d-a637-f6576a7f0245 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.402521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.402521Z digest=sha256:4de04f5d4293c25663fc88eb14218ffff9c243baa97e333022fcdc303c3ab267

Observation dab9472d-2449-4130-acf1-d9b6d776d216 · inbound

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding cites this paper.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.738153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.738153Z digest=sha256:c6f34e4ca5b6e7d26c7d1483c787705c539956b2b09c338a8b728b5137218945

Observation e27b375b-5317-4540-b42d-d6fe98846f67 · inbound

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding cites this paper.

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:46:19.505318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:46:19.505318Z digest=sha256:91393ee954886ea224161b342e6c3c71d5da86c23fd82c5469eff8ea53516a97

Observation b7efc35b-e33d-4610-9fea-8e2297ab618c · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.818052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.818052Z digest=sha256:d8d454aa39c18a5a5f3a83c08b757fa01597a97d79c609326114438a55d5b10f

Observation b0b2081b-5f53-4d21-8b42-08d2cc83303b · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.150095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.150095Z digest=sha256:ff975aec74ecce6bdcfce12e890b7a96c98770d259033b95c11869a4053fc908

Observation be7f6567-0ae3-4f92-81cc-1cc6f5346e46 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.322261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:3fe906abb18df544f86b589e3112af80ae9bf9e56b22e5f7d4ef4619108c4791

Observation 32f16455-d6c1-4bea-9e65-a38aff7fcc5c · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.170343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:a18d65cc47b65eb7946b35d596afe62808107e2a0ef491e67a5ca8c326f574a8

Observation eb07d3ab-fd95-49e8-babd-079361566bca · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.608836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:f1e3454f830bbd0af2388f98a9c22b253896aed4f0c3c73691f07eb58a2a070c

Observation 396f883b-4be0-4ea3-8768-009ca23e2b43 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.379342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:3572df25ecd6bdf5bfcb0d52e329c8ecd31586bfc2d2a1b30886ee9e435552df

Observation 6dc7a3a2-5fda-4960-a6a6-90e408212562 · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 111

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.401036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:f77aa87a9a069f324b5e9d7378e646fa068e92f2440eb088a59d0cb38a7e3627

Observation a4d5c3f9-8265-4d0c-8731-2cf740e7b81c · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 253

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.928923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f87fe684129ffa0a764d459a90c3f4c107753fccaa99e210dd7d148eef6e2ec7

Observation 73de4bda-be24-4798-8bc7-9b58a48d0114 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.683787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:1806503886c01a1af52514b1f761b3d955065136b937ec1d1ad72acc2bc4be46

Observation 7dc63e75-60e7-488f-827d-82080ef78b28 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.608318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.608318Z digest=sha256:97befe61c7c9486e65f0f93b4356df7a4221751fdede45f2bf7c992ed716972b