Pith. sign in

Paper Citation Record · LEDGER

Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2408.14023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.14023 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.119516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.765108Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8bbff39-2518-420d-8b3f-584957ae697c · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.536953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:a90fb7190e9507fa701c604e1e440ce86849ef625be1a59282fbf24de89cf747

Observation 8ce1ff7b-9ec1-483c-9524-5056486179ab · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.501093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:6344fa11b4c3366a2297130a335fd19bd0892d5d8cf7e5dcbb57e4e3421dea26

Observation 23d85641-bcfe-4bb7-b319-6c101585e709 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.716663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:c28f7f9ba7932d063b13f0a60649bb6c2a91f74ea3668ebc5cc3546f4157ddf3

Observation a1e3ce7b-6cf8-4833-9414-91be823de78e · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.714773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:cdbeac16acf32f1985815e8434161a9a575f0c6393873e8f8a4b8337e62a56cf

Observation fe4d3805-a816-4272-b56a-86137f4f33ec · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.255829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:aa0d2ba9b5ad02271a145388bcc08189995a3c90e7c0d3d0641687d9c55c6c7a

Observation 748aef16-b096-4671-98f1-b15f73cda2a0 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.119516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.119516Z digest=sha256:3fce83d5266746ad38565a4280d573bb9ba90cd4222946b442a1fe2b35d38043

Observation e3cfbcff-8b4d-487f-a123-e2ae1da5721e · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.907077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.907077Z digest=sha256:6c93c68eb4456a3d74c78ebf1348a2a71cdde266bbd2c5d7ac15bef90df933b9

Observation ce17acaf-da91-44b8-a33a-65f9f5f53e0c · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.981605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.981605Z digest=sha256:39a6d086de03298ee315b5d5906b6cbe083ff4c17e575bc157bc4ed7cdf478e1

Observation 4c265e10-9dcf-4610-9ba4-825985f4bb4d · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.775593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.775593Z digest=sha256:4ab7406e551fc7f861a3341e39b891feb55bbbbb605a123076793a9e182eb2ee

Observation 889779c6-2937-4cab-bb15-295b9ff8aeea · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.421786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.421786Z digest=sha256:e4aa85f75efe74274231067f167a8d09ad7eddb4ef4e3129dab30967dcfa4107

Observation 97b1818a-42de-4fdb-820d-985568e79bc4 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.431837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.431837Z digest=sha256:6bd65e71e6966c39198026954c181dad61fd37f75755df569639ab4e78167004

Observation 7952014f-ff33-40b3-bd24-7d4a7bfe2962 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.491781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.491781Z digest=sha256:d0d64de069b9cd22f7c8facf9acf96d679f1b7aeef2424db8d2df697aa987162

Observation d632ca7f-a4c0-4de8-9889-4d3d58abaa42 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.989110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.989110Z digest=sha256:57fba7f69d7310b7dcd40e53052dd4666f7c577fa7d00aaa66af608cc34f0cf3

Observation 29079ea3-7f5a-45e3-b673-b96d566aed81 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.868717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.868717Z digest=sha256:e946d70d846db99d87c8ccf0a6cc6c14abf0efef9e25557e63145d880e87ecdc

Observation 55a5ff53-44d3-4273-8475-82950400326c · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.340057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.340057Z digest=sha256:9966026075b5f1499d6ea5e6f69471aca0f8d26fee7881b9b2b2cf2705c8cb81

Observation bb93b653-36b1-4e29-a744-433a9651ab04 · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:00.121528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:00.121528Z digest=sha256:171dedaacefa46b043dc0bc74d8dc99ef2e8cfe25fc97bb335208f4d68d6ae01

Observation d8855dac-cc08-4c80-8e64-ec425fe59203 · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.951565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.951565Z digest=sha256:5de3e169218de39db78dd0c66f1b1636c61712717986efbded8439785bd66e8b

Observation 444ed85b-6e62-4477-8c1d-a427b44b15fa · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.727170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.727170Z digest=sha256:e4635c987f0b5c3a5347163777422ecf7af872839639abede21acc2e32a55d5f

Observation f875cfc8-e833-497e-87d3-73a3ddd9fb79 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.524954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:17b733aa5f5a343b6e72bb2fd3055352c27cba8e04f8f1ee73698eff04b9af34

Observation ed57979b-9d67-4aca-9efa-d64112ba6835 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.419248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:7b4dd9ceee102f8fdbc4bc87ae0ffc115a97cdd68e6e6e73dddce67c68e6a0d4

Observation 1c61831f-6a19-4f9f-a654-df1feb208d45 · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:04.334732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:8f0852caba9c5201470a4c50061a7df4c77c1b531991d00a6ecd132cd6a65341

Observation f2f9886c-ae29-49aa-9743-76d23ee107a3 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.648482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:13b217d148b12728c57818255d6c6f353f9fbc2ccf5b55f0b366e8a2d428f966

Observation fdd603bd-a208-4c6a-bfda-351152e3382f · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.039662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:3140b6f736f3a094f93b68474a078edd6903d333b01254d370005a19961a5ba0

Observation e0758a7f-de1a-4979-95b9-fb764a784d0b · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:56.556078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:991abb8774d4e7b00236972b6eebe6c1639d42df22b914bc5f00b629c2dbe973

Observation f8cf6d9b-4edd-435d-8e9e-65ffefa10bf4 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.454668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:7306e72cfaa291baf43952181bb2fd3a47f24630ae6b09c404262736381cfafe

Observation 0a27ff8e-79be-4bdd-9476-ea62c810ea37 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.707528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:e17a74a0971e8e20374dee8f1558bfe7241b89eff0dcea948fd3e06f1ba89a98

Observation 353b1242-ee1c-4f1f-8d1c-6fcd19ab717f · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.438058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:e9d37c2fb442f0cb8869ed2ab2f7e7f944c8d7aebd219f2558395b3824c22c71

Observation 4fd44adf-2d77-4b56-b812-061e54bb2403 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:14.924683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:b8ad972eb58f9ca9e3d55226c7f3b49f6117e18c17d06ee5b788eaef552fb5d5

Observation 60c46325-5641-4837-a59b-d9537c6f88d3 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.748177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d31ac09c8c81a4daa510723c54bfe463e7a05691514ff43474d1b1817d25dcc0

Observation 22a9fa5d-4c61-43d1-8f22-ee450655e785 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:08.028902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:79e47baa440985214416e100323a0dd94c01e78b10e5478d85f472aaaccbf5bd

Observation 433ae2e0-4eaa-49c3-a4f7-36f6b24270f2 · inbound

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models cites this paper.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.202123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:add1eb452b27825590ad0dad473ebca51ac3276e9183df702153024c50387b88

Observation 172a8c6a-851a-49fb-8c07-0e34cf02566d · inbound

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining cites this paper.

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:02:50.352612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:16:42.001361Z digest=sha256:36d9db5ea8b99e767c494587d333958ebb6515e52c5e8e2b3aad3933b625763c

Observation a60abe6e-cf91-4903-b0b6-abe486c2ae41 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.358625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:627cf482b569ed05db1d7511e04529e4924daed2f3fa752a6eab6d9e290e8a24

Observation b06a113c-fe5c-436d-904e-14032ee29aaa · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.056825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:36371c3178116c97f949e5e071ed2bce4a8260712007aa4a7231cdc9a13690be

Observation 612534ff-5173-43db-934e-1bda7fb6300c · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.974132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:fa00cf0a4ee7dc76232ba74b1b2581dcb93a44876e255d228247f03cbf007c0d

Observation 1cefe5d0-9b92-4e34-afdc-c43f7e417913 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.452455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:cbe19606734599aa4ec4d4f756b0fb34c8323b0d713667fd6000bf854fc64c85

Observation dbaf52bc-d04f-4e4f-a079-f604648b05b8 · inbound

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams cites this paper.

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:39:46.766518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:42:58.861599Z digest=sha256:abd51155e77ca0ba6fc9276bdcc5708fd7c61b53254df3804aae3ad2a9cb62ca

Observation 8611e1e8-588d-4426-9146-fe6284f1944d · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:21.335724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:6281a86778003f1a92ca346aaf53d0532f4e5ea8d404ea74ec65b38247a47584

Observation 440a1036-f026-47e1-8c1a-cf37f656f759 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:fc4690dc7dedb4ca2eeb704c04481b5c69aae40a137689ec97d2ce4780eab761

Observation 2bbef040-3c0b-47ca-98da-46922382d900 · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.181503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.181503Z digest=sha256:d5fb7e114822f4d365669b724b71674b6fde0cbe02e98e4ecac7d1e684ff0ba4

Observation dc306c73-3c39-4073-9433-30caa848ad63 · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.541193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.541193Z digest=sha256:751b20b6986a61cee1cdd22309c66704338d3a296d1335eb808b9dd99cd19463

Observation 2adaf088-b632-4f60-a978-61943e5b8a5f · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.215228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.215228Z digest=sha256:d0787ef1f95288d5062e4c7d652a32304536355e49bf7dc34056cda8dc013a7f