Pith. sign in

Paper Citation Record · LEDGER

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.15778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15778 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:26:25.786061Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71770873-c96e-41b7-8ea5-5f4f414eb406 · outbound

This paper cites Visual instruction tuning,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.595392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.595392Z digest=sha256:e5762b7a125d0cb228595e527f37e8a46d3786c816ee203b4a94d42ae7fb7722

Observation b60c17eb-72d5-4278-9532-11abf5a921e2 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.675835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.675835Z digest=sha256:bb49a60d7df8b5c7d269c57f794bf6cd0c5917364842bdbb2bcd82e5696bef1b

Observation bca80b31-fcc4-4c19-a038-2e67946393ab · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Moviechat: From dense token to sparse memory for long video understanding,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.762387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.762387Z digest=sha256:bdbdf8b92d9b972274fc77b2d07d377775789d3649487de86bd352d198a0d9db

Observation 516af506-3490-4cae-8328-df5f228a802f · outbound

This paper cites Longvu: Spatiotemporal adaptive compression for long video-language understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvu: Spatiotemporal adaptive compression for long video-language understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.851331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.851331Z digest=sha256:7241b26827237ea1099f034193ca0ecfcd3d4abbfa6588ee02ab6d96406f26c1

Observation fc751885-29c3-46e8-a23b-9341c4188ed2 · outbound

This paper cites Adaptive keyframe sampling for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Adaptive keyframe sampling for long video understanding,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.929607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.929607Z digest=sha256:77c20758f87c0e44d40fd9d23fd7fba8b33d2f094fb1ea843455b895e4893aa8

Observation b26a6ad1-2c4c-4c83-a035-d270a02de89e · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.014552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.014552Z digest=sha256:4ed59dbd08531603cd4b322ef259cd0d39754bc722ed4602ab9e4891cdb8e086

Observation b4514555-9fbd-427b-b890-8d02cb6bd112 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.098278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.098278Z digest=sha256:883099d74470deb7b481f13774e49c52b2dbb000819973d8b45e28c44e920634

Observation 619415b4-fd29-417c-a927-e3f6caaa5805 · outbound

This paper cites Multi-modal generative ai: Multi-modal llms, diffusions and the unification,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-modal generative ai: Multi-modal llms, diffusions and the unification,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.184455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.184455Z digest=sha256:7ab229ac8ac73c8b629df2e45ca3c123c5eebb701c7b73434855d5f725f86248

Observation 99a99d8b-abf6-454d-9c75-d4977f9faf5f · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.262670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.262670Z digest=sha256:f3cef205a66f82872bd5d234a97f66d796deeac3fe1b351ed86fe9e0f74da9f9

Observation 52d9003b-0352-48f2-97e9-ae08892790e7 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Fuyu-8b: A multimodal architecture for ai agents,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.319973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.319973Z digest=sha256:bb838dbca92c869877e8b7c091762dffcdff18d652688f0bf57072163adfe98e

Observation 186f6d8e-07d2-4b51-bde7-d135f99109ad · outbound

This paper cites an unresolved cited work.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.409429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.409429Z digest=sha256:84e98fad5744536293d80bf512a4fdbd2297ffdf27b9409553d4e81d093d6645

Observation 397a33f9-1e21-4be9-8587-1e01bb531ae1 · outbound

This paper cites Multi-sentence video grounding for long video generation,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-sentence video grounding for long video generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.490559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.490559Z digest=sha256:4ffddbfb8e3c07f347a33e6fc471f96cf18cd1747539af828952ae9696bfea9a

Observation 06667faa-010d-4fcb-98b3-ba8c654cf06f · outbound

This paper cites Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.568442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.568442Z digest=sha256:d7633261ae7acd1fe6952c8ca78f25a93dfb8f420b8be2282dcb55b2acc435b1

Observation 27f6580e-b61b-4b21-8470-779f462246bf · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.659409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.659409Z digest=sha256:f084a97533544ded7c271afae2874d2d41624afe1c2675c8cf7f4a6441a05c67

Observation b9a81adf-f114-4fdd-b6ea-507aa91a415b · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.726194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.726194Z digest=sha256:0c6218db02050832b5146583f2b789893d07298c01946b6ab432fefc6b7a5b56

Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.782850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.782850Z digest=sha256:2b209eb81ebbb8753fbc93ec052dd6e3dc48d3e043b530540ceacfc2ea3f33d2

Observation bc3e9312-03be-4c05-9bc7-dd2aa0f38df6 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.831412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.831412Z digest=sha256:9a8efdc6071a2c19029a07b2d4c4e0ee54e64932a7535bdedf5e538b8f0da720

Observation 5e94b226-070e-4c91-8f2f-c5abcb729756 · outbound

This paper cites Qwen2.5-VL Technical Report.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Qwen2.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.886162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.886162Z digest=sha256:83fadf044211750fe4ada4f8ff108fceb901a057a48012d7dd0f7718ed2530f4

Observation ffdb04a3-c6cc-41a7-8f42-bc62895b3193 · outbound

This paper cites Llava-video: Video instruction tuning with synthetic data,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Llava-video: Video instruction tuning with synthetic data,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.942904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.942904Z digest=sha256:3c4209926ddd213b83637133882a76dd2f6937adda291a10f55e9d011e5f19ed

Observation 1d97a38c-90ec-48ce-8061-3c336499d9a1 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Deep reinforcement learning from human preferences,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.981811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.981811Z digest=sha256:571b972564e14774498ba5fe83bbe824bb93d70aff10bb264348a11c375ff500

Observation 6ff3a8e3-9878-42e5-9d63-566d7077ad1c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Direct preference optimization: Your language model is secretly a reward model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.055475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.055475Z digest=sha256:4fd677123beff643ef0160344c5db2b8b8914ab369bb51127237eef9a001de17

Observation cab6b849-c83f-4440-976c-279ef2738788 · outbound

This paper cites Modularized self-reflected video reasoner for multimodal llm with application to video question answering,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularized self-reflected video reasoner for multimodal llm with application to video question answering,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.141517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.141517Z digest=sha256:2f33fcfde4cc1a5636c658eec9b1a4bb8ca07db2951b56f2c47c850ae057079b

Observation c2f565c9-4d53-4318-b85c-26a6049b825d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Learning transferable visual models from natural language supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.216497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.216497Z digest=sha256:1170e87ef5cc95febfdc4eb23199fbf7136c261321d455d11f153d2de79a2b56

Observation 349e2906-80de-4009-ae36-b0068cf5dddf · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.266600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.266600Z digest=sha256:b1fd7f4023787290e616b5927b3a6978edfeea1b681993dba9e72a00ecbf4ff9

Observation 7ee821ea-8269-4eae-b24c-152814391dbc · outbound

This paper cites Lvbench: An extreme long video understanding benchmark,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Lvbench: An extreme long video understanding benchmark,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.321124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.321124Z digest=sha256:ce9e2d3c67db4ed213ba69169569d33602f7ba69983e4bc62331f3c80c89c833

Observation bf72d852-b689-4c96-af23-6c98535cefeb · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Mlvu: Benchmarking multi-task long video understanding,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.405821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.405821Z digest=sha256:0ec6a5e0d140cb7f31efd2e305a001aa832bac20bc0c58320116d3bf8c75b814

Observation 157dd742-d96c-4a01-82fe-16c9341215c6 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.477899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.477899Z digest=sha256:4765792a16908101b9e3386f7a520c26f5c0c050881e5940b090a44a944bdec6

Observation e56562a6-61e0-413a-9acd-5a1a5c1133d2 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.559526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.559526Z digest=sha256:a71173b231404d12216a81b014a5040d6820c4547319282fbab09d13e74b5d7a

Observation 8e6a8668-c044-4dbd-8956-2902276bf543 · outbound

This paper cites Cg-bench: Clue-grounded question answering benchmark for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Cg-bench: Clue-grounded question answering benchmark for long video understanding,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.637198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.637198Z digest=sha256:20475f53b9a648fa8457368ca8ee553b27d39c4104aafa996ffcc8285b4a77c2

Observation 6b7562bd-dc6d-4bb7-b34c-a1c11e98e26e · outbound

This paper cites Videoitg: Multimodal video understanding with instructed temporal grounding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videoitg: Multimodal video understanding with instructed temporal grounding,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.723913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.723913Z digest=sha256:f95b18ebdcbb00b11a31d63bd8a85a27a87708972765c7a10f1831bdede730ac

Observation d76e0c01-b8d1-4867-8c1c-aa60e84153fb · outbound

This paper cites ordering,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding ordering,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.786061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.786061Z digest=sha256:a273dfa3ad6f4e9e37e25f0cbbfb6867b8d7ebe4e797745c563e61c33098293b

Pith citing papers

No inbound Pith citation observations are available.