Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 89 inbound Pith citation observations for arXiv:2406.04325.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.459440Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.677235Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 69acf894-d236-4b03-910a-a38fbbe25c9d · inbound
MLVU: Benchmarking Multi-task Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9c1b1b1a-f9be-4f69-9ab9-2203b31acd9e · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 13c3cee9-ad9c-4bf5-9fc0-a5837391ad38 · inbound
LLaVA-OneVision: Easy Visual Task Transfer ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 74c89a1a-f1f0-4b83-971d-a272a19e5b5d · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 37de3bd8-d552-455b-8ae0-d85ac3b9b3dd · inbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation adafa0e3-6f81-4449-ad51-6dc21c67c500 · inbound
CogVLM2: Visual Language Models for Image and Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39825b21-8a01-4bc2-bf4e-a2fbf81ed070 · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b2e0239e-d167-4485-84a0-1d2e663f2c9b · inbound
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eb74b757-d340-4bc8-96d8-8ae242490ded · inbound
SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81927198-2aef-4d3b-ba16-d860bcfcb006 · inbound
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411d20e4-f21d-4d13-a7f7-a189a9d5d80d · inbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ae6eec-bd68-44f2-ba7a-00b1f27df4a0 · inbound
VideoOrion: Tokenizing Object Dynamics in Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d115cd80-01c9-43a4-be4d-04a3b491d210 · inbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac0f485a-7c69-4684-b804-b43b6a35248c · inbound
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130a81ae-75d5-47e6-afa2-7b1689c72968 · inbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c023be49-699a-45cb-a054-d045b7b2efcf · inbound
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039c3b39-f01a-4617-964a-34ba3e6b9ec7 · inbound
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721f1b58-4bc5-4a1c-876f-64120cd3f48a · inbound
Progress-Aware Video Frame Captioning ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111ec42d-5624-46bc-9f4d-494adccba1d7 · inbound
HunyuanVideo: A Systematic Framework For Large Video Generative Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ddbd9cb0-6eba-4ab4-8f0e-b6cb83285c48 · inbound
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885d6d3d-6697-418d-8e0f-df75883c3c6d · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235ff406-5120-49e3-9bd1-921fedb6817a · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 514a8eda-260c-4fda-a60e-170b9fd0e581 · inbound
Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebfefe2-7ddf-48c5-b74f-a80380af84ad · inbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2949096-c0f9-4c23-990b-a3306ac34c6f · inbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f00bb27-f81e-463f-84e7-8ec8086f4590 · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd22d433-36fd-4676-9c06-eae22651bc0c · inbound
TimeRefine: Temporal Grounding with Time Refining Video LLM ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60fdd45-4d6b-417e-ba35-1b90bb74b59a · inbound
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a07dd3-0067-4586-a54f-f922941b0114 · inbound
Can video generation replace cinematographers? Research on the cinematic language of generated video ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d254a1c-332f-4cae-b067-0cf07920621b · inbound
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e07e8417-78d9-4864-b5c2-a8077e335d3b · inbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b28a0493-06ed-4eaa-bb9d-23cb296d62b9 · inbound
Do Language Models Understand Time? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c47207-6c7b-49ab-bc00-ebfba7a9d09e · inbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e65a34e-54fd-4d6b-99d1-356c8031ba8a · inbound
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e75adb01-ee7a-49bd-b510-13c046284058 · inbound
Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4c4aab-e7c2-4666-a51d-e9e39f8c8b7a · inbound
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29bee933-1e99-43d5-ba38-f71ccc284572 · inbound
Online Video Understanding: OVBench and VideoChat-Online ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e551242d-40a3-4331-bc06-41df8990ed44 · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55563787-0898-4098-923d-ce3997e27789 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7ade2e87-b39e-4fb9-a3b9-2a3bacd370d0 · inbound
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76dc3b8-f695-4eaf-9c72-ed9ca8af8034 · inbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd511f4-7199-41c0-8ae9-99e0872b6a71 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation faf5498b-79b0-4d6f-b1cc-31a0ebcfeb59 · inbound
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 431f0539-0b9d-447c-9712-5c0ea0be68f2 · inbound
Temporal Preference Optimization for Long-Form Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc1d275-3d71-404b-9f56-69b217851a25 · inbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0610ac8-318e-447b-a2a3-8986869d718d · inbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d94dd3c6-0fe4-4d42-9c67-1d61c89559f3 · inbound
$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84909031-5f6e-4f8d-8289-8991b2a172e7 · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48653a7-17e7-4649-b1f8-6273adf43fdc · inbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 299b4d8f-4f58-4ea2-b784-d23ec64689c3 · inbound
CoS: Chain-of-Shot Prompting for Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a3ad532-9edd-4f91-a7da-72196dccc641 · inbound
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f907f7d-2050-48be-abc9-27c2a2d25941 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c44d3654-2579-4e20-b1e4-fb6f56135dc3 · inbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3d83d8-9cee-4397-a0ce-37873dbe3792 · inbound
Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1efd48a1-dfb1-4daf-b9c6-e74ab80eb44e · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9cba1c9-195a-4afb-93b7-0eb6c5e9d4c7 · inbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e182584c-e789-40c6-b5c7-1d808da72705 · inbound
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a610f35c-47f6-427a-bd63-3ff7d1c30a28 · inbound
VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7dfa941-13be-4f45-b591-a253d06c7c01 · inbound
VEU-Bench: Towards Comprehensive Understanding of Video Editing ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 329b37c3-0136-49cf-a6d1-28dd703339eb · inbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b28086b4-d914-4043-8d35-b1de08144eeb · inbound
HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a20c8e5-9658-4394-8b3d-e13cbc5bd47c · inbound
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7edd419-47c9-47fb-babc-346f8edd4be0 · inbound
Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b005e4b9-97e1-43e2-939f-20470ea8ce43 · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b07fa147-f2c7-48ef-9f80-c9173bc1072a · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c02465-053a-418b-b21f-3b8822bafc83 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef9d158-91d2-4043-b920-430cefab504d · inbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2782dec-d4a8-494b-bbe4-8deb106d887f · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7381ecb4-c8b7-408c-aeab-a652d13bdaa3 · inbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9cd646-4ba3-44ce-8828-ca0f1aeea3ea · inbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6f1df4-7b18-491e-9890-07aaa8a4c393 · inbound
OutDreamer: Video Outpainting with a Diffusion Transformer ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67515cb2-5b29-4a06-be65-585b37f24353 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · inbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56eef71a-7d12-400a-8d99-5b6bc27ffb9a · inbound
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7255e154-2e35-4724-8474-3cd9e6b63110 · inbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457862a8-afb3-48f7-85d2-efa803b70330 · inbound
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea14db05-ebb2-4935-96bc-d8559c0b8a65 · inbound
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aa4d782-5645-4153-abea-5c2f8d9af55e · inbound
Sample-efficient Integration of New Modalities into Large Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bf05d2-7dc4-4c8a-942c-c4cdef0977f2 · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642d51a8-20ff-417c-9e70-b963f6116528 · inbound
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6e744474-39a8-494e-8bac-3deef550e724 · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c858b49c-ac43-48b9-b779-a0dc22e0868f · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6241f389-bd1b-4315-b1d4-789929997f9a · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 124ed411-0beb-4cc4-8191-f166e6c26330 · inbound
UNIVID: Unified Vision-Language Model for Video Moderation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d9cb0997-5ccb-4b5f-89c3-df9657029bc5 · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fab0ff6e-ec31-4328-acb1-0c8a0f266401 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9268b81e-016b-4aa9-9175-b18cd1c6f956 · inbound
TimeThink: Reasoning with Time for Video LLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddeb7df9-80c5-4bf3-9afe-bdd40ebea329 · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 581d0e6d-922b-493d-957d-1ec240bfe891 · inbound
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.