Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:27.201212Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 116 outbound references and 8 inbound Pith citation observations for arXiv:2504.16030.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:27.201212Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.405635Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T01:19:20.332342Z
100 of 116 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fd0846cc-809a-4cf8-86c4-31118228f770 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Marks, Chiori Hori, Peter Anderson, Stefan Lee, and Devi Parikh
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 257cfb55-f212-4089-8193-d20bcafc34f9 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d60b12ad-da72-44e3-bca3-0a1283a0944c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gemini: A Family of Highly Capable Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8503c045-d28f-44ce-b3cd-e1ffcd0e39a7 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b66b2b-b51d-4da7-bb04-aecaae912591 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e4853b-eebf-4d64-95ec-29d260a6bcd6 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2.5 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa3eead-eedc-4049-887c-697ccd1d3d82 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f23cd0-dbda-4da0-bb09-6a7f53641bbf · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Language models are few-shot learners
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a700740-b763-49ba-8d96-5002fc650141 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Sst: Single-stream tem- poral action proposals
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637d85f5-6c9b-4253-8e3b-9a2c22076285 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Quo vadis, action recognition? A new model and the kinetics dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a73a2a-f219-42cc-af70-7f311920b156 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale A short note about kinetics-
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3906aa8-2eb9-49e1-bea1-c5fa40ad6293 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c4cc69-e049-4d5c-b875-416653183106 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Videollm-online: Online video large language model for streaming video
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9cba1c9-195a-4afb-93b7-0eb6c5e9d4c7 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77644c62-c134-4fe2-9834-7ab34e3a176e · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b99d661-5a1b-4200-8165-41a5236dbc2c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850fc669-e29b-47e6-8e96-0bc22d9a53d1 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28bba35b-989d-44c4-8fc6-7864199704c3 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e651ea4-e587-453e-a528-de11543ad32b · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unsupervised cross-lingual representation learning at scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70153ca-b921-4f37-9612-fafd25b58018 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8233268-17be-4faa-9c2d-eec836006a46 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale An im- age is worth 16x16 words: Transformers for image recog- nition at scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfad65e2-bc85-4f24-a418-c4c601ef3d18 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8476e595-0b83-4c5e-8269-2ad8f2e8c03d · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dc06081-1aae-4958-b77a-dcdb6a00ab8e · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2eda36-8332-4475-9a88-c8091db910b1 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d46e76a-8a11-4884-a8e2-d70247e88d0e · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online action detection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bd417b-6f8d-4d82-9668-a01a5d2d271d · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 345faeed-0bb8-4911-a8ff-dbf4bde53adc · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale The ”something something” video database for learning and evaluating visual common sense
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1aec64-51eb-4c9b-994c-805a2a69db50 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Hello gpt-4o, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f364702-6318-4367-9f49-f75d0c5cc226 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Activitynet: A large-scale video benchmark for human activity understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bbc3a83-cd54-4a0b-b16c-5a1c2b391f4c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Training Compute-Optimal Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0aa7b7-c973-4720-877a-f412a50cbdda · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Multimodal pretraining for dense video cap- tioning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 428d89ef-547a-4ba3-825f-4f4f83c794ea · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online Video Understanding: OVBench and VideoChat-Online
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f299d2-6ecc-4193-a4bc-91bf59650387 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Zamir, Yu-Gang Jiang, Alex Gor- ban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4cdbe42-2727-42c3-aeb5-d2d4787d36cb · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Cag-qil: Context-aware actionness grouping via q imitation learning for online temporal action localization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641431e4-bdf7-480c-8305-e6399750a792 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Scaling Laws for Neural Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c304600-fece-4dc5-8a03-72e345afe0c0 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Dense-captioning events in videos
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deb05087-b1eb-4e89-ab60-1d6b2374109c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LISA: Reasoning Segmentation via Large Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db32a862-f260-4ab9-b921-d6165fa57ce0 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c787d8c2-8e6b-480a-8cd7-3178a0f78d2c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f7775e-a62a-46b2-80ac-5fe3d9b007aa · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LLaVA-OneVision: Easy Visual Task Transfer
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0306f9d0-0158-41f2-ba01-58846d135246 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8c2cc3-be2d-4146-ab18-8a688712a39a · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoChat: Chat-Centric Video Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87110916-75ce-4f5b-a40e-85b5ea15f5c7 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Mvbench: A comprehensive multi-modal video under- standing benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73670dad-9bbd-4051-b0bb-0b5e114ad508 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e60bc7-4ae4-4d0f-b46d-9410274097ba · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale A light weight model for active speaker detection
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ad1bae-9999-4bdc-8a78-abe754dffd57 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b345a33e-efb2-405e-b171-865a7a5412c5 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5d5368-5f9d-4747-8bf3-7eb8169b8a2e · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VILA: on pre-training for vi- sual language models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53da6c4a-db92-437f-bbd4-4015c36bba8d · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Egocentric Video-Language Pretraining
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da39fea7-c1e7-411a-8328-376f24a542ef · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Univtg: Towards unified video- language temporal grounding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2736919-88c5-4959-81bd-1faebf7becef · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Improved Baselines with Visual Instruction Tuning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbcf587-60c7-4ea5-b55a-a7d3ee96cf44 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Visual instruction tuning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a39f103-01bb-4617-8264-92f69a929d96 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamChat: Chatting with Streaming Video
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e984bf2f-6b36-497b-a987-99e4c1bbab05 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19029ba-75b2-4a8a-bac7-c146adf0e1f4 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e081aa2a-6b19-4c75-a395-b26121927e2f · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704b455b-9793-4287-9fb4-01d6fda1d00e · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ab3ac74-9a11-427f-b50a-af80fb355c76 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb594618-7bd0-4f8e-8580-f6e324dad4b3 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Soccernet-caption: Dense video captioning for soccer broadcasts commentaries
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d01ed2e3-3edd-48e3-a019-af951c800f96 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Introducing chatgpt
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4341a45-59f6-4dbb-bb11-777f7b1e4312 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale GPT-4 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d084c85-cd3f-4b17-a10c-f6facc010116 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gpt-4v(ision) system card
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c8a06b-c802-4e9b-93ba-cb1eb361dffb · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Py- torch: An imperative style, high-performance deep learning library
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a3bf00-bb5e-4d79-b235-3401fa466da5 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Perception test: A diagnostic benchmark for multimodal video models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e902d5-a7ab-415d-bdf6-4ba3878abb42 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Streaming long video understanding with large language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca7c4f55-dbfc-4924-a060-cae46446bc3d · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa23ca37-cace-4089-b2af-98460344288a · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Improving language understanding by gen- erative pre-training
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef305958-3548-40f9-a71c-2293c9743234 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Language models are unsu- pervised multitask learners
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b4012d9-27a3-45ec-8565-100beaa417eb · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Learning transferable visual models from natural language supervision
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28fbe0e7-2d3f-49b1-893d-06b3fa2deaf2 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Robust speech recognition via large-scale weak supervision
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 056449bd-b9ad-4b29-881a-b2fc58207442 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MatchTime: Towards Automatic Soccer Game Commentary Generation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379b02c1-19be-44e4-acd5-8c204cd94060 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5e1b67b-44c2-4022-a2ae-042f7f86450f · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55b5798-3b4b-4500-9a7d-8198f6e820ac · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online real-time multiple spa- tiotemporal action localisation and prediction
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation daeffac6-c7c7-4239-8447-2dca8966603c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8689b0fc-8e52-4dcb-9195-5afd5fff4288 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LLaMA: Open and Efficient Foundation Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a5b472-9a3a-481e-8ed3-550f9bbd6367 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63131da-0943-4629-b683-176fe23c49e8 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83e75d0-0d1c-48b8-a29a-e92780fc3161 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d822a3-1677-41ea-a25e-dce7d9195ba0 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction for- mat, 2024
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eb4a92e9-fd24-4b10-af18-29c4ca55e21c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a23ce1f-3f8c-438d-b923-afd5063047f8 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c1d9e19d-5e9a-4611-bc4c-8983f837b543 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Next-qa: Next phase of question-answering to explaining temporal actions
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 459444c9-7720-4835-9eb8-e9c9effc9025 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Streaming video understanding and multi-round interaction with memory- enhanced knowledge
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0fe113c2-9cf0-4d41-be15-037a91c09db0 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2.5-Omni Technical Report
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07193edd-8de9-4dfa-a029-9731f769c61e · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Ad- vancing high-resolution video-language representation with large-scale video transcriptions
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0803101-3fc9-4742-970c-411378aad026 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Vidchapters-7m: Video chapters at scale
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6f5dd80e-7e9d-4929-9b80-b19945ea7bc1 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 65debb99-407b-4073-adb0-cb6930d30c85 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2 Technical Report
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc2f568-8612-4aa5-8c20-724dcd182927 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f81d77e-263a-45a8-8b99-488b5add7216 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8b02ac2e-064a-4ef4-85b6-b71b2296a5c9 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc8537d-1432-4ac2-81a4-50ab1f6b6782 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699fbecb-687c-4872-acaf-50204469c765 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MERLOT RESERVE: neural script knowledge through vision and language and sound
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6305bd7d-de47-49d2-af13-1964a3e9a72c · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Sigmoid loss for language image pre-training
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 514febb0-d8a4-4197-9b0f-93e6126e6316 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8283c1e8-3d87-4f24-b9c2-f5773644ae1b · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdfee9e1-d589-47d1-82de-7faa174aa7ae · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation babcce6b-24f5-4b23-9bfc-00867f7882e0 · outbound
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79481876-9f37-4128-8e6c-513d58e8d9dd · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b7594dc-5632-4def-816d-b98de7c8a0bd · inbound
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cecdb1c9-4d81-4dd4-bb48-b83e4bef12d2 · inbound
EasyVideoR1: Easier RL for Video Understanding LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dbba1ab8-add9-4c98-a3ef-7ee2c9dc6af4 · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84dc0ba4-7c41-432b-8e86-6922425289cd · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 69338436-7dfc-4746-9633-f9fba57f9c20 · inbound
An Efficient Streaming Video Understanding Framework with Agentic Control LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aabb4aa2-8816-4064-8777-56c364f342e3 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d29b28de-f054-48a9-8b66-5a1aa809c1e3 · inbound
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.