Pith. sign in

Paper Citation Record · LEDGER

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

As of 19 August 2026, this Paper Citation Record lists 100 of 116 outbound references and 8 inbound Pith citation observations for arXiv:2504.16030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16030 v1

Coverage vector

measured 100 of 116 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:27.201212Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.405635Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:19:20.332342Z

Reference resolution

100 of 116 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved88
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd0846cc-809a-4cf8-86c4-31118228f770 · outbound

This paper cites Marks, Chiori Hori, Peter Anderson, Stefan Lee, and Devi Parikh.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Marks, Chiori Hori, Peter Anderson, Stefan Lee, and Devi Parikh

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.761886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.761886Z digest=sha256:8b8d451574e37c82edd9ea7b6c24a43e619d592fe2a2d934e6939f7563a8a9b9

Observation 257cfb55-f212-4089-8193-d20bcafc34f9 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.767470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.767470Z digest=sha256:836a036a0daa4689df7eae629f27d109013bc5561c9ec147efdc316ad270edb3

Observation d60b12ad-da72-44e3-bca3-0a1283a0944c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.771846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.771846Z digest=sha256:020eceb4960ac9888d44c00ac04f2cbbdf25786fa8b963b2243f4a0f2a29c1e4

Observation 8503c045-d28f-44ce-b3cd-e1ffcd0e39a7 · outbound

This paper cites Qwen Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.776536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.776536Z digest=sha256:41b1654a090c018844f81075c8a9c89326029eb0b9fe273dab0ffcf9a2f9a52e

Observation b1b66b2b-b51d-4da7-bb04-aecaae912591 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.780847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.780847Z digest=sha256:0c1799ae45e27d2668bb13e0fb56ab02dddda57d0ae25c40f116487877bc692a

Observation a0e4853b-eebf-4d64-95ec-29d260a6bcd6 · outbound

This paper cites Qwen2.5 Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.785092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.785092Z digest=sha256:36cc3ca302a2ca9c1f8f9c6b6976d5c8a70f9d0caa5b8933e449ec5c8efe58ce

Observation baa3eead-eedc-4049-887c-697ccd1d3d82 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.789286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.789286Z digest=sha256:46168f8c279f2ca588b02d5c3eed89f756919e540fc975c19dd6172e982268fc

Observation 02f23cd0-dbda-4da0-bb09-6a7f53641bbf · outbound

This paper cites Language models are few-shot learners.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Language models are few-shot learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.794090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.794090Z digest=sha256:17a755e5002476ad1bd3b3938fc39e1250b9b3b29aada00192ef05a6d3349caa

Observation 6a700740-b763-49ba-8d96-5002fc650141 · outbound

This paper cites Sst: Single-stream tem- poral action proposals.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Sst: Single-stream tem- poral action proposals

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.797881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.797881Z digest=sha256:c643520a3abbba1bab886aa1c3eb5ae87eaa6319ea3c744a6fdc7c1a5d96d878

Observation 637d85f5-6c9b-4253-8e3b-9a2c22076285 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Quo vadis, action recognition? A new model and the kinetics dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.801885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.801885Z digest=sha256:950f76bcf9d3a2317efabaa5edd897caceeabf9000b15a548f14362bc3364034

Observation 39a73a2a-f219-42cc-af70-7f311920b156 · outbound

This paper cites A short note about kinetics-.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale A short note about kinetics-

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.805640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.805640Z digest=sha256:3cc8e4cbb6d05d60e06a979797ed8f2350d1c4f2f622c0b1799e10f92a572955

Observation b3906aa8-2eb9-49e1-bea1-c5fa40ad6293 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.814227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.814227Z digest=sha256:b384cbb57130087611717b01bfc691d0e35776ef25228a56da79ce6114e04d18

Observation 17c4cc69-e049-4d5c-b875-416653183106 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Videollm-online: Online video large language model for streaming video

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.818806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.818806Z digest=sha256:b2abd01e25a4f5297e1e46791d5ea60c9b1e61a6d001143ddef5e7a2c06cb036

Observation e9cba1c9-195a-4afb-93b7-0eb6c5e9d4c7 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.823292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.823292Z digest=sha256:89c021933f2561572264357d3058977885fcf54b4d5819adb3923c6eeb00c469

Observation 77644c62-c134-4fe2-9834-7ab34e3a176e · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.828107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.828107Z digest=sha256:80a264f5a2ee01b0ba857146fe73d2b143b7128ab22b15f0d5aa06e9fa92d101

Observation 0b99d661-5a1b-4200-8165-41a5236dbc2c · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.832033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.832033Z digest=sha256:512fed88e76b63354654ea71492f160f5d7fc20aea4dc5582ee55c58dbc11949

Observation 850fc669-e29b-47e6-8e96-0bc22d9a53d1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.836191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.836191Z digest=sha256:67e791bfed13649593a1f91a7ebafe2fff7e21f97814fb47b4a48624b85ab40e

Observation 28bba35b-989d-44c4-8fc6-7864199704c3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.840455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.840455Z digest=sha256:035f2ce18e8138e80946b46573a13afdf8f9117c1b9651ad0d27c28691c18165

Observation 6e651ea4-e587-453e-a528-de11543ad32b · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unsupervised cross-lingual representation learning at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.844728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.844728Z digest=sha256:6bf1961cadff09aa03f1db7b8c931f8cdcc0fb329f91a709db7a07e2138f7cf5

Observation c70153ca-b921-4f37-9612-fafd25b58018 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.849493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.849493Z digest=sha256:ab880a9fe1fe19d296f9b93e5df609c888e071acc2f1d19d1ddfbce8c572d507

Observation d8233268-17be-4faa-9c2d-eec836006a46 · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale An im- age is worth 16x16 words: Transformers for image recog- nition at scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.853854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.853854Z digest=sha256:e3cc2f244dd4ce80f106d5e185d6ce882d6585938c8c463a7639d6e5b0983f78

Observation cfad65e2-bc85-4f24-a418-c4c601ef3d18 · outbound

This paper cites The Llama 3 Herd of Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.858399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.858399Z digest=sha256:394903db5f07554a66004c04ef798dfb04bf3d92ecdc8bc73c92a46852a86d24

Observation 8476e595-0b83-4c5e-8269-2ad8f2e8c03d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.862576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.862576Z digest=sha256:9a8a07e7520f9ea5904ea1bb98ba0d15513c366b71884e63215dbe59783d513c

Observation 8dc06081-1aae-4958-b77a-dcdb6a00ab8e · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.866575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.866575Z digest=sha256:850b6d7561ad8d998f43c53c033350e9d63da064818e47e58e04c3f3820f9863

Observation 9e2eda36-8332-4475-9a88-c8091db910b1 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.870990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.870990Z digest=sha256:2f3876abe7dbc8f4c16e997ab8ae7b8ac997033c0dad9468bfbd2f38ac612b81

Observation 4d46e76a-8a11-4884-a8e2-d70247e88d0e · outbound

This paper cites Online action detection.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online action detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.875352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.875352Z digest=sha256:f7a2a94efb6f987370d0029917e3d6fa008a1b0c6a93d61db94c88cd17b1c052

Observation f0bd417b-6f8d-4d82-9668-a01a5d2d271d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.880005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.880005Z digest=sha256:c3cf10963bfb09158d64944bda60fadf635474990f954d93c6f64a1e23247e27

Observation 345faeed-0bb8-4911-a8ff-dbf4bde53adc · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale The ”something something” video database for learning and evaluating visual common sense

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.884157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.884157Z digest=sha256:72a717cf4c6113a8ddb2c18bfe8e1f97d3ab507d5a09e9aa1b5af89ac892a41b

Observation 3a1aec64-51eb-4c9b-994c-805a2a69db50 · outbound

This paper cites Hello gpt-4o, 2024.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Hello gpt-4o, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.888082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.888082Z digest=sha256:ad77f8c341f1fcf0ba86e0d52de5eab18a8cb17316d5c508446370558a566ece

Observation 9f364702-6318-4367-9f49-f75d0c5cc226 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Activitynet: A large-scale video benchmark for human activity understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.891983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.891983Z digest=sha256:05520a1c590ee5ae131f406dec9285fb436085544c28c6c0f84957330147624a

Observation 1bbc3a83-cd54-4a0b-b16c-5a1c2b391f4c · outbound

This paper cites Training Compute-Optimal Large Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Training Compute-Optimal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.895809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.895809Z digest=sha256:b55add55c7b17383ef9b2cddeecab23f8dce17f70c0be72ddb4a0d4323a7609e

Observation 8a0aa7b7-c973-4720-877a-f412a50cbdda · outbound

This paper cites Multimodal pretraining for dense video cap- tioning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Multimodal pretraining for dense video cap- tioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.899907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.899907Z digest=sha256:2eaabd2b5e5a60a20ca3dfa0c3c6e6138506a2132ba767a82e012daeffca86c3

Observation 428d89ef-547a-4ba3-825f-4f4f83c794ea · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online Video Understanding: OVBench and VideoChat-Online

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.904234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.904234Z digest=sha256:0398c0434f8fc4198c10bfe2266a07208fada821d54a10a25aa55d4e8445d149

Observation f4f299d2-6ecc-4193-a4bc-91bf59650387 · outbound

This paper cites Zamir, Yu-Gang Jiang, Alex Gor- ban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Zamir, Yu-Gang Jiang, Alex Gor- ban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.908553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.908553Z digest=sha256:8baba247c892cfd0dd03b906a4cb6a55a051f033320b9ad6ddb0801c9a96fdc2

Observation f4cdbe42-2727-42c3-aeb5-d2d4787d36cb · outbound

This paper cites Cag-qil: Context-aware actionness grouping via q imitation learning for online temporal action localization.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Cag-qil: Context-aware actionness grouping via q imitation learning for online temporal action localization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.912817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.912817Z digest=sha256:b94ce907c9df94a3c93b419665e31d0ea2800db8ccb98f5ad68c6117f36b0deb

Observation 641431e4-bdf7-480c-8305-e6399750a792 · outbound

This paper cites Scaling Laws for Neural Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Scaling Laws for Neural Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.916707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.916707Z digest=sha256:8f59eca94590a1125a67a87bb55d2fa9f1e96dc2f546ec44f8f7ed455661a74c

Observation 6c304600-fece-4dc5-8a03-72e345afe0c0 · outbound

This paper cites Dense-captioning events in videos.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Dense-captioning events in videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.920757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.920757Z digest=sha256:72184d5dbdc619d1a012324c099611b2914e080ecfa8ad3fe69b476190a87856

Observation deb05087-b1eb-4e89-ab60-1d6b2374109c · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LISA: Reasoning Segmentation via Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.925639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.925639Z digest=sha256:afbabe5fa71d176d8a397c05582a3e01a796300ccd883fc752cbbc35ddb0d519

Observation db32a862-f260-4ab9-b921-d6165fa57ce0 · outbound

This paper cites an unresolved cited work.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.930291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.930291Z digest=sha256:d4bfcbffbb7b38993957833a0accef3cd88c435daec0d2df3a4a3e017ac2b789

Observation c787d8c2-8e6b-480a-8cd7-3178a0f78d2c · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.934804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.934804Z digest=sha256:60ae7e6a1a1c4fd08cf567b8731e32dd4f6c8367ba08601878f6609095db1048

Observation 95f7775e-a62a-46b2-80ac-5fe3d9b007aa · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LLaVA-OneVision: Easy Visual Task Transfer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.944573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.944573Z digest=sha256:479718b11bf6c3af5408aaec1b6ce12771e7c1691dbdff8affd834aa44313acb

Observation 0306f9d0-0158-41f2-ba01-58846d135246 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.949183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.949183Z digest=sha256:abfc401c9c0df8b2817e187a64c1d5574599dd5db85fc8ca035fb2b12ed032d9

Observation 9e8c2cc3-be2d-4146-ab18-8a688712a39a · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoChat: Chat-Centric Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.953754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.953754Z digest=sha256:56011a9233bda7221216fbc79b45f233931d777058767c9f36ed5265d60a57cc

Observation 87110916-75ce-4f5b-a40e-85b5ea15f5c7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.957684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.957684Z digest=sha256:e6017f72b4de689f874405bfdd9170ac904b38366d9d226b9a1fe804130ba513

Observation 73670dad-9bbd-4051-b0bb-0b5e114ad508 · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.961719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.961719Z digest=sha256:ff2518397f3f3551158de350bf18723ce4400a74de80fadf357c77be51ab9175

Observation 20e60bc7-4ae4-4d0f-b46d-9410274097ba · outbound

This paper cites A light weight model for active speaker detection.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale A light weight model for active speaker detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.966886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.966886Z digest=sha256:22495d8a7aafac6918b20596ee62c40ddf1bd53305543ce5252a5c6081d4fca3

Observation 19ad1bae-9999-4bdc-8a78-abe754dffd57 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.970697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.970697Z digest=sha256:ec4d6134bcfdae06d8cd17615a33cf33c766fb262406f0ed51367df9e41ac004

Observation b345a33e-efb2-405e-b171-865a7a5412c5 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.974973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.974973Z digest=sha256:b0a9bc7cd40f4a92468ec01a7e743a2f75f1c9d289ff6c43efb82d7f0b12bbbb

Observation dd5d5368-5f9d-4747-8bf3-7eb8169b8a2e · outbound

This paper cites VILA: on pre-training for vi- sual language models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VILA: on pre-training for vi- sual language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.979524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.979524Z digest=sha256:c36aedfa4efa86f7df5e770b4e7fd094a773ebbd7c500f41b39a2fc176bb0ad4

Observation 53da6c4a-db92-437f-bbd4-4015c36bba8d · outbound

This paper cites Egocentric Video-Language Pretraining.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Egocentric Video-Language Pretraining

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.983262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.983262Z digest=sha256:ff0741067831a37a19f6ab1f2a0d49936b5d5e629da941662f335582e8d2a707

Observation da39fea7-c1e7-411a-8328-376f24a542ef · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Univtg: Towards unified video- language temporal grounding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.987297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.987297Z digest=sha256:904b1a7501b771bbcdd9f9fef4568c92590135d82c168c37a65985b16499d20a

Observation d2736919-88c5-4959-81bd-1faebf7becef · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Improved Baselines with Visual Instruction Tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.991162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.991162Z digest=sha256:50106e80a253c350ed5ef29795ed95ea04d2fc9605470819718c3d16e18b1db1

Observation 4bbcf587-60c7-4ea5-b55a-a7d3ee96cf44 · outbound

This paper cites Visual instruction tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Visual instruction tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.995196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.995196Z digest=sha256:0456d4efa0bb6845d13ebc46591899b031671d345e6aac1ccc6ce89de14c5a03

Observation 5a39f103-01bb-4617-8264-92f69a929d96 · outbound

This paper cites StreamChat: Chatting with Streaming Video.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamChat: Chatting with Streaming Video

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.998921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.998921Z digest=sha256:e089661562ae76fc73c62310fc20db6a3b855ea5ffb1aca6b5b667959d083593

Observation e984bf2f-6b36-497b-a987-99e4c1bbab05 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.002786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.002786Z digest=sha256:07c8603a4f9e2660ac47b50aef942a616df6eb926171e900a4815b2d66c2deb3

Observation a19029ba-75b2-4a8a-bac7-c146adf0e1f4 · outbound

This paper cites Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.007316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.007316Z digest=sha256:b663ca9b4bf9ad19a48868dd2c091e933882d292d1e3606d3994bcebf0378a5d

Observation e081aa2a-6b19-4c75-a395-b26121927e2f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.011424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.011424Z digest=sha256:b97a7137ed40f9be691d7934e4378b2d0b6828b009abc08c07510b29b17a97ca

Observation 704b455b-9793-4287-9fb4-01d6fda1d00e · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.015320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.015320Z digest=sha256:ab1f9ed6b8f8f20ffa95ff2b2cfbdd0c2af9621055910b60f6d8d8f3331f4b73

Observation 0ab3ac74-9a11-427f-b50a-af80fb355c76 · outbound

This paper cites Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.020223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.020223Z digest=sha256:109299c19b52aa741d4bcb95ce824ff0fbc37e060ba0d982a574e3cc15e62bbf

Observation fb594618-7bd0-4f8e-8580-f6e324dad4b3 · outbound

This paper cites Soccernet-caption: Dense video captioning for soccer broadcasts commentaries.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Soccernet-caption: Dense video captioning for soccer broadcasts commentaries

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.024877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.024877Z digest=sha256:0097182cdf64528f00094b50ae586a656353e1413be1e90adc4f1bdc8e642817

Observation d01ed2e3-3edd-48e3-a019-af951c800f96 · outbound

This paper cites Introducing chatgpt.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Introducing chatgpt

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.029317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.029317Z digest=sha256:497afa1fce92bbeeb8246cd4c3bb6b29bf0aacef26bc25c40241b640962510a0

Observation e4341a45-59f6-4dbb-bb11-777f7b1e4312 · outbound

This paper cites GPT-4 Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.034669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.034669Z digest=sha256:635d0bdedf275fb6bb1ec554421d8f17d5b27b214ed424e22928bf35cc5ce4cc

Observation 9d084c85-cd3f-4b17-a10c-f6facc010116 · outbound

This paper cites Gpt-4v(ision) system card.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gpt-4v(ision) system card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.039213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.039213Z digest=sha256:a9e60033e8b35a13b25761a4761a05289837b8d8e24b30645c2e7fb04c7b771d

Observation 33c8a06b-c802-4e9b-93ba-cb1eb361dffb · outbound

This paper cites Py- torch: An imperative style, high-performance deep learning library.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Py- torch: An imperative style, high-performance deep learning library

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.043228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.043228Z digest=sha256:6bc9e4ada6db81e1bc10db5d961b8319006aadef4927beec1bb6cf457d74e9e6

Observation e8a3bf00-bb5e-4d79-b235-3401fa466da5 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Perception test: A diagnostic benchmark for multimodal video models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.047488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.047488Z digest=sha256:ea001a8069330afbd91937d6323a731074cbd0cca14b835cfd794ee6d9d2e303

Observation c5e902d5-a7ab-415d-bdf6-4ba3878abb42 · outbound

This paper cites Streaming long video understanding with large language models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Streaming long video understanding with large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.052049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.052049Z digest=sha256:0a54c0e2e94d7064a2e7f00fd30787b03774709e0e16f1ffc68178fc2889e7c1

Observation ca7c4f55-dbfc-4924-a060-cae46446bc3d · outbound

This paper cites Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.056202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.056202Z digest=sha256:7c1fcc966daab3bd0b6ba8fbf6c8cc89a73af30d9a82c9edbfa93155b1b2036e

Observation aa23ca37-cace-4089-b2af-98460344288a · outbound

This paper cites Improving language understanding by gen- erative pre-training.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Improving language understanding by gen- erative pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.060284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.060284Z digest=sha256:46994286f6eee91c1c80c0e130fd6122a81859c73d15fbab6f903ae15f0a8f1b

Observation ef305958-3548-40f9-a71c-2293c9743234 · outbound

This paper cites Language models are unsu- pervised multitask learners.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Language models are unsu- pervised multitask learners

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.067380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.067380Z digest=sha256:5709f9f8a2ad70f15a3c79408afb531f242259b9d97e93871c0897f6a96ecaff

Observation 4b4012d9-27a3-45ec-8565-100beaa417eb · outbound

This paper cites Learning transferable visual models from natural language supervision.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Learning transferable visual models from natural language supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.071808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.071808Z digest=sha256:bfb2a8b500bf8272d53bae8023781da0220dac1193290e8924a83eddc466a2d6

Observation 28fbe0e7-2d3f-49b1-893d-06b3fa2deaf2 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Robust speech recognition via large-scale weak supervision

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.693273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.075969Z digest=sha256:c43a03d4ea04ce5e60d78555a4d6782110314a1a0402815679dec520ab6bdc25

Observation 056449bd-b9ad-4b29-881a-b2fc58207442 · outbound

This paper cites MatchTime: Towards Automatic Soccer Game Commentary Generation.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MatchTime: Towards Automatic Soccer Game Commentary Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.080241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.080241Z digest=sha256:59cdd61da63e4bcce26ceda71fa3af40acab4b49a9eec2502d44596cda98879a

Observation 379b02c1-19be-44e4-acd5-8c204cd94060 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.678959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.085140Z digest=sha256:b7a1285d47a0c49a2283a8a0f0e71bd11c8e7780dc2a72b3cdc055ff9228df71

Observation d5e1b67b-44c2-4022-a2ae-042f7f86450f · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.089055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.089055Z digest=sha256:83857efd40d808ec2a44874052c6d16c2b92119dae8f407adb71543a1ab46da6

Observation f55b5798-3b4b-4500-9a7d-8198f6e820ac · outbound

This paper cites Online real-time multiple spa- tiotemporal action localisation and prediction.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online real-time multiple spa- tiotemporal action localisation and prediction

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.659905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.092797Z digest=sha256:3aafc8e26a0665b9f36fb0f8fab0c0f647d23a5dbff9dd9794051719f9f658b6

Observation daeffac6-c7c7-4239-8447-2dca8966603c · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.097192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.097192Z digest=sha256:e20c751e14654a600d1ffa69f47960d2b7e3ddd0f3953fdf864ed14f62869590

Observation 8689b0fc-8e52-4dcb-9195-5afd5fff4288 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LLaMA: Open and Efficient Foundation Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.101689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.101689Z digest=sha256:268abaaf92c102fcacfc889ef5d83dfa9606c6e91577e91d5687107a10f77604

Observation 61a5b472-9a3a-481e-8ed3-550f9bbd6367 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.106533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.106533Z digest=sha256:21f20ba1b009461ee0549dca223e92599f134c756ed1fe1c74c8e1af63554634

Observation f63131da-0943-4629-b683-176fe23c49e8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.111394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.111394Z digest=sha256:a6219fed4856314e54fc8823ef0718ca3d167bd960b94bc8a0bae1ece8517afc

Observation e83e75d0-0d1c-48b8-a29a-e92780fc3161 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.115890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.115890Z digest=sha256:0dc9cec54d549790dae63caeed9c55bfafd3208f76db3d5cf64d70b99d962077

Observation d7d822a3-1677-41ea-a25e-dce7d9195ba0 · outbound

This paper cites Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction for- mat, 2024.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction for- mat, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.644173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.120862Z digest=sha256:f3769f722e753875abd9373d24a23d8514d2a6e17f0ef650f201a20e85edbbb0

Observation eb4a92e9-fd24-4b10-af18-29c4ca55e21c · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.125134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.125134Z digest=sha256:7c9da4ce72e1e751619764e800f4a98ec98cbc1eb14a3c1bb441d06d2f938314

Observation 0a23ce1f-3f8c-438d-b923-afd5063047f8 · outbound

This paper cites an unresolved cited work.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:28.629649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.129671Z digest=sha256:1592bc0966882a8a96c8c7dae75002baf1efd74d87f924d92d62aa99a57ad05a

Observation c1d9e19d-5e9a-4611-bc4c-8983f837b543 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Next-qa: Next phase of question-answering to explaining temporal actions

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.615771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.133486Z digest=sha256:fdf36ac001c7c7e44287dd68876b9bf1afb800d1b60ad25c43e4fa6ce3fdba9e

Observation 459444c9-7720-4835-9eb8-e9c9effc9025 · outbound

This paper cites Streaming video understanding and multi-round interaction with memory- enhanced knowledge.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Streaming video understanding and multi-round interaction with memory- enhanced knowledge

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.601678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.138131Z digest=sha256:e2cd99c91a4ac82125af36f2c8f7f2b307a28a515282db444e2e215137ccdb3d

Observation 0fe113c2-9cf0-4d41-be15-037a91c09db0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2.5-Omni Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.142050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.142050Z digest=sha256:a4902caf4a46273aded89891e5c9428975e5cc6fccb17ec1ed54492b2aea8cce

Observation 07193edd-8de9-4dfa-a029-9731f769c61e · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.586709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.146546Z digest=sha256:395766b7a64b34297496acdb5d8ac75df661a1948d5804d59706952460e4e096

Observation c0803101-3fc9-4742-970c-411378aad026 · outbound

This paper cites Vidchapters-7m: Video chapters at scale.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Vidchapters-7m: Video chapters at scale

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.570350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.150371Z digest=sha256:6e59d082195b0b7a36ffebbdaf2f76ab1176a46c9ff4b2f6bc9d086639eb5b7c

Observation 6f5dd80e-7e9d-4929-9b80-b19945ea7bc1 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.555068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.154457Z digest=sha256:0a4d437d9d57a9813ebc175fa69b7a9080ffad853346d2d460503aaa9d0fc0eb

Observation 65debb99-407b-4073-adb0-cb6930d30c85 · outbound

This paper cites Qwen2 Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2 Technical Report

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.158788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.158788Z digest=sha256:db5d545e4a836ea6e456d04791f2ec6164e269ce6a03652916718d71f891a1fd

Observation 4bc2f568-8612-4aa5-8c20-724dcd182927 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.163742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.163742Z digest=sha256:6bf3d91f696261b2bcb42f2e8cd4022770c8d84dde1303b5778546b0699ab24f

Observation 3f81d77e-263a-45a8-8b99-488b5add7216 · outbound

This paper cites DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:16:27.554006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.168143Z digest=sha256:8a3dadfdb7f4c66c9aa3d8a130d38ce345224d49002436c9d43891222bb6730f

Observation 8b02ac2e-064a-4ef4-85b6-b71b2296a5c9 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.172576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.172576Z digest=sha256:0a421a2b1913ebf60badb4087f3db2b276911ba9ef4e8aa87087f8f3a0751c07

Observation fcc8537d-1432-4ac2-81a4-50ab1f6b6782 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.176811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.176811Z digest=sha256:3cf2708818348a0dae67c066111fd7811db60d27e31464715d8e4df3c6027bc5

Observation 699fbecb-687c-4872-acaf-50204469c765 · outbound

This paper cites MERLOT RESERVE: neural script knowledge through vision and language and sound.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MERLOT RESERVE: neural script knowledge through vision and language and sound

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.540022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.181405Z digest=sha256:7fe811bdf66743f1441d5237a006f80fd1b4e87eafeba0986a63a96493beaa56

Observation 6305bd7d-de47-49d2-af13-1964a3e9a72c · outbound

This paper cites Sigmoid loss for language image pre-training.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Sigmoid loss for language image pre-training

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.525984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:16:27.185356Z digest=sha256:c22b2a0eaccd383ac5bb3b8b830f8cfbb9dcf0b59c62d8073bbd7faea9391bdf

Observation 514febb0-d8a4-4197-9b0f-93e6126e6316 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.189191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.189191Z digest=sha256:288d1ba2d27aaa88ba8421e77a050100952e682c9b3640b8512c60ab10f2b2af

Observation 8283c1e8-3d87-4f24-b9c2-f5773644ae1b · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.193114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.193114Z digest=sha256:b9cc4a6f7c9006ff72d1d2a71b4ca689db69d7a12c79434f8298ced6598d03f6

Observation bdfee9e1-d589-47d1-82de-7faa174aa7ae · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.196961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.196961Z digest=sha256:61f2c3fb999edcc8f1b3dd586cc941b36a7107724b690eccaffee7fcb65398c2

Observation babcce6b-24f5-4b23-9bfc-00867f7882e0 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.201212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.201212Z digest=sha256:61db5734c7abc3807e57b27c75e0a228c5c0143d3b372bfa8c393da368fe495c

Pith citing papers

Observation 79481876-9f37-4128-8e6c-513d58e8d9dd · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.405635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.405635Z digest=sha256:1151a21e1901984f2915867fd20548d361eb248326a46fe40ba933ba0d1c6d34

Observation 9b7594dc-5632-4def-816d-b98de7c8a0bd · inbound

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? cites this paper.

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:39:06.101648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T05:36:09.208754Z digest=sha256:69e7b1d4b0d99cca0514a9dff0719935f47f6833dec4645f357f7382025a4b4c

Observation cecdb1c9-4d81-4dd4-bb48-b83e4bef12d2 · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:42:06.765151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:fe21965190306dc38d719daa0db2fa8a15ad028bb0dcb8eaefb30898a31cb934

Observation dbba1ab8-add9-4c98-a3ef-7ee2c9dc6af4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.144883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:8d5d1bba198039c9e67004682e7c92756cba979698ba1951a788466157433122

Observation 84dc0ba4-7c41-432b-8e86-6922425289cd · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.334651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:0a3a9636a84464f1ff439f41cfe45d389abca3d4283794920f4813a77c461fb2

Observation 69338436-7dfc-4746-9633-f9fba57f9c20 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.482051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:51b566eb5e6a3e164b1f473873a87cbe22ad97d5cec9214a7920846fdde417b1

Observation aabb4aa2-8816-4064-8777-56c364f342e3 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.899753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.899753Z digest=sha256:ab79a8aa426aca0006c65dd7370c4eaf55b1f4610898aeb1c7043d5849699b5f

Observation d29b28de-f054-48a9-8b66-5a1aa809c1e3 · inbound

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation cites this paper.

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:27.447842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:39:27.447842Z digest=sha256:8ddb5cd68e0dc36476ef4e40f9b3046bdaa3303c9698240977265851039a4048