Pith. sign in

Paper Citation Record · LEDGER

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

As of 19 August 2026, this Paper Citation Record lists 100 of 116 outbound references and 8 inbound Pith citation observations for arXiv:2504.16030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16030 v1

Coverage vector

measured 100 of 116 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:27.201212Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.405635Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:19:20.332342Z

Reference resolution

100 of 116 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved88
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd0846cc-809a-4cf8-86c4-31118228f770 · outbound

This paper cites Marks, Chiori Hori, Peter Anderson, Stefan Lee, and Devi Parikh.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Marks, Chiori Hori, Peter Anderson, Stefan Lee, and Devi Parikh

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.761886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.761886Z digest=sha256:3fa4cfe641b3061f72649d9cfde03a3a47bac8470c9449d2fe0ce9de9d1b780e

Observation 257cfb55-f212-4089-8193-d20bcafc34f9 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.767470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.767470Z digest=sha256:e1d4010f3971babf6a67a40e7a9e8bd2314d49f5d2e970d9a44061bad2beaed1

Observation d60b12ad-da72-44e3-bca3-0a1283a0944c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.771846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.771846Z digest=sha256:298d250bf6c1019cb1e289da274b9981e18b35c452b87d042756780b23e0d7f6

Observation 8503c045-d28f-44ce-b3cd-e1ffcd0e39a7 · outbound

This paper cites Qwen Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.776536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.776536Z digest=sha256:8c12f2fe32522e7f8ff9c72360dbdca97795be7303b7e19a3975d0474eea9ee1

Observation b1b66b2b-b51d-4da7-bb04-aecaae912591 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.780847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.780847Z digest=sha256:c1101b6b1162d1b6c51609d697b333e100424462549fa30726e0c96c8b2a9022

Observation a0e4853b-eebf-4d64-95ec-29d260a6bcd6 · outbound

This paper cites Qwen2.5 Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.785092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.785092Z digest=sha256:7dcfed83ce1ee433613e44ce0fac97c67c09decc67ef6398cece3069bd704fd4

Observation baa3eead-eedc-4049-887c-697ccd1d3d82 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.789286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.789286Z digest=sha256:be22f0143980f4c9c4b665fd8cbd370258ce8f11c85f2f915ecdda5752693dbc

Observation 02f23cd0-dbda-4da0-bb09-6a7f53641bbf · outbound

This paper cites Language models are few-shot learners.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Language models are few-shot learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.794090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.794090Z digest=sha256:cecaf004fa6a8a4396d2b62d157860d849a073276b76521eb3664e644c5d1821

Observation 6a700740-b763-49ba-8d96-5002fc650141 · outbound

This paper cites Sst: Single-stream tem- poral action proposals.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Sst: Single-stream tem- poral action proposals

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.797881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.797881Z digest=sha256:a62f7e4315d0c640d959e428a1aa44d94ab4b9907cedb0fdc77071684f8656be

Observation 637d85f5-6c9b-4253-8e3b-9a2c22076285 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Quo vadis, action recognition? A new model and the kinetics dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.801885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.801885Z digest=sha256:0d147aad845828b36004dae1f15c7b8d2979b28af4cbf7f4bfd76aae6ef38b65

Observation 39a73a2a-f219-42cc-af70-7f311920b156 · outbound

This paper cites A short note about kinetics-.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale A short note about kinetics-

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.805640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.805640Z digest=sha256:5dead8556dd163225e63be5aea0d8c8609f7414a1502795da3cf39d924238f0e

Observation b3906aa8-2eb9-49e1-bea1-c5fa40ad6293 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.814227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.814227Z digest=sha256:a293bedd72c723c680dd8656b665f7327a96bf51a69c4c75a3f90b600c622fee

Observation 17c4cc69-e049-4d5c-b875-416653183106 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Videollm-online: Online video large language model for streaming video

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.818806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.818806Z digest=sha256:0579cb1b9150dec30b0bb204982af70d4c15a2ab4f4241b88494267184c9f940

Observation e9cba1c9-195a-4afb-93b7-0eb6c5e9d4c7 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.823292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.823292Z digest=sha256:c716c06f04215f970c1821a6b909434a071a217a884e65c8ee7b0e690d6f679e

Observation 77644c62-c134-4fe2-9834-7ab34e3a176e · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.828107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.828107Z digest=sha256:0bf039e25e359fb61e92bf98ccadfc6b72e0e71c0119229e4a674b5e7bd6fa57

Observation 0b99d661-5a1b-4200-8165-41a5236dbc2c · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.832033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.832033Z digest=sha256:6e73f4fc0a0e07c1049afce273c4ef126d59597bab1473151a2724be6c797729

Observation 850fc669-e29b-47e6-8e96-0bc22d9a53d1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.836191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.836191Z digest=sha256:e7702654d3b6a5fcae7ac47734769a8612a7259f550f51ea21792a2510af8f12

Observation 28bba35b-989d-44c4-8fc6-7864199704c3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.840455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.840455Z digest=sha256:be543ea11cd45dcde0ec73de3cc89c56e7f0659d14f2b03bcb9b0d8d75fc78af

Observation 6e651ea4-e587-453e-a528-de11543ad32b · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unsupervised cross-lingual representation learning at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.844728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.844728Z digest=sha256:423f18b6958d038c7d26e805163493c3219df0954766621418919ae2e6fc0095

Observation c70153ca-b921-4f37-9612-fafd25b58018 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.849493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.849493Z digest=sha256:01adade9d8a1c84f4738583bfaaf8ea6258f9ff469597dbfa83105d4a8fdfd7f

Observation d8233268-17be-4faa-9c2d-eec836006a46 · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale An im- age is worth 16x16 words: Transformers for image recog- nition at scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.853854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.853854Z digest=sha256:04c61e8b49140a68ee7798e922f7e70bc0b57bdf148ed47c1a263be336d47dce

Observation cfad65e2-bc85-4f24-a418-c4c601ef3d18 · outbound

This paper cites The Llama 3 Herd of Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.858399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.858399Z digest=sha256:695588dde8bb3129b2df5fbc45427a8d01057a9f8156b6f5a44f80574aba439b

Observation 8476e595-0b83-4c5e-8269-2ad8f2e8c03d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.862576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.862576Z digest=sha256:c15a2d341b88d480daaa3c42bb0775d87d470e02267ba40d500de343ea9b59ca

Observation 8dc06081-1aae-4958-b77a-dcdb6a00ab8e · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.866575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.866575Z digest=sha256:9b28eec10165fc4c137d9f3220f39f6739ed701cc285a6d78d01f7f58dd8a726

Observation 9e2eda36-8332-4475-9a88-c8091db910b1 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.870990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.870990Z digest=sha256:bcd7f5760ede1580807f1cbbfe3fdfc04a61bdf5e5b1842a246f4e2054a810ff

Observation 4d46e76a-8a11-4884-a8e2-d70247e88d0e · outbound

This paper cites Online action detection.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online action detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.875352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.875352Z digest=sha256:40047dd5816d1689bc28f86db30f5ab358bb0c45b6a29d1d818d3a8d8ccb8946

Observation f0bd417b-6f8d-4d82-9668-a01a5d2d271d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.880005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.880005Z digest=sha256:f797e0c0b13d9c81add6966ae2e1685fe397bed857fe837f2a9d0ec5b57fa64d

Observation 345faeed-0bb8-4911-a8ff-dbf4bde53adc · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale The ”something something” video database for learning and evaluating visual common sense

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.884157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.884157Z digest=sha256:0bedae0b556fc8c63f21e4f4396fb95e2bd4337f024d3c7c9df27444d11407b5

Observation 3a1aec64-51eb-4c9b-994c-805a2a69db50 · outbound

This paper cites Hello gpt-4o, 2024.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Hello gpt-4o, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.888082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.888082Z digest=sha256:ad2bfbb707375c642602add6abe8cb4236124f887af9d434431360bd5a2e6e0c

Observation 9f364702-6318-4367-9f49-f75d0c5cc226 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Activitynet: A large-scale video benchmark for human activity understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.891983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.891983Z digest=sha256:32b5d0a1b4d72a09be5cfa1a0f73ac0924432c3bde1a2857d81decc5f198d3e6

Observation 1bbc3a83-cd54-4a0b-b16c-5a1c2b391f4c · outbound

This paper cites Training Compute-Optimal Large Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Training Compute-Optimal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.895809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.895809Z digest=sha256:17022dbf7d1aaacdc5c99e6de6b8d864e67f2dcbd3b6c1b3b18169d69bcbb3ed

Observation 8a0aa7b7-c973-4720-877a-f412a50cbdda · outbound

This paper cites Multimodal pretraining for dense video cap- tioning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Multimodal pretraining for dense video cap- tioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.899907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.899907Z digest=sha256:a1271ceba231e6763c7b3d05e5b3d8db68e0b83d5b2db4d978bb99567a2407a6

Observation 428d89ef-547a-4ba3-825f-4f4f83c794ea · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online Video Understanding: OVBench and VideoChat-Online

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.904234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.904234Z digest=sha256:0e1b7a52fb914a9bf2195cc2815eece9787d35985542717523d9330bd9388fca

Observation f4f299d2-6ecc-4193-a4bc-91bf59650387 · outbound

This paper cites Zamir, Yu-Gang Jiang, Alex Gor- ban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Zamir, Yu-Gang Jiang, Alex Gor- ban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.908553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.908553Z digest=sha256:0462aad464c380b54e4718cf4ec3e722e399c9453bfbe75dec526bebc9b33a02

Observation f4cdbe42-2727-42c3-aeb5-d2d4787d36cb · outbound

This paper cites Cag-qil: Context-aware actionness grouping via q imitation learning for online temporal action localization.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Cag-qil: Context-aware actionness grouping via q imitation learning for online temporal action localization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.912817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.912817Z digest=sha256:c299960d99723016ed934a34e7c0ecc2d4aca40028b7fbb3f87007478caf6121

Observation 641431e4-bdf7-480c-8305-e6399750a792 · outbound

This paper cites Scaling Laws for Neural Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Scaling Laws for Neural Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.916707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.916707Z digest=sha256:1d7d8c640cd1dc14cc78885191c830a2e47dda41ae33f017a76a2933a0e3a9b0

Observation 6c304600-fece-4dc5-8a03-72e345afe0c0 · outbound

This paper cites Dense-captioning events in videos.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Dense-captioning events in videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.920757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.920757Z digest=sha256:2d5124253ebde9ab9f2c5b8ba701bd3e4a319e4e500b81ae62edbe75c70ad453

Observation deb05087-b1eb-4e89-ab60-1d6b2374109c · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LISA: Reasoning Segmentation via Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.925639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.925639Z digest=sha256:b4c2294395199dfb52da44677afeee5752dbd0bce19b38f240e2ed6c9e7e9dea

Observation db32a862-f260-4ab9-b921-d6165fa57ce0 · outbound

This paper cites an unresolved cited work.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.930291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.930291Z digest=sha256:b065edeabf2f2d2712f00a03a74b3721e450c9af0383168362d1cb45f527d111

Observation c787d8c2-8e6b-480a-8cd7-3178a0f78d2c · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.934804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.934804Z digest=sha256:bed1d0d789c820c035cd5af8b1131f5417c238865e3d9b62fed2b7ff27ca91db

Observation 95f7775e-a62a-46b2-80ac-5fe3d9b007aa · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LLaVA-OneVision: Easy Visual Task Transfer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.944573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.944573Z digest=sha256:0c04d613f5cdbc7b301939ad1e895ab958ed8ee27ee188dd3030592e7b863bd4

Observation 0306f9d0-0158-41f2-ba01-58846d135246 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.949183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.949183Z digest=sha256:2d37bdb76da462220e3d0f89c159ae82acce34333705905ccbff49cb0de2e20b

Observation 9e8c2cc3-be2d-4146-ab18-8a688712a39a · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoChat: Chat-Centric Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.953754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.953754Z digest=sha256:f423445a83509141b053191525f48dc9d70b4e6b59c1fba4da7cec30f62b039d

Observation 87110916-75ce-4f5b-a40e-85b5ea15f5c7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.957684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.957684Z digest=sha256:45e1fa991313927375cea34be4187ecfbfef691e05b64581aba4ea778f523901

Observation 73670dad-9bbd-4051-b0bb-0b5e114ad508 · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.961719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.961719Z digest=sha256:4a865d815bf2a44aa630f22a51e639daa000b9256e7011bf87b18a018c2c622a

Observation 20e60bc7-4ae4-4d0f-b46d-9410274097ba · outbound

This paper cites A light weight model for active speaker detection.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale A light weight model for active speaker detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.966886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.966886Z digest=sha256:0366c795389ae440e31a995a61c20a701394f0c584c9baa2e98ea458112fbadd

Observation 19ad1bae-9999-4bdc-8a78-abe754dffd57 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.970697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.970697Z digest=sha256:3f92e7c90f28108c42c9bd7a7ec2d67324d57f94a2584804b0ef452ce2d6faae

Observation b345a33e-efb2-405e-b171-865a7a5412c5 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.974973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.974973Z digest=sha256:adb249b10e78a1075d954f9d21c4facd33bebecb1b79bdd555755bbc2f9967ab

Observation dd5d5368-5f9d-4747-8bf3-7eb8169b8a2e · outbound

This paper cites VILA: on pre-training for vi- sual language models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VILA: on pre-training for vi- sual language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.979524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.979524Z digest=sha256:2ce20e275eecc05d877ca171605a3989c30131d5ff5ae64188b7a5daa5335b89

Observation 53da6c4a-db92-437f-bbd4-4015c36bba8d · outbound

This paper cites Egocentric Video-Language Pretraining.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Egocentric Video-Language Pretraining

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.983262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.983262Z digest=sha256:5c27501e3d6621afc00967c690b2e24e3cd346e32dcf17ad56c57fbfdb6b2f7b

Observation da39fea7-c1e7-411a-8328-376f24a542ef · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Univtg: Towards unified video- language temporal grounding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.987297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.987297Z digest=sha256:941e920c484c8eaca187c2ebfe88b36a772f2259262aa95069eacc3c63356fc4

Observation d2736919-88c5-4959-81bd-1faebf7becef · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Improved Baselines with Visual Instruction Tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.991162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.991162Z digest=sha256:7bf3a1944e16549d8c565019d62d08ba035600dae0830ec3bfb29c4e0cb348bb

Observation 4bbcf587-60c7-4ea5-b55a-a7d3ee96cf44 · outbound

This paper cites Visual instruction tuning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Visual instruction tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.995196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.995196Z digest=sha256:f47825802e968eee9ca51f8a8126c179681083eb25493b541dfc0d83db662882

Observation 5a39f103-01bb-4617-8264-92f69a929d96 · outbound

This paper cites StreamChat: Chatting with Streaming Video.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale StreamChat: Chatting with Streaming Video

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:26.998921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:26.998921Z digest=sha256:23f692ea321230ee2008a77850d24bf242005b1168da9619ffdb3182419cb3ce

Observation e984bf2f-6b36-497b-a987-99e4c1bbab05 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.002786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.002786Z digest=sha256:41eb720cbbcb8206cc2a88fe36b4147551419aef5aaefdb203bbdc78a4220f6d

Observation a19029ba-75b2-4a8a-bac7-c146adf0e1f4 · outbound

This paper cites Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.007316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.007316Z digest=sha256:0deeaeee89d88b4b59cd01f6defac090df25eb7cee41427451ad222e79a1ca5c

Observation e081aa2a-6b19-4c75-a395-b26121927e2f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.011424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.011424Z digest=sha256:ff0de6720c8b254fa472eea1223d4102e89d6ee0b1d22e30a7c80ab5639a868c

Observation 704b455b-9793-4287-9fb4-01d6fda1d00e · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.015320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.015320Z digest=sha256:c26e9f46c68d4bd8ae6b91dc7278a54ec219b5c2090e93d9d97518cf96927524

Observation 0ab3ac74-9a11-427f-b50a-af80fb355c76 · outbound

This paper cites Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.020223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.020223Z digest=sha256:23d749115d983fd920be1c7bad33fb9c9453d3cdb49498fb6d239315e3dc727a

Observation fb594618-7bd0-4f8e-8580-f6e324dad4b3 · outbound

This paper cites Soccernet-caption: Dense video captioning for soccer broadcasts commentaries.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Soccernet-caption: Dense video captioning for soccer broadcasts commentaries

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.024877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.024877Z digest=sha256:bbdb4cb67630d494258c33309bb9a1a5dafec20d268b52565d026b32e5584402

Observation d01ed2e3-3edd-48e3-a019-af951c800f96 · outbound

This paper cites Introducing chatgpt.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Introducing chatgpt

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.029317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.029317Z digest=sha256:c8ec6f701affb55009cef462824d8e5a29ac3d53c8633bf2a34a8d013893c9a3

Observation e4341a45-59f6-4dbb-bb11-777f7b1e4312 · outbound

This paper cites GPT-4 Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.034669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.034669Z digest=sha256:5c7d0fb6a7dac3078c7267a56c9ee3aff8ba4a7cf405c274faff40a99cde44af

Observation 9d084c85-cd3f-4b17-a10c-f6facc010116 · outbound

This paper cites Gpt-4v(ision) system card.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Gpt-4v(ision) system card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.039213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.039213Z digest=sha256:5997094fef382e55dca16df70822c1abbba626f23de9b9932eccf5576d855fa9

Observation 33c8a06b-c802-4e9b-93ba-cb1eb361dffb · outbound

This paper cites Py- torch: An imperative style, high-performance deep learning library.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Py- torch: An imperative style, high-performance deep learning library

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.043228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.043228Z digest=sha256:3dc0bb9a4a1c5834698b0c3e515080ad93ca45dc4e176c001edf6818aa67a4f0

Observation e8a3bf00-bb5e-4d79-b235-3401fa466da5 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Perception test: A diagnostic benchmark for multimodal video models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.047488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.047488Z digest=sha256:d3dd5515ee50f7486220ae81c602b67a4b90d0d7e46ec9bcf3dafd723de7ceb4

Observation c5e902d5-a7ab-415d-bdf6-4ba3878abb42 · outbound

This paper cites Streaming long video understanding with large language models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Streaming long video understanding with large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.052049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.052049Z digest=sha256:b3d8d83791c0aedceb0bead100b7056fedf6d4f660232a6e7952934b5fa164b2

Observation ca7c4f55-dbfc-4924-a060-cae46446bc3d · outbound

This paper cites Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.056202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.056202Z digest=sha256:61a91b7b059919c6a086a452f9c254b795625153d8bdb45f875de76d857086a3

Observation aa23ca37-cace-4089-b2af-98460344288a · outbound

This paper cites Improving language understanding by gen- erative pre-training.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Improving language understanding by gen- erative pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.060284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.060284Z digest=sha256:f8260e9e9992d9fb4270b445e4ea9a670aa799b5454216d3d62e9d44188700b0

Observation ef305958-3548-40f9-a71c-2293c9743234 · outbound

This paper cites Language models are unsu- pervised multitask learners.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Language models are unsu- pervised multitask learners

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.067380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.067380Z digest=sha256:ef422d12a1a86f9058a7082c92145895d60bc356838b695b7e602dfb45f74d4c

Observation 4b4012d9-27a3-45ec-8565-100beaa417eb · outbound

This paper cites Learning transferable visual models from natural language supervision.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Learning transferable visual models from natural language supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.071808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.071808Z digest=sha256:18e724cd903fc77b282b7eb43c525c99890ab15502608ad80fc4ff3d63c4b8e3

Observation 28fbe0e7-2d3f-49b1-893d-06b3fa2deaf2 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Robust speech recognition via large-scale weak supervision

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.693273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.075969Z digest=sha256:700d910a320ee4de65cbd102cfc7ce6a823d133738f313b0f603c60c556b906c

Observation 056449bd-b9ad-4b29-881a-b2fc58207442 · outbound

This paper cites MatchTime: Towards Automatic Soccer Game Commentary Generation.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MatchTime: Towards Automatic Soccer Game Commentary Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.080241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.080241Z digest=sha256:8a427fd576186f987762b71ba30f0d7ee4738f09348f325e4fcfb136639420c2

Observation 379b02c1-19be-44e4-acd5-8c204cd94060 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.678959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.085140Z digest=sha256:df5c282b1b109b58979d2c934f2b215e340267a156443804908e1fef514fff9e

Observation d5e1b67b-44c2-4022-a2ae-042f7f86450f · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.089055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.089055Z digest=sha256:242d4487ab365831eeb0aea64ccfa9c13c02a1895331c6989d1c248d5309f953

Observation f55b5798-3b4b-4500-9a7d-8198f6e820ac · outbound

This paper cites Online real-time multiple spa- tiotemporal action localisation and prediction.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Online real-time multiple spa- tiotemporal action localisation and prediction

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.659905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.092797Z digest=sha256:68fc4c5d2dae172b1fd1f16a863e939868d30646c4074193ba6d37268db66067

Observation daeffac6-c7c7-4239-8447-2dca8966603c · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.097192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.097192Z digest=sha256:c913f5e45e3cf768173474541a2779e6f9a1df21ef33079f6accdec12000f943

Observation 8689b0fc-8e52-4dcb-9195-5afd5fff4288 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LLaMA: Open and Efficient Foundation Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.101689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.101689Z digest=sha256:8582e67ae01c3b0b1261f8709f36e3827eabf966bd585a6d52daabf127633e6f

Observation 61a5b472-9a3a-481e-8ed3-550f9bbd6367 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.106533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.106533Z digest=sha256:b020ff24b813058bd1cf02b635eef59495eead81ad3a56701eb1ec5e5107dc03

Observation f63131da-0943-4629-b683-176fe23c49e8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.111394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.111394Z digest=sha256:d18153e898807e0290e929f477bf3d9c014c2f21e42422144bbd39173dd5c976

Observation e83e75d0-0d1c-48b8-a29a-e92780fc3161 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.115890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.115890Z digest=sha256:fdc17e9db12a48e6cfc845fd76fe8aa7fc5f9be8b50b309e4e56ea3c908789e0

Observation d7d822a3-1677-41ea-a25e-dce7d9195ba0 · outbound

This paper cites Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction for- mat, 2024.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction for- mat, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.644173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.120862Z digest=sha256:c1d5310f90be5854e8c26d7583f59c0f9f1250e9ad1c10379d1559aea5c37f11

Observation eb4a92e9-fd24-4b10-af18-29c4ca55e21c · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.125134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.125134Z digest=sha256:3ded89acba6d74927837bd4afd12750652c6039b4f6222dd307fe9f093ae553f

Observation 0a23ce1f-3f8c-438d-b923-afd5063047f8 · outbound

This paper cites an unresolved cited work.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:28.629649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.129671Z digest=sha256:786ad9fe80426e8a27e4ec3ff30695b756e3a6a525ffb64e52e3d14c31bad1a5

Observation c1d9e19d-5e9a-4611-bc4c-8983f837b543 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Next-qa: Next phase of question-answering to explaining temporal actions

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.615771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.133486Z digest=sha256:9c5469928783742bc2f4930e17699864233e28e08c101478bad6746166040d42

Observation 459444c9-7720-4835-9eb8-e9c9effc9025 · outbound

This paper cites Streaming video understanding and multi-round interaction with memory- enhanced knowledge.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Streaming video understanding and multi-round interaction with memory- enhanced knowledge

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.601678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.138131Z digest=sha256:fd45e6b05186e255dc5c1c2996860359ffbeaad622e99284ee0e7a5f6d1fd084

Observation 0fe113c2-9cf0-4d41-be15-037a91c09db0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2.5-Omni Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.142050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.142050Z digest=sha256:49ce030c4d07788c8bbea222654b8a2f4ab9e79cbbf5288ea2b90e7f6250c1a2

Observation 07193edd-8de9-4dfa-a029-9731f769c61e · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.586709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.146546Z digest=sha256:c1b493f9608a22d61df70a01dad36e184e1567540eb19f5325db882ab9237297

Observation c0803101-3fc9-4742-970c-411378aad026 · outbound

This paper cites Vidchapters-7m: Video chapters at scale.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Vidchapters-7m: Video chapters at scale

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.570350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.150371Z digest=sha256:a3c653d0fdb29bd52f5dd96b9a7809a3b24e603de4db618226dfe8788d7839d1

Observation 6f5dd80e-7e9d-4929-9b80-b19945ea7bc1 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.555068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.154457Z digest=sha256:7037ea42642dc697cb0a75177c7632b5e1b13d034776a34174c1a8c498969b0e

Observation 65debb99-407b-4073-adb0-cb6930d30c85 · outbound

This paper cites Qwen2 Technical Report.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Qwen2 Technical Report

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.158788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.158788Z digest=sha256:ed7a17cf52d5a850f7a65bb43af33a7bcdfd4809c985c687f838f4666dda3dea

Observation 4bc2f568-8612-4aa5-8c20-724dcd182927 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.163742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.163742Z digest=sha256:0beb6b4ccb10b235ba661151ec89f1e8d40a028a3faf36cf7c2ac26ea1ca5735

Observation 3f81d77e-263a-45a8-8b99-488b5add7216 · outbound

This paper cites DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:16:27.554006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.168143Z digest=sha256:f30d2b07d50117fffa160a51d758aa0a42c908d195b8c31bccfafce76ffdbe84

Observation 8b02ac2e-064a-4ef4-85b6-b71b2296a5c9 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.172576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.172576Z digest=sha256:8471bee78f0f0c0466cfa86e7fd436ea44cb5ecaa8a8df9422cd8010593a31e4

Observation fcc8537d-1432-4ac2-81a4-50ab1f6b6782 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.176811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.176811Z digest=sha256:f558f956309240218e6c6107809d3fcf7f6424d19a849c11078a65b19d7cd580

Observation 699fbecb-687c-4872-acaf-50204469c765 · outbound

This paper cites MERLOT RESERVE: neural script knowledge through vision and language and sound.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale MERLOT RESERVE: neural script knowledge through vision and language and sound

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.540022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.181405Z digest=sha256:093469d38b7c050c45bd25c0eab4d28ffe38535475d555165750fefe286037c1

Observation 6305bd7d-de47-49d2-af13-1964a3e9a72c · outbound

This paper cites Sigmoid loss for language image pre-training.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Sigmoid loss for language image pre-training

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:28.525984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:16:27.185356Z digest=sha256:36a78e55dc49445a86c0b12b1505ebb6ead26d9e9bbd49abf0ad2d43f5087a97

Observation 514febb0-d8a4-4197-9b0f-93e6126e6316 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.189191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.189191Z digest=sha256:93fc27e27b9a3b3d6825012b7f34dac52ba741172d62015a458ba1bfe39754f3

Observation 8283c1e8-3d87-4f24-b9c2-f5773644ae1b · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.193114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.193114Z digest=sha256:2701c646680e6ebc23d2ad952d736f5eed40057ff07378dd11b5ee7b5f26a61a

Observation bdfee9e1-d589-47d1-82de-7faa174aa7ae · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.196961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.196961Z digest=sha256:1529ff37afef5e4a3f7de70f1f924cd3bbee9bf19e423ad17da68140d1a846f6

Observation babcce6b-24f5-4b23-9bfc-00867f7882e0 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.201212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.201212Z digest=sha256:fb7493a00e0a7e570f679bfbcbfc4c06891b9623f4358464c3393c4860f38596

Pith citing papers

Observation 79481876-9f37-4128-8e6c-513d58e8d9dd · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.405635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.405635Z digest=sha256:1c709138ef343fab025c00694ae6fdd598136e79ba71e0b9487dcf7d04e58ab9

Observation 9b7594dc-5632-4def-816d-b98de7c8a0bd · inbound

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? cites this paper.

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:39:06.101648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T05:36:09.208754Z digest=sha256:5ad385fa935fab0a5347edaf53e3bbfd03ed8533a93a642a72f173a8a3057d5d

Observation cecdb1c9-4d81-4dd4-bb48-b83e4bef12d2 · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:42:06.765151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:fe264ab7eb3fcca6ec7cc548b09ae311c1ae75f1cd0d8ae58db42fbdb49776e1

Observation dbba1ab8-add9-4c98-a3ef-7ee2c9dc6af4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.144883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:74cb499cea9e65a60dfa22055ddbecaed7cacd037878239d7b2a3e52f4801c14

Observation 84dc0ba4-7c41-432b-8e86-6922425289cd · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.334651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:a424a9b8d21c1dd6e4755f7ee8fee6127769d26ccb79c1dd8ab944a21340b8c9

Observation 69338436-7dfc-4746-9633-f9fba57f9c20 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.482051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:16f14707f5dedfd547c762b7aa87cf80f83087fe92477c006710c1931d74c21c

Observation aabb4aa2-8816-4064-8777-56c364f342e3 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.899753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.899753Z digest=sha256:6da9284b4e72042e9a9ee142dfcd4920edf853345bc1b6d39b281e83bb77e8d5

Observation d29b28de-f054-48a9-8b66-5a1aa809c1e3 · inbound

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation cites this paper.

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:27.447842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:39:27.447842Z digest=sha256:6f5ef594fa8b2461e2d651287b63b756a9b573f4dcceea186ad82c6ec7de99a4