Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:44.490450Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2505.11326.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:44.490450Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75cc6bdf-361b-4106-9598-db20beabe6fb · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Whisperx: Time-accurate speech transcription of long-form audio
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f03ec7-e3b7-4e61-8478-188164d96183 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f98d959-d7c1-407e-8ec2-e042c4a3f884 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Collecting highly parallel data for paraphrase evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ec80b8eb-3278-4975-bb24-63c30c98d033 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Videollm-online: Online video large language model for streaming video
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cf206b-517a-4e58-9020-752a3809d13f · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Scaling up soccernet with multi-view spatial localization and re-identification
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fe29c422-2803-4947-9b12-29889e1c19d5 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7232378e-932c-4555-aab4-93690da115c0 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6849114-f0f8-4624-ba96-25c03349e43d · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Lora: Low-rank adaptation of large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab8fa84-8620-4326-a328-c5bc05163bfa · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dc99ae8-e853-4700-ba2d-df6efd0f2259 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models The Kinetics Human Action Video Dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c910d08-713e-4632-b3a9-38aacb82405e · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Kuehne, H
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f983df-8579-40c0-badd-30d8c695e420 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Revealing single frame bias for video-and-language learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7a362db-5f54-4548-b73e-99ceb56abe63 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a941e4-affd-461e-ba72-7bc96ba969f5 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Video-ChatGPT: Towards detailed video understanding via large vision and language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c723c2a4-4c43-482c-9f4b-a56295a245a5 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d797e867-3737-451a-b6b2-bfb6ac20aec2 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Sentence-bert: Sentence embeddings using siamese bert- networks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4620d1-88e6-470b-95ea-b89dcbff0d32 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Ego4d goal-step: Toward hierarchical understanding of procedural activities
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 26f36893-dc16-4247-8980-214f94fc6d29 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7c91f7-c8bc-4515-9515-da0647bdb275 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models LVBench: An Extreme Long Video Understanding Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fbeb405-5b44-461d-a139-de0c6372cb3e · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 563f63b5-921f-4b96-827e-5d52e4954507 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c0ae33-b22c-4b93-ac37-f72554f67797 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Video question answering via gradually refined attention over appearance and motion
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4aae3c04-e1d7-4d26-8935-6e9bb160baa2 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Msr-vtt: A large video description dataset for bridging video and language
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6854d76f-e17e-4701-949b-73fa9386dad4 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4187126b-7494-40be-bdfb-29e760759ba6 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Eliciting in-context learning in vision-language models for videos through curated data distributional properties
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d8794ff5-f4a5-4c5e-a1fd-1b8f530ed218 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 72b39727-9394-4d29-a9f6-046ca197b88b · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d01f6e-a1e8-46ca-b135-c1b904cdd358 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfdbc82-a9f5-4d8c-a563-a0de92771cbb · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Streaming dense video captioning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 819f5db5-769b-4bb4-b26e-1449ad8caa05 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Streaming dense video captioning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f3643ec8-d4c5-4917-bc67-5acdd74abe71 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Assistant: Now remove the indicated component that’s damaged,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2cfa6a8a-4c17-49a7-aa41-814f6dd7ccb3 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models User: Oh, this thing?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 54a7b08f-9c91-48d5-b906-b79e2f593a37 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Assistant: To the right
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d68f24af-eb19-486d-80b3-4a095169399a · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Assistant: The small cube
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 257df93b-1121-4904-89de-6ab1c86d47d6 · outbound
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models Assistant: Yes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.