Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:56:52.768365Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 15 inbound Pith citation observations for arXiv:2411.14794.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:56:52.768365Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.517813Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T09:27:44.037400Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8641a781-0bc4-434b-95b4-e18fcbe6abac · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection The claude 3 model family: Opus, sonnet, haiku
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c7120d-96d8-4c19-b264-b0f7c5a054fc · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9928193e-96fb-4f86-b3cd-bbec22ea4f5e · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Qwen-vl: A frontier large vision-language model with versatile abilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5f70f527-dabd-48af-b58c-8164af78dc4f · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411d20e4-f21d-4d13-a7f7-a189a9d5d80d · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c826ae8-d066-4145-a3ae-3be54bc84222 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f138d19-5ee3-4f91-a669-39d6e8e99120 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acdb216b-1a99-4de0-9126-731c5d0f7826 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Flashattention: Fast and memory-efficient exact at- tention with io-awareness
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fcde2b7b-9746-4fe8-af0a-e564eeb36d4d · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Uncovering what why and how: A comprehensive benchmark for causation understanding of video anomaly
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5c488a9c-0c72-4fd1-af2e-56508565206d · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video-of-thought: Step-by-step video reasoning from perception to cognition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3e32a39b-7ba5-426a-8594-0cb1455c03c6 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Llama-adapter v2: Parameter-efficient visual instruction model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8824cf70-bb5b-47d3-8cc9-1f159f4bccd3 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Gemini: a family of highly capable multi- modal models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4ed70c2a-779c-4006-ac02-1df47fe19691 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aeba5340-9534-41f7-ae2c-1b13d8cf77fb · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e5ff6060-0f1c-41ad-bd32-4125cac1a322 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection TVQA: Localized, compositional video question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a62c21cd-f56b-4679-bb28-c78ed3a17e27 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9cbef0-0a5d-48d3-8e1d-af6f58e817b3 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b54c52-e865-41f0-97d0-99607df05060 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe201c6f-78df-4246-b3ce-32f7ab7fbec4 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Videochat: Chat-centric video understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a43366cd-085b-415f-b018-0bc61cc01e14 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bcda31d2-0721-47e3-9a19-f4d4f735df4a · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Value: A multi-task bench- mark for video-and-language understanding evaluation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7fc5ed11-3574-496b-9e00-0021885e443b · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Llama-vid: An image is worth 2 tokens in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 625876c8-24c2-4cdd-a518-2901bb5f4f51 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Improved baselines with visual instruction tuning, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e40631-1345-4043-a567-36346264d6b9 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Visual instruction tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5dd7e9f7-39d8-4654-bfa1-f3c8c5ab97d1 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c5877974-3f6e-49c8-af9f-5770b982b22d · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Sqa3d: Situated question answering in 3d scenes
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72688aa1-8467-44c6-b669-fe6870ff248d · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d1814e-8a34-4ea8-91f0-1389ffac799e · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation af1654c0-9205-4ab0-abad-b21ed12915eb · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Introducing chatgpt
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 11f4dc2e-d1cd-422f-a0bd-39e5a1614101 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Gpt-4 technical report, 2023
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d89e79-df55-4fc7-bfba-eae03d9b5478 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection GPT-4o system card, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c454eb37-54e2-4bf1-8fe7-159b30ac9e2a · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Learning transferable visual models from natural language supervision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f9992a-c4de-498e-9dee-bfc9b9bb0aac · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6de5819c-e30f-4ad1-ad5b-b15849cc3b50 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05a166d-f99d-4c64-8c8b-de1e10ab1756 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c3a6d5-6636-4762-aa94-d650846fe23a · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Tvsum: Summarizing web videos using titles
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c97b8a4a-104a-4e43-af7a-31fcdbfa7a28 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Vatex: A large-scale, high- quality multilingual dataset for video-and-language research
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d942e3ad-8fa4-42f2-9bb1-217320a85e95 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70657d85-1f6b-4de7-b88f-5099a1b45e62 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection V?: Guided visual search as a core mechanism in multimodal llms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321f0513-c418-4969-9e4c-59befaff0969 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Not only look, but also listen: Learning multimodal violence detection under weak supervision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 493aed7b-a2b1-494b-b165-eb140b959826 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Next-qa: Next phase of question-answering to explaining temporal actions
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation af2f41ea-8249-43a7-8801-207942f8d6fa · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Can i trust your answer? visually grounded video question answering
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 55bb20a7-592d-4497-ba32-88c662b998ac · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b20491cf-043c-4fd9-8ce2-7f055c0f8b0a · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection SEED-Story: Multimodal Long Story Generation with Large Language Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb340ce1-f8a2-4ce9-875f-2b0b92abfdff · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4948f57e-595c-41f6-b83f-413d9df28bdc · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b747631e-2b9b-48cf-86e7-bca5338605a3 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7c92aac0-0429-4f31-a787-7c6324d4435f · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video-llama: An instruction-tuned audio-visual language model for video un- derstanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6d64b167-4ce5-4b1d-be79-72c8c1212010 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Long Context Transfer from Language to Vision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2ecdde-335b-4768-81aa-61fda51fbb4c · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Llava- next: A strong zero-shot video understanding model, 2024
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9d4596-9f2e-4a06-9fc4-8c1c1b20c3af · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Multimodal Chain-of-Thought Reasoning in Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242b06e8-a6d5-4135-afdd-b1764cd8371e · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Towards automatic learning of procedures from web instructional videos
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 327e4af4-8173-4c53-8f94-6074114f51d2 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 43550815-a086-4052-ade4-f7b55daf3735 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 33395fe7-6540-4c64-a065-2714a17788a2 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection For example, by the cause of certain images, the result of certain images is obtained
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e7bf41c0-ff9a-4ad8-bc15-55a8b3e40d68 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3f8b17e1-37c8-411b-8408-11f7efa482a7 · outbound
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Subjective Question,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3712b52b-dec3-41ff-9eb8-98ae17197a82 · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c2d0ea23-9410-47c9-b020-8e6ac73192ec · inbound
CoS: Chain-of-Shot Prompting for Long Video Understanding VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d8e53a-eb28-4c15-82dd-5b8c67579206 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 06a0620f-8b22-45c4-b8f8-45e47763aa1d · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbd45b1-9d19-425b-bd23-9273d7ea6f04 · inbound
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daafd92f-7778-4657-841c-37cdd492c9a2 · inbound
MINERVA: Evaluating Complex Video Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8a6322-451b-4ed4-ad9c-d2caa6fce31d · inbound
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2771756b-e8fa-4632-881e-95c545e86d39 · inbound
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c58fbad-de91-4d02-b625-61d4030b97fa · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b07519b-5ef0-457f-b17f-6d27a24ad23e · inbound
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55638bfe-9742-4f24-9d1a-33ba2c263953 · inbound
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83fac34c-6574-4a22-a63a-a92e08f38943 · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c3a023-4234-4918-89ae-f0b15cb0194e · inbound
Video-ToC: Video Tree-of-Cue Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b8bc99ed-566d-4acf-9f27-d1f693e9db29 · inbound
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 258700ec-3943-4844-af85-1f3c86fba121 · inbound
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.