Pith. sign in

Paper Citation Record · LEDGER

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 15 inbound Pith citation observations for arXiv:2411.14794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14794 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:56:52.768365Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.517813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:27:44.037400Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8641a781-0bc4-434b-95b4-e18fcbe6abac · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection The claude 3 model family: Opus, sonnet, haiku

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.484063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.484063Z digest=sha256:2a8a62069fae500ab8a2903ebbcc53824a9cc2a7837cc3dc70ced01602a84309

Observation 85c7120d-96d8-4c19-b264-b0f7c5a054fc · outbound

This paper cites Qwen Technical Report.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.489974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.489974Z digest=sha256:4213fbd624ae789a2d97643e4da0187f1332f32d6b5c987f5588b560d252ab22

Observation 9928193e-96fb-4f86-b3cd-bbec22ea4f5e · outbound

This paper cites Qwen-vl: A frontier large vision-language model with versatile abilities.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Qwen-vl: A frontier large vision-language model with versatile abilities

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.659971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.495757Z digest=sha256:3f384a06e5c22485a0856e01f35c316d5b09693dc5e5104aaa10cbae4e6b2594

Observation 5f70f527-dabd-48af-b58c-8164af78dc4f · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.501461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.501461Z digest=sha256:49fcad9ab6fca2dd6682efcd3f95c7d3b3aa79a15a22a2edf80783241b4e7a48

Observation 411d20e4-f21d-4d13-a7f7-a189a9d5d80d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.507092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.507092Z digest=sha256:495c1cb942c39e5228074bd0c6b65e9fd528a19b069be307487db4174a8e7409

Observation 3c826ae8-d066-4145-a3ae-3be54bc84222 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.512869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.512869Z digest=sha256:5a39ff849517415063aebfaf272b189271a4d8b8699d424736c779964af27a69

Observation 0f138d19-5ee3-4f91-a669-39d6e8e99120 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.518861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.518861Z digest=sha256:78cd8a6c089cb69e4fce1f22c4b1d48b3065d5b6614f9a6788553342fb76b923

Observation acdb216b-1a99-4de0-9126-731c5d0f7826 · outbound

This paper cites Flashattention: Fast and memory-efficient exact at- tention with io-awareness.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Flashattention: Fast and memory-efficient exact at- tention with io-awareness

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.644194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.524130Z digest=sha256:b706dec165ab5bcbd68c94bad7e1a0b36a0c41cbf56a0981bfef06a3a53a25ad

Observation fcde2b7b-9746-4fe8-af0a-e564eeb36d4d · outbound

This paper cites Uncovering what why and how: A comprehensive benchmark for causation understanding of video anomaly.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Uncovering what why and how: A comprehensive benchmark for causation understanding of video anomaly

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.627782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.529670Z digest=sha256:93a2b2737618f87c1f4d9be6d4cf7f1fe4550fab3bf510c8f28f7c1c64c19574

Observation 5c488a9c-0c72-4fd1-af2e-56508565206d · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.610913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.534213Z digest=sha256:e84150bc51333a75b14b7f5f05eb4fc20f43b6e4715f3f92fc80f9ab91f0a82a

Observation 3e32a39b-7ba5-426a-8594-0cb1455c03c6 · outbound

This paper cites Llama-adapter v2: Parameter-efficient visual instruction model.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Llama-adapter v2: Parameter-efficient visual instruction model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.593851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.539049Z digest=sha256:2814ae7c65ae83b0ffe6f9972d7ca3b0896ad1fcad8023ecf75d669bfa4badf8

Observation 8824cf70-bb5b-47d3-8cc9-1f159f4bccd3 · outbound

This paper cites Gemini: a family of highly capable multi- modal models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Gemini: a family of highly capable multi- modal models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.577004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.543952Z digest=sha256:69c132a9324c998384f2580ccd08df03e7d6e6a4b63c7477bdf587d21a431d5b

Observation 4ed70c2a-779c-4006-ac02-1df47fe19691 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.561367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.548350Z digest=sha256:f29bf3428ad6691db2686f51117472b73ad937fcafbdea5b0f2e68fbf9a31cf8

Observation aeba5340-9534-41f7-ae2c-1b13d8cf77fb · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.545940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.553184Z digest=sha256:86ab29c047687248998f1d57e9b3240f9835ff3647da7fd4543b8cc82ba10f2f

Observation e5ff6060-0f1c-41ad-bd32-4125cac1a322 · outbound

This paper cites TVQA: Localized, compositional video question answering.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection TVQA: Localized, compositional video question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.530820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.558216Z digest=sha256:8fbdc22a2e5bbe26cf5ac39d3129e24dc8c17dfdcbf71d363d77ac6fe5443890

Observation a62c21cd-f56b-4679-bb28-c78ed3a17e27 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.562938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.562938Z digest=sha256:fe063d77c59a22fe6eab6b19e226d2c62aae26e6238abe78c1a00a630532582a

Observation 6a9cbef0-0a5d-48d3-8e1d-af6f58e817b3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.568199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.568199Z digest=sha256:e4b15c2e386a5375ac2dbc4b8dea4dfe2d0d7596947001c7417307401831ecfb

Observation 60b54c52-e865-41f0-97d0-99607df05060 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.573093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.573093Z digest=sha256:d4d2270274a62b925533199bd7e0567014e019b58e86bf76515a83dffe0f3f4a

Observation fe201c6f-78df-4246-b3ce-32f7ab7fbec4 · outbound

This paper cites Videochat: Chat-centric video understanding.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Videochat: Chat-centric video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.504659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.577885Z digest=sha256:023b51e0afd0314a156812f11f835ddb1cc57d118991138e01ab035cc79c4260

Observation a43366cd-085b-415f-b018-0bc61cc01e14 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.488803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.582677Z digest=sha256:ec776fc4b69090c4d185861ad265d0c74ca14fcef7cfffdc5538bf67540a669e

Observation bcda31d2-0721-47e3-9a19-f4d4f735df4a · outbound

This paper cites Value: A multi-task bench- mark for video-and-language understanding evaluation.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Value: A multi-task bench- mark for video-and-language understanding evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.472359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.587317Z digest=sha256:6ae294b17fb701b5df5a1eca06572f150ca9bfbbcfd637cd2456a11aa99a5ee9

Observation 7fc5ed11-3574-496b-9e00-0021885e443b · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Llama-vid: An image is worth 2 tokens in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.456210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.592265Z digest=sha256:80b8819b500fc925fa22f1720394e84118ef6677e0f68185020bf75546f44b84

Observation 625876c8-24c2-4cdd-a518-2901bb5f4f51 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Improved baselines with visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.597136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.597136Z digest=sha256:04b45b40b06d6bf37d1087039b223faffb22723f2e9c9f44f2b8011be2ac560d

Observation e7e40631-1345-4043-a567-36346264d6b9 · outbound

This paper cites Visual instruction tuning.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.430872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.601683Z digest=sha256:d4e7dae9bd6d1eb9012564c3adfab2e563a84b27330ec3561541664ac336589e

Observation 5dd7e9f7-39d8-4654-bfa1-f3c8c5ab97d1 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.415490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.606265Z digest=sha256:1aadeb699414c60e78c76f50c54069cebe0c870eb61010e6fcdb0ec719e421e8

Observation c5877974-3f6e-49c8-af9f-5770b982b22d · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Sqa3d: Situated question answering in 3d scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.399479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.610785Z digest=sha256:975eaad184e5754e73e24396afdab1d05ef30d9e803eeb13d3c4bfae272783b4

Observation 72688aa1-8467-44c6-b669-fe6870ff248d · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.615499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.615499Z digest=sha256:58c28bb4e67909e6ef6b89f783d3b6e84c8463e7f20ba2849026e532ecc7f526

Observation 44d1814e-8a34-4ea8-91f0-1389ffac799e · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.372250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.620103Z digest=sha256:850e4ca568de1131e98bde8667f3e81ca5f6b7c471b51cfff8209a8d8e4c0ba5

Observation af1654c0-9205-4ab0-abad-b21ed12915eb · outbound

This paper cites Introducing chatgpt.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Introducing chatgpt

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.356703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.624543Z digest=sha256:fb16a6c0bbcfbc898c4e36e34ab1171445054d869cf3dbe6962124207b2bd1e7

Observation 11f4dc2e-d1cd-422f-a0bd-39e5a1614101 · outbound

This paper cites Gpt-4 technical report, 2023.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Gpt-4 technical report, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.628955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.628955Z digest=sha256:35d79f6a3bfb25fa83735072c0c1e273874a451352b2b2a12587ce3c4900cfcd

Observation 44d89e79-df55-4fc7-bfba-eae03d9b5478 · outbound

This paper cites GPT-4o system card, 2024.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection GPT-4o system card, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.330580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.633463Z digest=sha256:07ff992159b23c32604f2c983a26d28dfb2bebccbf1cb81751154679feae1bb7

Observation c454eb37-54e2-4bf1-8fe7-159b30ac9e2a · outbound

This paper cites Learning transferable visual models from natural language supervision.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Learning transferable visual models from natural language supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.638288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.638288Z digest=sha256:6f0f637b81ed5504e51e5b427f7c8d4f3164763b5b72ecaa3b539d2e871133e9

Observation 30f9992a-c4de-498e-9dee-bfc9b9bb0aac · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.303671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.642869Z digest=sha256:e8c5016eeb5c7ba1753c687e9db2a68f732e31e49d12324fbd16ed15dbb8da6f

Observation 6de5819c-e30f-4ad1-ad5b-b15849cc3b50 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.647475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.647475Z digest=sha256:1fd0d9a7d867653ef850c94ad49730644559c31a6b5697aa560025cd6069361e

Observation b05a166d-f99d-4c64-8c8b-de1e10ab1756 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.652576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.652576Z digest=sha256:3cfd351dd2669e1141228ea670b12959c19c5887f506c5c1d74de5017fd30cba

Observation a9c3a6d5-6636-4762-aa94-d650846fe23a · outbound

This paper cites Tvsum: Summarizing web videos using titles.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Tvsum: Summarizing web videos using titles

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.657398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.657398Z digest=sha256:d7d3cc85627e4332a250043d58a290170c7ea3f57b093eac55cdd623af38a78f

Observation c97b8a4a-104a-4e43-af7a-31fcdbfa7a28 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.276685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.662417Z digest=sha256:aeab5e3a052b7f09d5785e8fe0515d40e8fddf3a16e153b4f584cf2c481f218a

Observation d942e3ad-8fa4-42f2-9bb1-217320a85e95 · outbound

This paper cites VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.667304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.667304Z digest=sha256:419c990f1720f0fed75bda2483745f04e788003dd0a4b293cdde3fded2de4085

Observation 70657d85-1f6b-4de7-b88f-5099a1b45e62 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection V?: Guided visual search as a core mechanism in multimodal llms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.672777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.672777Z digest=sha256:48039bb14578d7670feb1eb9ed4022c3e5bc420afd853c47c63c9d54bb282823

Observation 321f0513-c418-4969-9e4c-59befaff0969 · outbound

This paper cites Not only look, but also listen: Learning multimodal violence detection under weak supervision.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Not only look, but also listen: Learning multimodal violence detection under weak supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.250506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.678184Z digest=sha256:8bd60549c86bdfab3f71d2a8daa7285e625659158105a5eb845efd4a8590c57b

Observation 493aed7b-a2b1-494b-b165-eb140b959826 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Next-qa: Next phase of question-answering to explaining temporal actions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.233401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.683334Z digest=sha256:c764f78e0f6b207ebf6459280ed2ae77f7bddba43b14e56f0989a235aa1f940d

Observation af2f41ea-8249-43a7-8801-207942f8d6fa · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Can i trust your answer? visually grounded video question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.218358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.688424Z digest=sha256:af2500e41ada42b4bed989af1bd6ca5f9337b964e3d66493ed47bf29188e40c6

Observation 55bb20a7-592d-4497-ba32-88c662b998ac · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.202097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.697206Z digest=sha256:a53df3b05e7bc5279bce3d09747f484a02bf3455ca1173deb0cfb8e0d89767f1

Observation b20491cf-043c-4fd9-8ce2-7f055c0f8b0a · outbound

This paper cites SEED-Story: Multimodal Long Story Generation with Large Language Model.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.702184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.702184Z digest=sha256:d8cf280bd04f73a30a7311b44f8e738d763c4b4ed337a08eae8020b2e6cfe9ac

Observation bb340ce1-f8a2-4ce9-875f-2b0b92abfdff · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.707406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.707406Z digest=sha256:15b64154a9112e3f033e7fa56cf5dc31912119d05fcdb210a74e41800748bfc7

Observation 4948f57e-595c-41f6-b83f-413d9df28bdc · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.713928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.713928Z digest=sha256:db822a649fc7b1bf5f30748f2d0636d10c49abfa310157655a6bd55fcea3fd31

Observation b747631e-2b9b-48cf-86e7-bca5338605a3 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.184416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.718907Z digest=sha256:4389bbc1d6c391f0585acecdd53a659e5c32aa9d7e23e87988cd456033804a06

Observation 7c92aac0-0429-4f31-a787-7c6324d4435f · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.168123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.723619Z digest=sha256:f1948c97b2560a7baf27b3726952137017f97ef22d074b686c2a6aecd549cacf

Observation 6d64b167-4ce5-4b1d-be79-72c8c1212010 · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Long Context Transfer from Language to Vision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.728364Z digest=sha256:5950e94d26d757a4036a613e3d37adbdacc1ad52aef1e099a75000641c91f358

Observation aa2ecdde-335b-4768-81aa-61fda51fbb4c · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Llava- next: A strong zero-shot video understanding model, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.733167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.733167Z digest=sha256:655a6b72fd1de76a4ca71eef2cdd6dd820a47237fdebca4ccf143ce2d89ee37b

Observation 4c9d4596-9f2e-4a06-9fc4-8c1c1b20c3af · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Multimodal Chain-of-Thought Reasoning in Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.738306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.738306Z digest=sha256:0810d1f0278a978d97bd118b47c11ba2defb19ee1b94e1309f59c0743a647ad9

Observation 242b06e8-a6d5-4135-afdd-b1764cd8371e · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Towards automatic learning of procedures from web instructional videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.141081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.743262Z digest=sha256:77a090c2a7944447ec4e5f95d92fea112e1066ed62ebb7ea4abba9a3715bf4a5

Observation 327e4af4-8173-4c53-8f94-6074114f51d2 · outbound

This paper cites an unresolved cited work.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:56:53.123560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.748535Z digest=sha256:e9e3fd69582c4c7538fb3884d3e68a5993e73d0cb3d5a70174ff210dbaecd120

Observation 43550815-a086-4052-ade4-f7b55daf3735 · outbound

This paper cites an unresolved cited work.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:56:53.107584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.753453Z digest=sha256:fadeb49cc4ddd0eb3154d9e46b339607ec365ba08ab2ce1480f66f5e694c6cd3

Observation 33395fe7-6540-4c64-a065-2714a17788a2 · outbound

This paper cites For example, by the cause of certain images, the result of certain images is obtained.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection For example, by the cause of certain images, the result of certain images is obtained

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.091825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.758011Z digest=sha256:1d0b15a778ec713e695877892da1f79d3d398a0f3efbba5196db68634ef86060

Observation e7bf41c0-ff9a-4ad8-bc15-55a8b3e40d68 · outbound

This paper cites an unresolved cited work.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:56:53.073933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.763381Z digest=sha256:2e113f3ed8c6e1395f2710231119bbf3a0979185c773fb20cdc4b7d179c256fc

Observation 3f8b17e1-37c8-411b-8408-11f7efa482a7 · outbound

This paper cites Subjective Question,.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Subjective Question,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:56:53.048253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:56:52.768365Z digest=sha256:1656daf56756a526f99d0184b9a337ad1bf5e00433bc074614e3ec5e0716b6a2

Pith citing papers

Observation 3712b52b-dec3-41ff-9eb8-98ae17197a82 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:44.041073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:e5b8ba781e9a238dd4ad6846b4897b4a1f95c6a72409a40b7006993631de1be9

Observation c2d0ea23-9410-47c9-b020-8e6ac73192ec · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.912764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.912764Z digest=sha256:af654cde8f92b575ec203e54f1e758e839830fdad7eac8531c5a620f1b173660

Observation d5d8e53a-eb28-4c15-82dd-5b8c67579206 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.266805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:39fb6c44c34112f73194f6eae97a1884558978f7f6b7157e33abc0069c06f755

Observation 06a0620f-8b22-45c4-b8f8-45e47763aa1d · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.517813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.517813Z digest=sha256:6f17cb06069dee4cd32a6b61169eaaa5bfb7b9db78b1003eaecc19fa88873885

Observation edbd45b1-9d19-425b-bd23-9273d7ea6f04 · inbound

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification cites this paper.

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:20:20.286412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:20:20.286412Z digest=sha256:4b24d295fff193680111fbded9aa147dc254ff800799f947fce5eb243ee846b8

Observation daafd92f-7778-4657-841c-37cdd492c9a2 · inbound

MINERVA: Evaluating Complex Video Reasoning cites this paper.

MINERVA: Evaluating Complex Video Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:19.069310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:19.069310Z digest=sha256:1e7c53bc9c666343f8af4a0c73c9539ea8f8cdb21e4918a6aa71b5e01d735a24

Observation 4b8a6322-451b-4ed4-ad9c-d2caa6fce31d · inbound

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR cites this paper.

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:03.727067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:03.727067Z digest=sha256:665017f5dff40e40e4267c4cbbd88af3e7ae5755906840f44ebde3d26b9566b0

Observation 2771756b-e8fa-4632-881e-95c545e86d39 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.372866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.372866Z digest=sha256:83077eb7682496cf182186d218329b97919ac92f1f230a8c33909fef0df3728f

Observation 7c58fbad-de91-4d02-b625-61d4030b97fa · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.852974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.852974Z digest=sha256:34090dc2a91683ca96598b509181e503403b5bfa17efb6ed025f0bc732f33c0b

Observation 7b07519b-5ef0-457f-b17f-6d27a24ad23e · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:54.039135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:54.039135Z digest=sha256:a29c5203502d2f76a3db98947cc694bb384c9138faecc29bf4b6dd325a25b308

Observation 55638bfe-9742-4f24-9d1a-33ba2c263953 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.618589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.618589Z digest=sha256:2522a91652ad88dbb21ed7c6226b849102750c9a82bf3571fb84ab0b5e856b22

Observation 83fac34c-6574-4a22-a63a-a92e08f38943 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.372667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.372667Z digest=sha256:13634330506ac133c4fd8711d7137b2fa6bc618dce617ae2aa903057bca7b726

Observation c9c3a023-4234-4918-89ae-f0b15cb0194e · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.642866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:de22ddfb09432407660751c1777601218a5c16e65a5baf5849b0e579c1454b24

Observation b8bc99ed-566d-4acf-9f27-d1f693e9db29 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.051371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:5ddea85502f480d362ec3f91396bb26514847d750e3b85d019f49cdc67473180

Observation 258700ec-3943-4844-af85-1f3c86fba121 · inbound

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning cites this paper.

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T23:10:24.087228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:10:24.087228Z digest=sha256:d2bb35e75b6c4d59c55575a40412477f05ba17b419b918c0d209cb7bedadbb0e