Pith. sign in

Paper Citation Record · LEDGER

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

As of 18 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2505.08455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08455 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:33.853704Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved52
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b89677d-3b55-4f0b-b0fa-46f60380ab2e · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Robovqa: Multimodal long-horizon reasoning for robotics

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.239931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.444017Z digest=sha256:71e1d4dfcd36c30e2f3297a4d700017f85f3eb421f49c955287febc05fc7b9a8

Observation c878ab8e-0a4b-4286-9bb6-0c705479fbd3 · outbound

This paper cites MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.450018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.450018Z digest=sha256:5640146aabf0662d621b246fad3b64d1b8b20fd44889afec9493995dc42abefc

Observation 236615c5-5e79-4ada-ba59-d2b445ddf7cc · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.455775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.455775Z digest=sha256:669c22661e3e6e340a505c2a9fec25d922f4e6291dd3975044866600c8814d35

Observation 144edce9-8fb6-4610-bc43-0d037e06a919 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Openeqa: Embodied question answering in the era of foundation models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.224949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.461137Z digest=sha256:3e32505af81217db4a1c099c44173b0311cde0e708638f27a088ff3f7ff7ba29

Observation d80e3b03-69f1-440e-8910-a0d0c85dc28a · outbound

This paper cites Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.466110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.466110Z digest=sha256:195b8bdf1e7a81cc21cdb6e6d72fe25a6cd48e8f625c3648074e91e9e3f100a5

Observation 6f59abb2-0476-4a6d-a816-b4485f291313 · outbound

This paper cites A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.471828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.471828Z digest=sha256:d02eed8d69339a89dc09981800180868248cdd3394d7753a2b867e0f810339ba

Observation 45022173-519c-455a-bc87-57177bbf8318 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.208399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.477225Z digest=sha256:79a22463d4f3241ea3913e660ba71cb790cab2338045ff633a13537f7d0d6486

Observation 6cb698aa-a7bd-4323-8b10-e43c707f888d · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.483140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.483140Z digest=sha256:dd373f28a73e303e038536ff6ca89ba6f8842098fd3cfa5295edcdce2257fc70

Observation 46c22944-33a7-4cec-936e-4a58d7efeb44 · outbound

This paper cites A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.489753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.489753Z digest=sha256:496f628645fd9aa5c1ebc5ab37b2cdd110cc98d3408db4ea4d8673b516356ffd

Observation 81b998e8-84cd-4ebd-82ad-3ab02563bc02 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.495030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.495030Z digest=sha256:247f6817155f4510b44e8885762b1bbfc41f155fc085fd3ca424ff30e8c74e92

Observation e50418e7-bf9e-4d51-86fd-1264daa20d9b · outbound

This paper cites Learning transferable visual models from natural language supervision.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Learning transferable visual models from natural language supervision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.193579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.500010Z digest=sha256:0d545a248e5a137335c42f7ab3dfd67d05e9ae756d70bf0818cbedfd4700cd4b

Observation cdff0e5f-ee66-4d28-a96b-3f5716613f57 · outbound

This paper cites Sigmoid loss for language image pre-training.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Sigmoid loss for language image pre-training

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.178154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.504583Z digest=sha256:e6d717cdbadbbb60617e4f5b341e7ac30f72e955fa8fd612674f47c80a4fa80a

Observation 82f532e7-71c2-45e0-b07e-e60c41ae0a9c · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.509378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.509378Z digest=sha256:d20e767333f6e3faf3b2e04878df254267e1fd42a5ab5a5c0b37ddf8b93b5ed3

Observation 96598b34-9e8d-4787-afd8-98e1cfe4e53e · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.514987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.514987Z digest=sha256:777b85ad04f03c69deff05c402cef5b5e0ea5a364e795435547f510c0dbb67b2

Observation 7270f04b-5650-4b02-bbef-ae2a4a921ff8 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.520336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.520336Z digest=sha256:eed616a17da31df1f47283877d8e21c0b8cfa5fce2aa6458287e3480a45f3739

Observation 605c03e1-405b-4afc-b7ee-1ac3815bde4b · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.162399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.525739Z digest=sha256:c3b58c20a83d954841d263aa60dc678d6ac5d919670ef481376e16a982f0bbfa

Observation a418e650-accc-4d06-8ae2-41cfbd230895 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.532560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.532560Z digest=sha256:549bf57379c2ddbd63737ba00549b61366e446890240a97b84195063c2c78d71

Observation ae912172-be3e-4f6d-ac41-6e33bfe252ab · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.537472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.537472Z digest=sha256:0299d7a99f849046bf612adf3432deb3b1d353bd3e9fb11e42fda6af969eb296

Observation 50caf98d-008e-4c68-b7f7-a27bc88dd2a9 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.543100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.543100Z digest=sha256:e1103427fb3f06b09ea7e6903b14edd4b2bbeb001deebee02ae1d7e1b6e6d75e

Observation 9cf4d981-c735-4225-bb99-f5f4764d190b · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.147759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.547817Z digest=sha256:a4e343e56e18df53bdf09ad947ba46e6c021d820fcb7f72e4400a65eae25e5a0

Observation bb79dd26-bed0-4924-b083-bc95c8cfb7b3 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.552216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.552216Z digest=sha256:44f5118b59e0e0cae62d1c738aac839cf68e0641a930b9194178ebad6d35099c

Observation bfdc15e9-adde-4fe9-a13d-624b11eb2769 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.557150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.557150Z digest=sha256:f3399bcb6eeb93a7805ad763302c066a629faed140bbbf504f6783be6ca73f8d

Observation e5439363-9c59-44d3-a269-b8b7db28832c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.562745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.562745Z digest=sha256:fe5213a72ed2de32914175a2a1eb867148bc9ed67d95566c7c926a802de12ec6

Observation 578520e5-b3bb-4afe-9cc2-8bb64ce22f25 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.567661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.567661Z digest=sha256:de6119650e62d698969ea0493ac4511cdec5adf19ad555fc655e5d3f26222b62

Observation 43c1aa7a-e644-415c-9837-9a358b3c61b3 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.132123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.572655Z digest=sha256:8e4b0dca722ac7bafacc4635198cab114cd9b135b059c0c2023b3d2b0fc52202

Observation 1ebae743-b61f-48fa-b400-fbdbd49df6e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.577918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.577918Z digest=sha256:907ebb3c09c9683d8c90dffdc83ab1b6ab11c399347659d406fc9df7ffbf024d

Observation abc00dbc-3b25-4895-8253-99514193ed0c · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen-vl: A versatile vision-language model for understanding, localization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.116420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.582957Z digest=sha256:806662f228de76f8c95fac79aa427b1e8ea4b99ae4914ccdc7dc9bc2fbfaced2

Observation ed7d0b5a-a7ca-4ac5-b33c-f162c980f986 · outbound

This paper cites Qwen2.5-VL Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen2.5-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.587609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.587609Z digest=sha256:fd53068a81edf2201f4ed89ec0cc59045dc1cc4b51165133b1e10f94342c029d

Observation 8d2f1091-ee92-4f41-b884-c575b4d27450 · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Intentqa: Context-aware video intent reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.084753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.597957Z digest=sha256:9a647ce7b866afe2fdeda5ab8989ab2f9376b1f72fe601992ea09743ad2a24cf

Observation 852a0e58-1b22-4e8e-903b-0ef34fe7db90 · outbound

This paper cites Rex- time: A benchmark suite for reasoning-across-time in videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Rex- time: A benchmark suite for reasoning-across-time in videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.069675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.602270Z digest=sha256:77513ab20744b408d9224e00a50167a100dd7a5021c44c3724090fe67af122f4

Observation 95eb2e73-826e-45ec-ba3f-6e5aa90e6e69 · outbound

This paper cites From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.052199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.607218Z digest=sha256:66fa5b6374fbca1872326b52eab6c1691c1eb5117c07e6cb5f677caca06b051c

Observation a9ef2c91-79b1-4a9a-ab86-7279b66c08b5 · outbound

This paper cites Long Context Transfer from Language to Vision.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Long Context Transfer from Language to Vision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.611835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.611835Z digest=sha256:23bf0ea6bde888d1d940380482d5814bec74f292557127ea898b375145bfa972

Observation 906e0ef9-b0e3-40b6-a7d0-d3b554e4d1d6 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.036159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.617442Z digest=sha256:cea4e8faf752373161f0cb41bea0c06883c47aee1767bfad0acdf5a3ef559d03

Observation c725de8b-4d3c-42b8-bbb1-e77eda694019 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.622058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.622058Z digest=sha256:b3887f9afbb6936242e92995ce8db2292b3617b0aea136e698ba2d234b53615e

Observation ec3dd096-487a-4ce7-8741-b02ab8275c7c · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.626864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.626864Z digest=sha256:c76064b4bc450fae27c08027613e897689a4dcb3a6feb4cc672798ba3afc6fb6

Observation 01a1d4b7-fae9-465e-a4fd-a4c44af26b47 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.631527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.631527Z digest=sha256:ac186263a1ece654ac0c82eaa830106e7846a2ed0ec6dde97624775958636ec0

Observation 4e1f5f85-d232-4aab-8ea3-c116692f5b5e · outbound

This paper cites Video understanding with large language models: A survey.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video understanding with large language models: A survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.637682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.637682Z digest=sha256:ae82c214bfe46153e1cb185ad4dbc3ad70cb2757101153a3a1a5a53b0525c7be

Observation 920f45ec-7d6e-4994-a205-3e17288a260c · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Foundation Models for Video Understanding: A Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.642396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.642396Z digest=sha256:34c1e44027dd8c3d720909c5a4175a8efc43ffe692df1806f6c5e1a5338a47b1

Observation 6ae3780d-0c83-40ce-a81c-fa06d730fe3f · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.647351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.647351Z digest=sha256:015bbe3f66dec356204ccb0f7b7a9241ca209f3d9107ca4f60837527b6ba3e5e

Observation eca70cb1-cdc6-4f87-93d4-b83935edb694 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video question answering via gradually refined attention over appearance and motion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:35.008796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.653452Z digest=sha256:b641db2cda055734b6fdc7368fc6d66c891b4f77202a60d29b5ba0895dc82ded

Observation 2a023add-8aef-4e6e-8b2f-07d2cf103131 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Next-qa: Next phase of question- answering to explaining temporal actions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.990702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.658131Z digest=sha256:72b9f4db064101d7824b34e8f1f843859804b59c3e2291cfef491e537acbeb57

Observation acea9151-c1c2-4783-9c9f-1d659178ccc0 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.975540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.663554Z digest=sha256:c4413ab0df3543d12af2c8bff75e77c7c7b05f3a05a1ff36b4642f49c899fccf

Observation e75fe3f7-da28-4ef7-8baf-7a88741e82dc · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.668469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.668469Z digest=sha256:b77bbf7c7e21c7d4939eba6f00206dc0a86e8824555b8f7879024db3b3103aaf

Observation e39db794-266f-4f71-b8c6-afea1340d505 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.674311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.674311Z digest=sha256:c5daeb9b341df2a5ece0d2ed78dacff7eb909558c2fb71f8182585655f854d48

Observation 0efe9d6d-0a96-4d3d-b9df-09d0c1139a53 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.679506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.679506Z digest=sha256:54aae21aba1d1b523a3dc0cfe867a670ac15f9a872d37bad96c987fced994be5

Observation dd700fdb-2a23-43c3-9a61-04b000dd0e50 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.684509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.684509Z digest=sha256:56b9e4b5b66f29224764ca82b0a477096c999ce650a956d770ba00a39b230975

Observation 861cd7f6-85bf-478e-bea8-bf3dc51f9482 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.689913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.689913Z digest=sha256:89b2784c9ad96479f173b5f038c40668e51b08e3ad0e089c212592920c125cd2

Observation bb3b54af-c5be-4614-b2b4-575ef17ec51b · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Longvideobench: A benchmark for long- context interleaved video-language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.959836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.696191Z digest=sha256:a3c2aceab3ad60306c0797b9f6ebda706fc597acf41a31c49001665099660694

Observation 4af261fa-79dd-48c9-a700-bc67770e9b37 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.944426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.701111Z digest=sha256:ec099db28c5cac9d504591a68f38ca45b8c9f5cdcfb1df1cb9463d3e93e1e5ec

Observation d0eb0d0b-aa1a-4a6d-83de-645e6b7d1a3a · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.705892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.705892Z digest=sha256:949e7dff6ef126e43842acd5be32d43c5b732e76e5efbb97d759adab8b0b0221

Observation 3d5e342c-f4bb-4d8b-a56c-51bb49b5abcd · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:59:33.711050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.711050Z digest=sha256:560efc760d51cf4bce58c8352f9252252ae9abb2af3322af8696d7c7d0de4424

Observation b934793b-b950-48ce-93e9-5c77b1a3a9c0 · outbound

This paper cites Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.929308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.716606Z digest=sha256:ea21281f6b939157fff75024c6813b60240ac138e61211306f37f6a889708e2b

Observation fa4837b5-47b0-42d8-8d65-e4e6cd901d7a · outbound

This paper cites MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.721333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.721333Z digest=sha256:f3a9cc40d84597b6c25deac70df8cf51c6197948765ea1aa69f120b7b8c5c564

Observation d78f15c0-b7e3-4db3-8dc0-380def64992e · outbound

This paper cites ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.726301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.726301Z digest=sha256:b0084802b65bddc22d7377ccf5237c53258456624e57ee658a67e97e4064480d

Observation ac8e6e87-15f4-430c-9521-2b29d9083cbf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.731448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.731448Z digest=sha256:36fb09b2f52046ee27dcc93168dde4995b44d86908fcb1bf21687223eca7c493

Observation 8580e34b-d30c-4b28-b1e6-c0848ed85041 · outbound

This paper cites Deep reinforcement learning from human preferences.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Deep reinforcement learning from human preferences

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.904330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.736464Z digest=sha256:b4da4fee5a186018e4b63b8782586a2e5ad28213b192746a0b96f7d15fd7a49c

Observation 89df4697-23fa-4319-9127-b731d1cc5aac · outbound

This paper cites Proximal Policy Optimization Algorithms.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.741587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.741587Z digest=sha256:5b98e292b497d136b4825f4edd87a6887f25291713c640e9126d9d8f30f08c01

Observation 97814645-f452-4faa-8251-db811fa094db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.748580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.748580Z digest=sha256:d988281e65f124c9757c79813b989f8a1d2d92f89a718e093935b95e023ced29

Observation 24680908-ecc4-4b80-bf19-e826a679e737 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.753498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.753498Z digest=sha256:977547a37a9cef89c12ee7e6fc7f5f5e3491ddbb2f14efed3dfb0a7e71794aed

Observation 9e1edf25-7ef5-4916-a861-67fdbd79e8fd · outbound

This paper cites Learning to summarize with human feedback.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Learning to summarize with human feedback

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.889209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.758432Z digest=sha256:e10612f82d3016201d1e36d1324cd774e74008d5e81fa6f9ef2adfb4db1c59b9

Observation a8c36977-6c37-45b7-9b12-17e04a61dd27 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models WebGPT: Browser-assisted question-answering with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.763413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.763413Z digest=sha256:6baa39ae1910edbea7a90cd78b58aa1ac3347fbb948151453f892053cf8ad830

Observation 95e00432-1208-4893-87e9-0d48cb358aa9 · outbound

This paper cites Decomposed Prompting: A Modular Approach for Solving Complex Tasks.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Decomposed Prompting: A Modular Approach for Solving Complex Tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.768463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.768463Z digest=sha256:28f02ba4c6a2ea84ff060fa6ba0dc213bbdbcac52f526b9fe23fe0f6efa4bd5f

Observation 7501a49a-7ab1-4c2a-b32f-8c0ba6a17baf · outbound

This paper cites Cross-task weakly supervised learning from instructional videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Cross-task weakly supervised learning from instructional videos

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.774187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.774187Z digest=sha256:8ccba499a2fcff5101d8fa449d39f7cb556809f083d0087a7460ef52cbdd9c9d

Observation bfcfd53b-1f87-4dc8-b886-63934c8076a9 · outbound

This paper cites GPT-4 Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models GPT-4 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.779074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.779074Z digest=sha256:6f53da5d70f23d39ba91fa4b0dc519bc51da5b044175d5a938dbd5f70f6f6e73

Observation 1d7f9d09-8164-4e31-b0ff-cf7cb4912456 · outbound

This paper cites Procedure planning in instructional videos.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Procedure planning in instructional videos

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.862986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.783601Z digest=sha256:6cad0a5ab2c6f5448115b9b206f60682ba150ffc8900f715d852bb66264af969

Observation 45816025-77a8-4996-8343-f5f185a6bce3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.788339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.788339Z digest=sha256:2b40f65ee7f5440abb0d4a6fe46e10d152d0a90a23ca73f72514caee1d48917c

Observation 6b696177-b007-40e3-a160-eb62669a7c28 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.793129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.793129Z digest=sha256:fd7cca3ee915cca2d4d114a9792a6527d655cd037b36ba8e0e1197221d583de1

Observation ca8f6bae-b1e5-4f52-8979-691d0a85fb06 · outbound

This paper cites The Llama 3 Herd of Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models The Llama 3 Herd of Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.797702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.797702Z digest=sha256:31993d3d97cfefe4dd5a2de0eea5b9ad4b873239c60ab7cb9a19700e8efd2bac

Observation 577db0fc-0d97-4733-8a58-afaccc14980c · outbound

This paper cites Identifying and mitigating vulnerabilities in llm-integrated applications.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Identifying and mitigating vulnerabilities in llm-integrated applications

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.847612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.803407Z digest=sha256:aabad0a9fe4cd8840e9dc37df2ef07dd01cd914d2893eec40925ec71c46c476f

Observation cf3fe687-b0a0-444a-87ea-e87f80ecc16e · outbound

This paper cites Qwen Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.808139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.808139Z digest=sha256:a2b03f43f72634448e85db1e290aace6d6299e7be4016eb6188dd8f54c9d1cd6

Observation 047be0b7-5194-461e-9856-6be2066e33e2 · outbound

This paper cites Qwen2.5 Technical Report.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Qwen2.5 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.812991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.812991Z digest=sha256:d906cd2d0681ccc10e17279aeec07bb0dc6481c76b1229b291e39acb5a7da101

Observation de0015e8-0aeb-4867-97b3-f8b93ab7d2bc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.832119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.817963Z digest=sha256:24124862e38e856d8ba5e5cc8995339e395dc2e7427df98ca9ae1bdad143603e

Observation 5ef27645-4dab-43fc-85a8-edf929a3dbc1 · outbound

This paper cites Improved baselines with visual instruction tuning.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Improved baselines with visual instruction tuning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.815963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.822529Z digest=sha256:ae3a6774e95c6c86a5eb8e0a733256438c6ec2fc42abe4856f396cfe2cb1e913

Observation faf3c1a1-325b-4625-ae12-0e58912ab2c3 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.827196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.827196Z digest=sha256:3dc3c21176cf71178ff7f2d15a2791f99ee1cf06f91e42a33cb51bb60ba5af78

Observation 53afbc61-2f7a-425b-a164-06134bde3f79 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Unmasked teacher: Towards training-efficient video foundation models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:59:34.799739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.833255Z digest=sha256:86730a4bcec51cf2ea6f8f69e09010f92a491e294a0dab352bd9134156fda7db

Observation 443f75ea-042d-40ab-bbbf-82c5d5f71f07 · outbound

This paper cites Self-alignment of large video language models with refined regularized preference optimization.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Self-alignment of large video language models with refined regularized preference optimization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.838263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.838263Z digest=sha256:2cad61d7acf926150eb564e605b2aa3647d68c3437d13caa0b50bcd9db0f4781

Observation e95008de-bb64-4ce2-8fda-886c9f83364d · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models NVILA: Efficient Frontier Visual Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.842742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.842742Z digest=sha256:03604abb489f241cb24e890d9425d8cc9e3e591f45e6739a2b8c2656237bb67d

Observation 7df600f0-39b1-42a5-8f74-b0b2ab60c761 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.848016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.848016Z digest=sha256:56f64adbe4aab5d8fca1c9285f19532a02e3a95d1cbb98e32b159926c6446f67

Observation c2d2c112-4f3f-477f-9dd0-beaee1a5d3d4 · outbound

This paper cites GPT-4o System Card.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models GPT-4o System Card

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.853704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.853704Z digest=sha256:394c81f08eb05a3362c962bdc0ff9811fc11683417a635e6befc22418a32dac0

Observation f3caf5fc-588c-41b9-9866-da8cedf6f0d6 · outbound

This paper cites an unresolved cited work.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T21:59:35.100239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:59:33.592973Z digest=sha256:76b728c48c016e1c16edd40d77a545d97a9bd568b0f0519430edefa910333330

Pith citing papers

No inbound Pith citation observations are available.