Pith. sign in

Paper Citation Record · LEDGER

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

As of 9 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2509.08538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08538 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:32:57.059594Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved87
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e356cbf-7a71-4122-b51a-a651befb776a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.355191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.355191Z digest=sha256:58fd10055dd414e3861b9e54eee2bf1e1d35ef2b1484e858a2e4e8a7e19ef7d9

Observation bc4d416f-196f-496d-ac35-4956244fe320 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.362792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.362792Z digest=sha256:d71105c26100b80ca9cba65e18a9610398eb8f6b6ca2074fe036d3f3af490cff

Observation 7bbc4ae2-2b48-449f-8b42-69a845ab6de4 · outbound

This paper cites Qwen Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.368080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.368080Z digest=sha256:0f5fe76e58afc6f88f981560f068aa561bd44efc2972127cabfc21448abc842c

Observation 506debf6-d33e-4759-ad25-68f24fa4c7d3 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.373710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.373710Z digest=sha256:5590016946d6a6c7ab2253da9362d925d7ec73130c077e208ee467743706635f

Observation e9e3d4af-39a4-47d7-803e-99950a1b1502 · outbound

This paper cites 2020.The visual story: Creating the visual structure of film, TV, and digital media.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2020.The visual story: Creating the visual structure of film, TV, and digital media

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.385920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.385920Z digest=sha256:fc25b3d7f8d6efec8c42b85e5c6c5a9029924c9d325f3cac0c6bf9e4caa7ecf9

Observation 6c802f18-24d5-4edf-b6d1-356ca4c3cd7e · outbound

This paper cites 2005.Figures traced in light: On cinematic staging.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2005.Figures traced in light: On cinematic staging

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.393082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.393082Z digest=sha256:696a2b9f243a9defbd1bcd910636bfaee09bfeb380ce3badc602651e4c26b41c

Observation 7c56042b-1401-4670-9243-4994f562161f · outbound

This paper cites Bordwell and K.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Bordwell and K

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.399095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.399095Z digest=sha256:f5026ae525d2f99d132debc0563b5bd75d40a403b8748070c292b8ee408dafc9

Observation 48aedb5d-29d7-4710-be97-319161b2aa83 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.405345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.405345Z digest=sha256:694b83068ae4123e639392aeab921cb9b7eb6554655a882835453217108dfba6

Observation cd5213d4-144f-4941-a1e6-c02f6c688d28 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.413191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.413191Z digest=sha256:b04778a1ad7fd310f63e5ca7b241e4bd9f07e621e84308f30ca063d5f8e10f63

Observation b5c93435-f166-4389-9de1-3406a681970f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.424333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.424333Z digest=sha256:07d4e129e3389c6fe109dfa2532fc84278d51c44439dd0dd1dddb2c73afb114d

Observation 163f588a-820f-4e44-96dc-363e3f0a1951 · outbound

This paper cites InternLM2 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models InternLM2 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.418696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.418696Z digest=sha256:55ca8547eae2a3243c3d72c5ec7771606501140158b4e90e2852db57a50ea535

Observation a75f2100-c4ac-4857-8313-27be9a211d3a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.437988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.437988Z digest=sha256:19d813fb14e057f322d5e03416d50e0af836a2910d7c8dcd3393394bface9b6e

Observation 89a7aa44-fea8-48ec-afbb-113bec72c9a3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.430841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.430841Z digest=sha256:81970b54eefed37bf4f46d1c16416cbe63ba3079a28d941e8814e519f2675ed9

Observation c689760a-0b09-4ad3-88ed-e4cb4f398809 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gonzalez, Ion Stoica, and Eric P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.449752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.449752Z digest=sha256:fe7b18ca7e6edcd73502652c85cb6a9318358af2eb4ab58980e06a886bdb9033

Observation 129077e5-ac55-4b2d-8b39-f1a3bbc40f0a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.443650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.443650Z digest=sha256:c864b36b74e10978ed31250748f6b22e2dded901f2cbc17de2f01be904786e0e

Observation 55d36d7d-628d-4b31-8ccd-196bb6d8b1d3 · outbound

This paper cites 2013.Human information processing: Vision, memory, and attention.American Psychological Association.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2013.Human information processing: Vision, memory, and attention.American Psychological Association

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.461933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.461933Z digest=sha256:6d50cb7779692f394531b53ea4f93d41df056bdda67cafc0984f88f7dfcbc6bc

Observation 1ff7a2c9-2aae-46fc-a09c-759ecf7c37ab · outbound

This paper cites VidHal: Benchmarking Temporal Hallucinations in Vision LLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VidHal: Benchmarking Temporal Hallucinations in Vision LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.455277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.455277Z digest=sha256:7a094f8d018f368f314cc11179bdb998d16c3b07d5fc0fef7d41a45668322816

Observation 01e836b6-5343-49bf-8ddc-8e925eb88abb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.471710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.471710Z digest=sha256:fc8e0b7ee8b0c1c7b494406d06f6fdb89845e4051b4b9996c194eaf4617eeaef

Observation c3ba4ab5-7b46-4cf8-a28f-7f50626fb6bf · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.482556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.482556Z digest=sha256:04ce2e844ef4791efcfb0888a78185daa499151bcf36d5e763235d458894d540

Observation 0b62f425-e764-4385-a3c0-7c7e35137384 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.477678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.477678Z digest=sha256:f54172c157ff3e871608a5def82e4c29d24565b32ca1314bdda570d526e7fca1

Observation 5a94453e-fb10-4388-a852-2c8bff38d7bd · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.493996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.493996Z digest=sha256:ad4c4f9a9ca9cae2a11443c16a8ad8803fbddc1b37092e991f43e989773f0a82

Observation 9b19fdc3-d003-42a6-b2f2-2260e3534d61 · outbound

This paper cites DeepSeek-V3 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models DeepSeek-V3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.488472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.488472Z digest=sha256:6462a58c654889b7d23cb7bf3aae54a8788e97a7fba80f5b0af1fdbba1aba656

Observation 01668f78-84b6-4797-9512-062dda1e2696 · outbound

This paper cites 2002.Mise-en-scène: Film style and interpre- tation.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2002.Mise-en-scène: Film style and interpre- tation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.505747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.505747Z digest=sha256:9a082eaeed14148abef0ae3f4e6d911e8af070556cc30df1284873f267ec1c76

Observation bac15c89-d678-4e26-aca4-202fc2396257 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.499149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.499149Z digest=sha256:ba45a358f5041a8121298e61cff26a07a031fc8182d198941bee5365c02296e3

Observation bdf5e9e9-13c3-476d-bb22-a90c3b2bf7e8 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models ImageBind-LLM: Multi-modality Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.516191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.516191Z digest=sha256:f91bdb52485989fd4e64daff2a77db58b7c32f42c009c5f7eeebe2e84741ba17

Observation fde065e8-3b2e-42e2-93f2-e98e9cf51af7 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.523017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.523017Z digest=sha256:ca37cb19f535bbd2953f7e1173bfd0414d4d95bbedd4f60241b8573485f2fd01

Observation 985e3870-8e8e-4c16-9e1b-8bf5a67d78e6 · outbound

This paper cites GPT-4o System Card.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.537462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.537462Z digest=sha256:2888316287556061ee4587e7ef3b4b9e65e2a2228263fc08b78728cfadd731e6

Observation 379ccc35-1a65-4357-98b5-e12bd82b8bac · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.532290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.532290Z digest=sha256:bbbe2e7616f7918e22ebae09457e911e84dd8ca86a3e0185e282f1fccca504e0

Observation 87382e3c-6127-4ec1-a8b8-a85f99667977 · outbound

This paper cites A Comprehensive Survey on Visual Question Answering Datasets and Algorithms.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T20:33:54.768847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:32:56.548575Z digest=sha256:3b91927e5f2ef5cfd16deb948fae990aad24ff75dcde8e3271cc9031297d1244

Observation b7e168e3-710a-4620-8786-2f19419ce369 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.543391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.543391Z digest=sha256:de394d4e0cc220d465a96f7a41169841a19aa6420ba534a6b4a2df18fd6332b3

Observation 26f2bd1a-148b-4313-acf8-1c91b77cd277 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.563400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.563400Z digest=sha256:7fbc5e480510c61841ee1af668b86bfb0c97e7198795c30b077bdaedf9672df6

Observation a6efcdce-56d0-46e0-ba0b-dbc53ae73aa1 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.554860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.554860Z digest=sha256:80fc84610231fda138a9c9d2032728bf6ab1be2875968d68186b90f6d18d34e8

Observation d8d35c71-7537-4e24-ad74-7058134d5267 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.574322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.574322Z digest=sha256:4ff89ec390b802ae6a7fb5519f1c4db141eea86f1f4609009e1849a377725104

Observation 0c100652-0ca4-4266-9173-dd453fb46011 · outbound

This paper cites Berg, and Mohit Bansal.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Berg, and Mohit Bansal

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:32:56.567981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.567981Z digest=sha256:3c8a680122f409df8a6c1cee0c02cc5f077cb56f32a97f1fc17a34e1b3c78994

Observation a133ce72-f350-4e46-b0ab-30c9f0976478 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.586319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.586319Z digest=sha256:570328d3a6b9e2444bc59bb1153eb17b3cc1f280e1505511d31084660de7e83c

Observation 6536a079-2501-4bae-aa6f-40347e8612a4 · outbound

This paper cites VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.580184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.580184Z digest=sha256:3cc6128dd14f9128678def33f7ed108d42c55d026e6276fe73bc33f6f1e698bf

Observation e5f377ff-4841-4fa9-a423-5bb7e32e9fdc · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.597946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.597946Z digest=sha256:6f0440606c0388346041ad690b7b07626b2aac282fe3a246cb88b6faeee02722

Observation 46b695c2-4d8a-437a-b8c4-2b14d639210a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.609106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.609106Z digest=sha256:26732aa296d9e53605bfb8b7af9859ba4190f71b37c11c2e70ecd1f96231caef

Observation ad4666fc-3b11-4d3a-9e02-7d8192d1a559 · outbound

This paper cites Making Long-Context Language Models Better Multi-Hop Reasoners.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Making Long-Context Language Models Better Multi-Hop Reasoners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.603444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.603444Z digest=sha256:bdcab2781b0e4da19c8a3f5f414946f905b8f388f9f189959ae8f697ec003e73

Observation 1354d313-04a2-4e86-a48a-d63ce58b7a6e · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VILA: On Pre-training for Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.622377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.622377Z digest=sha256:a72cd38331c52dc1339d35bbd299552ecc2675f226377e3873f86a1d83636aa6

Observation bb973a2c-59d6-40c4-ace3-c32c41566e5f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.615854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.615854Z digest=sha256:3b6c42269a124ca36711befd5763fd32ff6657b7529a0a41588294f01d2e7179

Observation d376f94d-7f43-4ee1-ae00-3a6a48c7a26b · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.638414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.638414Z digest=sha256:88ea387f8f0c3bc8a0fa3ee61066bd882b2a8a65ffff105359257f770ade2dc2

Observation 495627fe-3d36-4200-a04e-5b3679e4326d · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.628654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.628654Z digest=sha256:f60634eed196618f9ba40fd71a92fd00c55829f08e0da020322f8dc7277cf498

Observation 06da0f1b-aa4f-42fb-8eea-179b37a4ad6f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.656822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.656822Z digest=sha256:15d5b0edad6171a9a763f8e834f09f1b635ccf074360ea1b162b9741b1d600ab

Observation f6b3f7e5-12bb-4488-a772-6471d83c52a9 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.649625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.649625Z digest=sha256:859af82089b3b1015957e4fff2e7247ea6871b69927dcb9854dc950b29cb20dc

Observation 97bd2deb-23c0-4f14-9651-fec031ddc604 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.670143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.670143Z digest=sha256:0f3775798b941f8dfb6bde1debac89373d521075872d970004c36f43414c61d2

Observation 8bf8dc53-b87c-4285-bdcf-5ad10395a3ea · outbound

This paper cites PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.662679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.662679Z digest=sha256:b5adb776b2532388fb332f0a5f61f57af8906919bc917e1ed3a3bf3631430dc6

Observation 58f373fb-7795-410a-b6db-af1f1b950951 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.689614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.689614Z digest=sha256:8c86bd5ccea4225f6119a2d41ab2ad8a18948033092d221bc764bbb4e6fd201d

Observation fd07ad05-d3d8-4c9c-8428-c674b264f753 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.676255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.676255Z digest=sha256:eef9c48b1540d768703077f03d2e5c7b8fe86871760c677aeb0b92abe78c784f

Observation 8291684d-8999-4361-96c2-dda099eae3b5 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.684255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.684255Z digest=sha256:d9e45a0fba92af4b5c6a81e7a6c83ae5abbe43add2a8ba132355b5bb57dde8ab

Observation 25003f71-9556-4097-860a-2ba84db3ad0d · outbound

This paper cites O’Connor.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models O’Connor

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.708218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.708218Z digest=sha256:22eea33ab62213b108dd25bf2ed2a6768455f7a409d885d06ae4326696aca869

Observation 0ee60542-d7c1-4097-af91-440d060e6b87 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Foundation Models for Video Understanding: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.695590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.695590Z digest=sha256:56ba1c4d2223b1fed17a56ec9df1390a421ca33867dbbd96293ff98170dedaf4

Observation 2f6e6c7d-485e-4426-acb8-9bdb093bfc71 · outbound

This paper cites 2016.From Human Attention to Computational Attention.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2016.From Human Attention to Computational Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.701856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.701856Z digest=sha256:fb3188476c1ba8ccc7099ab53bc34e195217a70945c097ab3ca654a8e891ac55

Observation 18ffc1c8-dede-49d3-9eb5-78e45b38b78a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 58

Resolution
verified exact
doi, observed 2026-08-04T20:33:54.420702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:32:56.726333Z digest=sha256:3bc24533eb37a89386810f19f8de0dc5dc20c2e0c2bbc41e8d4ddcaca793eff0

Observation a9428a4c-7806-414d-af13-5fe53049cac8 · outbound

This paper cites 1999.Foundations of statistical natural language processing.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 1999.Foundations of statistical natural language processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.714924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.714924Z digest=sha256:81ccc7c3d907622c46e366e4c4e33371e074c06c51bd0a525830cc09973e2646

Observation 0caae250-134e-44c9-b60d-100de4137922 · outbound

This paper cites Large Language Models: A Survey.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Large Language Models: A Survey

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.719979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.719979Z digest=sha256:64bd6fb2158073e22626ba88cc693665e913c456a64e5ed9f52709b2aa695af6

Observation 3e437a63-3153-4989-8783-835832e93c72 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.753734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.753734Z digest=sha256:dadb5180ac9b76cdd3136cc9657e532977483fccb8515cb2f137589c1b305d47

Observation 21a2c4cd-3046-4363-94bf-87e42334fa85 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.736581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.736581Z digest=sha256:fcd610f98034b6c0eff11aabe7f292bf4c198e0c937444913a7022ff5b979f82

Observation cf239a5f-b4c5-488b-84dd-7fca0def9139 · outbound

This paper cites GPT-4 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.744472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.744472Z digest=sha256:17bf1d77adc57f0ba2e58d48e3f1bc253f5210561a17040815c7a6f8e5596cfd

Observation 4fad3c1e-f926-404a-96d2-75968684131d · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.787102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.787102Z digest=sha256:ea0cc2bab825bf3ba68a3763e6b8fe572beee8745b6eb2c37fda70b89e713df8

Observation d2e64a15-bf18-4644-86c8-68f0a76d935e · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.761348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.761348Z digest=sha256:42d235b32ead663495689e982472a007edffb21771f48fd297e24f9894d55400

Observation 52cab061-e964-47fb-af3e-dff4cd902496 · outbound

This paper cites Girshick, and Ali Farhadi.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Girshick, and Ali Farhadi

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.771366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.771366Z digest=sha256:94a0de4b4a53ba0375410c35332b0f1994aff004170f3eee10ba8edc52b9cbf3

Observation 48ab59c9-9d8a-4f19-843d-204ab322525c · outbound

This paper cites VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:33:54.254429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T20:32:56.811561Z digest=sha256:d9acca2b6e3f3f87f6fc5a4f0da9433b2d421cc0da41d055e634d25b95f7ee00

Observation b7efc35b-e33d-4610-9fea-8e2297ab618c · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.818052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.818052Z digest=sha256:c95cd3885cb1ab2e3f0b70e9656b3d8bd36bc22e0597e4d3e8305c3bb51981f5

Observation 8502143b-8ff9-4926-ac1c-d057501e47a2 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.791245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.791245Z digest=sha256:4740636f0d82b3b6049b8bd8a95250b3ea007e58377029dcc5a1d5927f61970d

Observation cd7a7a2f-5127-4713-9ec0-14b0dc18d603 · outbound

This paper cites A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.800975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.800975Z digest=sha256:b1a524bbd11d11686992cf01d3833653f1ed5502b9ffddbaf37fb01d165224f2

Observation ba898d62-0a29-4e56-8bf4-2fa5ddf24e8e · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.860424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.860424Z digest=sha256:6db8abc343d04cba0ae67b9d85631e671cef0fb51b533b1e2fe0058ee6181301

Observation 2a774b62-1def-48a4-8bf3-6a5b37fdb8cb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.881841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.881841Z digest=sha256:25e7162adb6371e6faf77cde6467a5cf012cd493da8c44e83fb1a12b7cbb18b6

Observation fd36a8e4-b0fc-46e9-98bf-99b284787fc0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.838540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.838540Z digest=sha256:416f81b5c47946e5a30a4bbe8317be5e7175f66943c67a849631c07de128c0c5

Observation 18729337-7b79-43b7-ae5b-c8c879c8914c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaMA: Open and Efficient Foundation Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.919983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.919983Z digest=sha256:a6a2d415010f0f633dc87c8c78041953795ee2d711e8bbed878cc54fc68dd7c3

Observation a0306755-04af-44b5-b69f-1995e547ec4a · outbound

This paper cites 2013.The Oxford handbook of sound and image in digital media.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2013.The Oxford handbook of sound and image in digital media

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.929596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.929596Z digest=sha256:47003bae6fb7922ccfe6a2d308c6b490a648a6c449b2e7f305c3c8b07ab5c70d

Observation a0173072-d145-4abc-903e-ca8a286a69e4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.938002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.938002Z digest=sha256:e1a45824b157e4c021b7d352e8d5c4e8473be36604037ed7708cf948d62270ed

Observation 0c613ce5-c7c3-4050-9a71-e8656ada9263 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.907053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.907053Z digest=sha256:03394545b120d1c3461b7f57a9b87de154cb881b312d46703a61761852f44131

Observation fd98d0cf-6b39-4998-a054-ef3e5fcfe619 · outbound

This paper cites 2009.Multimodal Sig- nal Processing: Theory and applications for human-computer interaction.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2009.Multimodal Sig- nal Processing: Theory and applications for human-computer interaction

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.914794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.914794Z digest=sha256:3b101c7475e6968605678840127126395dfbe8de1f9cbeebe5ec2c576b23fb09

Observation a55a2557-5730-4d44-8d43-a7dfa88ea0b1 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.967932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.967932Z digest=sha256:6886df8246f912856bfc6a1d39526fcdf91e3bbd0a1b56775accc642988124ab

Observation a1cf3dde-b5ac-47fb-af1c-7cd6dcf65ebe · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Vript: A Video Is Worth Thousands of Words

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.975944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.975944Z digest=sha256:a79dc6e501d09e81e32c74873af1da3b438056348b76bb42754591217cbbd110

Observation b4a199f5-70ff-4355-8060-0d6ce4e43a0d · outbound

This paper cites A Survey on Multimodal Large Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Survey on Multimodal Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.981480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.981480Z digest=sha256:5e9592c4a6edb719601458325737551660900cf851f971cfde503978c31da725

Observation fbe2bbd1-7f04-4423-aaf0-1df9d433dd98 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.945883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.945883Z digest=sha256:9e45d7afbedc7d7441197470e79546fdf234f2a02f743e332af89af2a2bfbef1

Observation 14760259-ed39-4b68-b20a-974d045ae3cb · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.952259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.952259Z digest=sha256:b9966cf4884450049f021aae1f53476236a432e5a0abdbf318e9453b09bc9369

Observation 13186908-81bf-4add-9d15-bf2d0dd2677d · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.960144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.960144Z digest=sha256:aaa358875a79c9e4f0c59b66b09d2bfbced2f618a54b8d640283d04db913f817

Observation cc89babc-b975-4d58-9f7b-b9b38196ef01 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.016518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.016518Z digest=sha256:28cb8572eb410edb4abb5f94a3888f2dd382549d06a0eb7e2df906bb9582b247

Observation 96c2936d-b272-44d5-a83f-ba6c9623791d · outbound

This paper cites DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.027292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.027292Z digest=sha256:e93cfe969cbff2265ad4d11d0b56c6fc9410457ce937b55485a7c942ee89b697

Observation 9f0fd19d-6798-437b-a629-a7d532ad0a5c · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.038704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.038704Z digest=sha256:fd3c048933488cc0ce2f71fea185eb75a42c93f77253f09bb9af7e6da449714e

Observation 288f0ac9-7ebe-4473-9e8d-6d6f16129ce4 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.987009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.987009Z digest=sha256:a15f58e8234278830a5713c7155ccd1f1c03a97b4e6ec9d4e06d3c8f5ee22c32

Observation 25887e44-7304-4c90-8082-a4d8b91cef5f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.994039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.994039Z digest=sha256:43df743e7e4eff7c19ffe2834cf3aed9287c3d33e4ad30bc3659da5d63bc3cb4

Observation 3b0cb272-5819-462f-a3d7-d55ee8d43571 · outbound

This paper cites arXiv:2409.16597 doi:10.48550/ARXIV.2409.16597.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models arXiv:2409.16597 doi:10.48550/ARXIV.2409.16597

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.000075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.000075Z digest=sha256:1138f77fb7fdd912687b120277261b5a8e6f10f1c271dacd508b710a6cc38976

Observation 78f725de-5138-4e5d-8427-d6f0628d5ddb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.007950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.007950Z digest=sha256:641334b2bfe15852580c60ec09244278e8c47354f99c20189f49f36787c2dd8c

Observation 2999ef6d-299e-4698-8d5e-77d45e0f5f99 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.050146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.050146Z digest=sha256:567d73bfac7f8c843e8bad63699cd938569db02f5e98a1a839647bb8f2e8ab77

Observation fea97aaa-4d6d-47aa-b7ce-35c4e290a4f6 · outbound

This paper cites (00:38 – 00:46) (CIV) Q: Which individual was the one who spoke in the video? Options: A: A man in a white vest was speaking in the video.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models (00:38 – 00:46) (CIV) Q: Which individual was the one who spoke in the video? Options: A: A man in a white vest was speaking in the video

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.059594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.059594Z digest=sha256:2da32182102e7264b22670b67ed91bbd91b90f8a3d25868cd83f4129023d6c06

Observation 7303eb7a-f8ea-4545-a508-6cda5b5daacb · outbound

This paper cites In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016

Reference 2016

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:32:56.778550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.778550Z digest=sha256:7262c73c1ef34972e065a5ee7ca927dc9c7e412bba1832ab258a59dc2ddb58c4

Observation 7c1dc52b-a6f6-401a-90ea-ad035364cf8f · outbound

This paper cites arXiv:2312.17432 doi:10.48550/ARXIV.2312.17432.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models arXiv:2312.17432 doi:10.48550/ARXIV.2312.17432

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.895347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.895347Z digest=sha256:7cdc1e815504021a1ae19742a424952488ad881222ed5464136a42e4f3d9d639

Observation 1f7800d9-7adc-415d-b298-9ecdd297d6bf · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.379260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.379260Z digest=sha256:f964d8b2c6fe72cf68d48a96b26bf0f5cb60358fc2df2d5331d28a12ad091fc5

Pith citing papers

No inbound Pith citation observations are available.