Pith. sign in

Paper Citation Record · LEDGER

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

As of 17 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2509.08538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08538 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:32:57.059594Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved87
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e356cbf-7a71-4122-b51a-a651befb776a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.355191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.355191Z digest=sha256:66ab2528c40a8023c2253da889ea3981ca9eeec9898891c41efc50daa49c54c8

Observation bc4d416f-196f-496d-ac35-4956244fe320 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.362792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.362792Z digest=sha256:efac31fca3f7e07f87ac33f2cd138cb38e9c3056a2909308ff92270202e4877f

Observation 7bbc4ae2-2b48-449f-8b42-69a845ab6de4 · outbound

This paper cites Qwen Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.368080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.368080Z digest=sha256:ec23cfddbc809d63dc64088eaadc46ef39c24fff3fb1ab3ceb29a261da97f187

Observation 506debf6-d33e-4759-ad25-68f24fa4c7d3 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.373710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.373710Z digest=sha256:ff8447cee6afbecd5a15217ca83d73839aaa12822187a625cc2ccdb220b32d48

Observation e9e3d4af-39a4-47d7-803e-99950a1b1502 · outbound

This paper cites 2020.The visual story: Creating the visual structure of film, TV, and digital media.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2020.The visual story: Creating the visual structure of film, TV, and digital media

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.385920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.385920Z digest=sha256:f3f479f950a744fa80a651399b6353cb57a33010ca24072739073c8cf7620c2b

Observation 6c802f18-24d5-4edf-b6d1-356ca4c3cd7e · outbound

This paper cites 2005.Figures traced in light: On cinematic staging.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2005.Figures traced in light: On cinematic staging

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.393082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.393082Z digest=sha256:1b136fc720eeadce4ee1420e835ffa630e03b1070a91659f211c0780aee30f48

Observation 7c56042b-1401-4670-9243-4994f562161f · outbound

This paper cites Bordwell and K.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Bordwell and K

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.399095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.399095Z digest=sha256:30857454541b6740d02272c01caa6e10d447aab5d0743f6216908b50e4e1b2eb

Observation 48aedb5d-29d7-4710-be97-319161b2aa83 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.405345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.405345Z digest=sha256:b89933d91861158a4b21e26aabfff3948b6a115ff00d0cd562e44187a64e977d

Observation cd5213d4-144f-4941-a1e6-c02f6c688d28 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.413191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.413191Z digest=sha256:f0ba4bd39622007911fbc06771e2906b196342b81c7ccab27898804594dc6714

Observation b5c93435-f166-4389-9de1-3406a681970f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.424333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.424333Z digest=sha256:3e67f2b86d4cfc3daede26b6ef680f2953ea297a96bb7837e54582dc1be443aa

Observation 163f588a-820f-4e44-96dc-363e3f0a1951 · outbound

This paper cites InternLM2 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models InternLM2 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.418696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.418696Z digest=sha256:68814aa09a4560ce96d2996104d80f2bb65f29729de8299374bc76fdb2dc03d9

Observation a75f2100-c4ac-4857-8313-27be9a211d3a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.437988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.437988Z digest=sha256:2db492e21db9329d530a5f09f0906a2dfd04dfee463924506f655c861376d801

Observation 89a7aa44-fea8-48ec-afbb-113bec72c9a3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.430841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.430841Z digest=sha256:6f454f556bea432ae5ed13b21229f9dade37f988b8f5bda3ceb1e408131f13fd

Observation c689760a-0b09-4ad3-88ed-e4cb4f398809 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gonzalez, Ion Stoica, and Eric P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.449752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.449752Z digest=sha256:683bee2b1435f0d66c8b349bf18d01ae4a9af3251e21eea7c80148d100effaed

Observation 129077e5-ac55-4b2d-8b39-f1a3bbc40f0a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.443650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.443650Z digest=sha256:139629b9d6279efa1ec0d7e5d919432d6c1ee7ff05ec549d30b24b86d70d75ec

Observation 55d36d7d-628d-4b31-8ccd-196bb6d8b1d3 · outbound

This paper cites 2013.Human information processing: Vision, memory, and attention.American Psychological Association.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2013.Human information processing: Vision, memory, and attention.American Psychological Association

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.461933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.461933Z digest=sha256:89fb1a6b25967171dc0e2c04ec3756bb543e7abeee75a4fd37384a68dab964c0

Observation 1ff7a2c9-2aae-46fc-a09c-759ecf7c37ab · outbound

This paper cites VidHal: Benchmarking Temporal Hallucinations in Vision LLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VidHal: Benchmarking Temporal Hallucinations in Vision LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.455277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.455277Z digest=sha256:2cade410d6a7d8b1ef7f2c42f37037d3c206182bb69f649c69f38786dd98c51e

Observation 01e836b6-5343-49bf-8ddc-8e925eb88abb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.471710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.471710Z digest=sha256:7d0fa68a667115b3aef565d421e915e6bfc392eb2cba3d5b691a8cf77e146c67

Observation c3ba4ab5-7b46-4cf8-a28f-7f50626fb6bf · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.482556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.482556Z digest=sha256:75fd834e30c5f85f9074b6593a24959a91923a1853e6ea0968f5d08b86c91068

Observation 0b62f425-e764-4385-a3c0-7c7e35137384 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.477678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.477678Z digest=sha256:66a16dcdded32a0fb0f841a93f7d83ac44813469f49d6ff3b31c9112d531cfbe

Observation 5a94453e-fb10-4388-a852-2c8bff38d7bd · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.493996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.493996Z digest=sha256:215b87fd7669af57be773b9ca869a06c952358c541fb3ef407a98327a49f618c

Observation 9b19fdc3-d003-42a6-b2f2-2260e3534d61 · outbound

This paper cites DeepSeek-V3 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models DeepSeek-V3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.488472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.488472Z digest=sha256:a81c30ef67f90149bc2bf3a6e7a8d23d40d5e09e268e335b674ecfff08e06d81

Observation 01668f78-84b6-4797-9512-062dda1e2696 · outbound

This paper cites 2002.Mise-en-scène: Film style and interpre- tation.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2002.Mise-en-scène: Film style and interpre- tation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.505747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.505747Z digest=sha256:4d0ca288149037b5c514060db8dd570c4448fd14a0783f4997a4b6226d57d151

Observation bac15c89-d678-4e26-aca4-202fc2396257 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.499149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.499149Z digest=sha256:a2544cb2f2b4648285f8f4fbba4cb2c321714e1577a7630703f8abacc016bcc5

Observation bdf5e9e9-13c3-476d-bb22-a90c3b2bf7e8 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models ImageBind-LLM: Multi-modality Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.516191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.516191Z digest=sha256:2a7f0d9f24c7c586c94b9a34796694a057d491df7c53a4219c027d7ac559aa28

Observation fde065e8-3b2e-42e2-93f2-e98e9cf51af7 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.523017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.523017Z digest=sha256:a4445b801b919726092b3e1df52d7876b4373823e4d6deecde3aed45cf7c7851

Observation 985e3870-8e8e-4c16-9e1b-8bf5a67d78e6 · outbound

This paper cites GPT-4o System Card.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.537462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.537462Z digest=sha256:161a3e1dfd6f19ca6e3e96c37bf80fdefe9bbc0c9d255742c57b062f32f6e5f8

Observation 379ccc35-1a65-4357-98b5-e12bd82b8bac · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.532290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.532290Z digest=sha256:1648b7b56743e9c618e3dacd0923cc8d3ab9a855174ba31231b947a01e2fcc98

Observation 87382e3c-6127-4ec1-a8b8-a85f99667977 · outbound

This paper cites A Comprehensive Survey on Visual Question Answering Datasets and Algorithms.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T20:33:54.768847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T20:32:56.548575Z digest=sha256:3ee369e67b69c2401d0364fbec300c85789ef849066c8dbe3a079e52c3610e09

Observation b7e168e3-710a-4620-8786-2f19419ce369 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.543391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.543391Z digest=sha256:761979c315679175302ef30425be8ca94d643d6de0f61ce2db06a6a4d72fd772

Observation 26f2bd1a-148b-4313-acf8-1c91b77cd277 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.563400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.563400Z digest=sha256:46e0382b80102278a7e03245fda6e3343fb7dd551eaf9ffa9ceaa4014c3ee0fd

Observation a6efcdce-56d0-46e0-ba0b-dbc53ae73aa1 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.554860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.554860Z digest=sha256:e177613032640d57e2162806e6e02cc2efd4c435404f4fdaf8a494c88b155d8a

Observation d8d35c71-7537-4e24-ad74-7058134d5267 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.574322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.574322Z digest=sha256:435f749dea598b9d28009f6ea3072483b90f435df197d16e907cf4095bcf1b00

Observation 0c100652-0ca4-4266-9173-dd453fb46011 · outbound

This paper cites Berg, and Mohit Bansal.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Berg, and Mohit Bansal

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:32:56.567981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.567981Z digest=sha256:878c75ddc8714b85a8ee8be93536adb441623273fd9379a2709563486a493fa2

Observation a133ce72-f350-4e46-b0ab-30c9f0976478 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.586319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.586319Z digest=sha256:76a9e5c1e860d7026d540dd88116fee935f5695790dbe499a4ef0f010651c459

Observation 6536a079-2501-4bae-aa6f-40347e8612a4 · outbound

This paper cites VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.580184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.580184Z digest=sha256:340071f6e654457b31ee6a0d7f212f68b38a32f452efa46d531f8ed76f95ff9f

Observation e5f377ff-4841-4fa9-a423-5bb7e32e9fdc · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.597946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.597946Z digest=sha256:a4734d203d06bf04dc635104bbb7bbbee4a27a4d068677145d9a7d89a726fa1f

Observation 46b695c2-4d8a-437a-b8c4-2b14d639210a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.609106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.609106Z digest=sha256:b458e04c3de225331337bb8d7a0695ffa71b3f99ac93201ca3d0bc0266341629

Observation ad4666fc-3b11-4d3a-9e02-7d8192d1a559 · outbound

This paper cites Making Long-Context Language Models Better Multi-Hop Reasoners.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Making Long-Context Language Models Better Multi-Hop Reasoners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.603444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.603444Z digest=sha256:2304216896f39c9103f6c04e6ce60348875a8a3f1a9ca13a7211317444a54e57

Observation 1354d313-04a2-4e86-a48a-d63ce58b7a6e · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VILA: On Pre-training for Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.622377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.622377Z digest=sha256:a181f0714501b50ad5e718d61ba4a6d9aef73ec108cd5ba7cc8a72b50967a8d5

Observation bb973a2c-59d6-40c4-ace3-c32c41566e5f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.615854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.615854Z digest=sha256:dcaa649b18cbef45fe488dede841554ce22b1adbb0724576de8cce380a33c6ab

Observation d376f94d-7f43-4ee1-ae00-3a6a48c7a26b · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.638414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.638414Z digest=sha256:8cfdb3d970cb0220e7a820af1a33ef0ec64e9387c3ae97fabca351fe984b7847

Observation 495627fe-3d36-4200-a04e-5b3679e4326d · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.628654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.628654Z digest=sha256:df603d65e05520bf20297b3c32637f6dba2b4c533acc967a22c14a8189087248

Observation 06da0f1b-aa4f-42fb-8eea-179b37a4ad6f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.656822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.656822Z digest=sha256:91586d5e08648c8625356cef98a4717ed86aba1a6737a13a6a480700c909a13c

Observation f6b3f7e5-12bb-4488-a772-6471d83c52a9 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.649625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.649625Z digest=sha256:4f49669659f52797cb05500a81b6ec5bef798cff30cb79ef59d3cf52b77bad6a

Observation 97bd2deb-23c0-4f14-9651-fec031ddc604 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.670143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.670143Z digest=sha256:afac8c17b66cce3081f871af2015b12ee8403a8babfcfc15ceb5d8d430a23d6d

Observation 8bf8dc53-b87c-4285-bdcf-5ad10395a3ea · outbound

This paper cites PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.662679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.662679Z digest=sha256:e766fc9e762f81c096bec3324f5c60d0cf37bc67ac67a24f5d48b27ecd4bfe30

Observation 58f373fb-7795-410a-b6db-af1f1b950951 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.689614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.689614Z digest=sha256:18228cc7f262354e01392bc4af0c64233c7aa5e12c2182ef2a6c0550bb7d5f79

Observation fd07ad05-d3d8-4c9c-8428-c674b264f753 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.676255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.676255Z digest=sha256:6f0bc0f0aded3a959a248f79bd9d968767750cea46edb036b03c96c32ce8cab1

Observation 8291684d-8999-4361-96c2-dda099eae3b5 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.684255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.684255Z digest=sha256:9e3c0db7de3b137d82074c3d83ffe511d68b57b00b29c365c8410987f0f241bc

Observation 25003f71-9556-4097-860a-2ba84db3ad0d · outbound

This paper cites O’Connor.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models O’Connor

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.708218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.708218Z digest=sha256:75c2c7361369d1b9647c642898ec0e66bcc22b8ef31bad472008d8fd04f6c733

Observation 0ee60542-d7c1-4097-af91-440d060e6b87 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Foundation Models for Video Understanding: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.695590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.695590Z digest=sha256:dd4442afcfade2e53e656581a4fccf1b975b19e1e0c3170767edab6172df60f8

Observation 2f6e6c7d-485e-4426-acb8-9bdb093bfc71 · outbound

This paper cites 2016.From Human Attention to Computational Attention.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2016.From Human Attention to Computational Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.701856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.701856Z digest=sha256:f8100db47e8f5d0e83b264a7f4945a47f96065b33e78ba68bb04db28c87f0936

Observation 18ffc1c8-dede-49d3-9eb5-78e45b38b78a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 58

Resolution
verified exact
doi, observed 2026-08-04T20:33:54.420702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T20:32:56.726333Z digest=sha256:8c4857c3e2d8ec1d1050c136b702881c1dd0a375314d4099ad6c7f6225d937f7

Observation a9428a4c-7806-414d-af13-5fe53049cac8 · outbound

This paper cites 1999.Foundations of statistical natural language processing.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 1999.Foundations of statistical natural language processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.714924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.714924Z digest=sha256:18818014cb248e08a0079cec674d3e6572e5ea9029982166f08495f190f02e0d

Observation 0caae250-134e-44c9-b60d-100de4137922 · outbound

This paper cites Large Language Models: A Survey.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Large Language Models: A Survey

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.719979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.719979Z digest=sha256:c15f6bcc010b05b6515640ac08cb47b07774f3a52d6a292da431fa9195322ff0

Observation 3e437a63-3153-4989-8783-835832e93c72 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.753734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.753734Z digest=sha256:20d66154537666173617f51f6dae3222f9063296b6c1d885dd430a5df8c320e3

Observation 21a2c4cd-3046-4363-94bf-87e42334fa85 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.736581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.736581Z digest=sha256:d03d7e63b454ea855d1167f220568ef7657acbcf5e4520be51acffa2200dea32

Observation cf239a5f-b4c5-488b-84dd-7fca0def9139 · outbound

This paper cites GPT-4 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.744472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.744472Z digest=sha256:ffbd91929068d1a5a2c976ae27b3b1bc806d688c741e7d84c0f4b3345ba44020

Observation 4fad3c1e-f926-404a-96d2-75968684131d · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.787102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.787102Z digest=sha256:58dc23932e9b950741554bd9bee4c76b0afb8569710f33e31088760da35e063d

Observation d2e64a15-bf18-4644-86c8-68f0a76d935e · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.761348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.761348Z digest=sha256:5232be13ac61afca15452486f8d16c1374d21a509c16690037a6af64487fb6d2

Observation 52cab061-e964-47fb-af3e-dff4cd902496 · outbound

This paper cites Girshick, and Ali Farhadi.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Girshick, and Ali Farhadi

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.771366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.771366Z digest=sha256:4edec4c1c84299579ff1bf3c7dca7585563491f2207ec8e623e93e16ddf48fa1

Observation 48ab59c9-9d8a-4f19-843d-204ab322525c · outbound

This paper cites VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:33:54.254429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T20:32:56.811561Z digest=sha256:ce2847a97d25acde36305377f2e8c36e44ab5a9cc3a25c5e41066fe379961377

Observation b7efc35b-e33d-4610-9fea-8e2297ab618c · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.818052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.818052Z digest=sha256:6e1b1f00a9bd1845cf8ef498b81d6876dd9c058acc8814f4abf71b1ee7af7b54

Observation 8502143b-8ff9-4926-ac1c-d057501e47a2 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.791245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.791245Z digest=sha256:e4bff98fcdeb4ee7307d2477bccc494c903b52b854195591e47448d5dff861ce

Observation cd7a7a2f-5127-4713-9ec0-14b0dc18d603 · outbound

This paper cites A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.800975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.800975Z digest=sha256:1edfaf151ee5ed359d47a6dbb348eb3a7ee84b0edfc43364259696e720a61362

Observation ba898d62-0a29-4e56-8bf4-2fa5ddf24e8e · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.860424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.860424Z digest=sha256:43fee975594dd38721cbb680dce981a5cd8907c707f8ad154458fc5ac0e6ddcd

Observation 2a774b62-1def-48a4-8bf3-6a5b37fdb8cb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.881841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.881841Z digest=sha256:c9e69685669706e24bf52e626dd76530f04cc3d989bbe8db620c222ea721a79f

Observation fd36a8e4-b0fc-46e9-98bf-99b284787fc0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.838540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.838540Z digest=sha256:012aa10f57b7c7533349ad268247d0d15a143ac22fbfe499e3d01c12d043a88b

Observation 18729337-7b79-43b7-ae5b-c8c879c8914c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaMA: Open and Efficient Foundation Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.919983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.919983Z digest=sha256:42f12b26af37340d32dd9279676ca4ed05f7d195793432dc0d7ee6eb3575d9bd

Observation a0306755-04af-44b5-b69f-1995e547ec4a · outbound

This paper cites 2013.The Oxford handbook of sound and image in digital media.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2013.The Oxford handbook of sound and image in digital media

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.929596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.929596Z digest=sha256:aab22d1a727cb1f6126aa6d9c63932f760413531e39b28f45b3c16eb24a5ab4e

Observation a0173072-d145-4abc-903e-ca8a286a69e4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.938002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.938002Z digest=sha256:4592ab14e367c271440ee3fc12327be459bac6190fdac3ce74526a0f835dd6b3

Observation 0c613ce5-c7c3-4050-9a71-e8656ada9263 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.907053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.907053Z digest=sha256:0e609cb806921618f3e8acf5b330f4ddf7892f838fa36ab2ae13b3ae5f964bb7

Observation fd98d0cf-6b39-4998-a054-ef3e5fcfe619 · outbound

This paper cites 2009.Multimodal Sig- nal Processing: Theory and applications for human-computer interaction.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2009.Multimodal Sig- nal Processing: Theory and applications for human-computer interaction

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.914794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.914794Z digest=sha256:b44d2d163f5142a09cfd24c968c2715134962453334a44b6769e65da40d0fa03

Observation a55a2557-5730-4d44-8d43-a7dfa88ea0b1 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.967932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.967932Z digest=sha256:d22c815afcdca1ea7ddb4f85c6d10b5a300c4e10dd83ece130edd257dab04070

Observation a1cf3dde-b5ac-47fb-af1c-7cd6dcf65ebe · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Vript: A Video Is Worth Thousands of Words

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.975944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.975944Z digest=sha256:96aac738948077d5f588166516b130057644d4754f9509d9be1771396a61abda

Observation b4a199f5-70ff-4355-8060-0d6ce4e43a0d · outbound

This paper cites A Survey on Multimodal Large Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Survey on Multimodal Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.981480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.981480Z digest=sha256:354fd6b1c7d618e14e1a4ce3594cd309b2401f0d06c1bc80d4fb8654f53475bc

Observation fbe2bbd1-7f04-4423-aaf0-1df9d433dd98 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.945883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.945883Z digest=sha256:0f7732e72f40f450082efd57f3e375c842444178f77f730808a604741197eb27

Observation 14760259-ed39-4b68-b20a-974d045ae3cb · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.952259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.952259Z digest=sha256:956dc26183c5cb35e6ffea7dfa2deb370d92a9689fbe3d0f1f3d3b65b04ed43b

Observation 13186908-81bf-4add-9d15-bf2d0dd2677d · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.960144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.960144Z digest=sha256:ff34eb8fddda85521d5cbb2ccc3e77602558fe04015f1ba3a8172b7cc18abb0d

Observation cc89babc-b975-4d58-9f7b-b9b38196ef01 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.016518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.016518Z digest=sha256:c375a58ed5ae3266347fac14980aaa71a31e7fe09fd93fa017fbd5e5f7f8c3e3

Observation 96c2936d-b272-44d5-a83f-ba6c9623791d · outbound

This paper cites DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.027292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.027292Z digest=sha256:3e06811c5a6167633c709e3edae591d71efca0707cd09929a1a632ce716ae0d7

Observation 9f0fd19d-6798-437b-a629-a7d532ad0a5c · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.038704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.038704Z digest=sha256:8f36e8d3a4f54eb1fa883ee560a414b86a5a4361773d6a2714c53f6548daa235

Observation 288f0ac9-7ebe-4473-9e8d-6d6f16129ce4 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.987009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.987009Z digest=sha256:fd13fde73d3b54077e6f11c37ec976100286b41a1468d86eca4a21cf2896e274

Observation 25887e44-7304-4c90-8082-a4d8b91cef5f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.994039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.994039Z digest=sha256:d3d256ecd2e46538d1eb3aa1eced52dee973b650682f3bd4351d24b1a63fe52b

Observation 3b0cb272-5819-462f-a3d7-d55ee8d43571 · outbound

This paper cites arXiv:2409.16597 doi:10.48550/ARXIV.2409.16597.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models arXiv:2409.16597 doi:10.48550/ARXIV.2409.16597

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.000075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.000075Z digest=sha256:ae8d0625f0713b4e14efba050593961bee6bdadc656f20c7f5e7feee36f5ad8a

Observation 78f725de-5138-4e5d-8427-d6f0628d5ddb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.007950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.007950Z digest=sha256:12f4f2ae081a2e26442f4aad19661512567636200fca02feb4405cffbc702b2e

Observation 2999ef6d-299e-4698-8d5e-77d45e0f5f99 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.050146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.050146Z digest=sha256:40d4c68a063c628d606e8d0b1eb7f39157c4b75ed30de7ce1742e5ac13adf92e

Observation fea97aaa-4d6d-47aa-b7ce-35c4e290a4f6 · outbound

This paper cites (00:38 – 00:46) (CIV) Q: Which individual was the one who spoke in the video? Options: A: A man in a white vest was speaking in the video.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models (00:38 – 00:46) (CIV) Q: Which individual was the one who spoke in the video? Options: A: A man in a white vest was speaking in the video

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.059594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.059594Z digest=sha256:361f0b2f39a4cc810730e5a40b9e88d44dd7cb3555cc873836a93e6e0107c71d

Observation 7303eb7a-f8ea-4545-a508-6cda5b5daacb · outbound

This paper cites In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016

Reference 2016

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:32:56.778550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.778550Z digest=sha256:f02117fd2e34c778faba8cc382fa63f400ee5d5d53b7f111cbd599c97fa6f73f

Observation 7c1dc52b-a6f6-401a-90ea-ad035364cf8f · outbound

This paper cites arXiv:2312.17432 doi:10.48550/ARXIV.2312.17432.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models arXiv:2312.17432 doi:10.48550/ARXIV.2312.17432

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.895347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.895347Z digest=sha256:c10667cd75d0809c6145b314580157ccd4e6f669cc9bc8c3e5044e666a85597c

Observation 1f7800d9-7adc-415d-b298-9ecdd297d6bf · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.379260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.379260Z digest=sha256:1588db99f03df3c53c256fd39772c53bf0f48456fac92f178033b2faa6d86a97

Pith citing papers

No inbound Pith citation observations are available.