Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T07:12:52.433947Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.11078.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T07:12:52.433947Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a4e19fb7-a16d-4bec-a8db-83079d6338f2 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video In- finibench: A comprehensive benchmark for large multimodal 8 models in very long video understanding
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5289ad-e642-4ab7-8f38-abb0500efc6f · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db61d9e1-24cf-4e61-b9bc-b82472097d3c · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Routledge, 2 edition, 1988
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00b81ec2-478a-4f2e-81e9-cce4e09940af · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2c481b-374d-4079-910a-bf8984b6c114 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c3468da-0e6b-43a7-89d7-c623dca194f4 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e6ea0f-6d94-4544-b843-d8b360fb7079 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video What’s “up” with vision-language models? investigating their struggle with spatial reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a04b7b-c35f-4110-a7e2-7d9fd560e213 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video TVQA+: Spatio-temporal grounding for video question an- swering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40459037-d1c3-4952-b4ea-26fea0891468 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Evaluating object hallucination in large vision-language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa54dc5-274f-465d-ad72-908d6a20d098 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Note on the sampling error of the difference between correlated proportions or percentages.Psychome- trika, 12:153–157, 1947
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dab5a83-5af3-42ec-93f1-f2d889791963 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Plot twist: Multimodal models don’t comprehend simple chart details
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387d0c24-268f-4ff0-b2bd-d84b93400a9e · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Qwen2.5-VL Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b448a5-4725-47db-a3b2-1cd7321fc10a · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Probable inference, the law of succession, and statistical inference.Journal of the American Statistical Association, 22:209–212, 1927
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1814fd4d-d517-4f64-8d66-cdfea2180a12 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video ChartInsights: Evaluating multimodal large language models for low-level chart question answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc7fce04-770a-4527-8d84-c237c0f8e3f0 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Can i trust your answer? visually grounded video question answering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd02eab-610d-404f-a04c-06e910111c30 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video LLaV A- NeXT: A strong zero-shot video understanding model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9bdedd6-e32a-439f-b16f-fa84fdaa75c1 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Large language models are not robust multiple choice selectors
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7188285d-ddb4-48c1-b106-af8b91bd5e50 · outbound
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Xing, Hao Zhang, Joseph E
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.