Pith. sign in

Paper Citation Record · LEDGER

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.11078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.11078 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T07:12:52.433947Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4e19fb7-a16d-4bec-a8db-83079d6338f2 · outbound

This paper cites In- finibench: A comprehensive benchmark for large multimodal 8 models in very long video understanding.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video In- finibench: A comprehensive benchmark for large multimodal 8 models in very long video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:646643ccf0c913d8d759a9d49bd99e1c0f2f5d65860173cf6a3a3308a543995e

Observation 2d5289ad-e642-4ab7-8f38-abb0500efc6f · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:32d33696e71070a3b84c964146f62b703b6e26ab6ce438c52ccfe08c330d45d4

Observation db61d9e1-24cf-4e61-b9bc-b82472097d3c · outbound

This paper cites Routledge, 2 edition, 1988.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Routledge, 2 edition, 1988

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:3659104a24eb1c083a3840fb4384ec51b0df5955832ccc04dae290f852b9ba55

Observation 00b81ec2-478a-4f2e-81e9-cce4e09940af · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:bc39c0be2be726095f3d725303fa3f0e27964db3f087fb466b68de3a6f957664

Observation de2c481b-374d-4079-910a-bf8984b6c114 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:47ef9b44c6229f079dc5c66ae7852f9130cff1d25c30848fc73724c530d543cf

Observation 3c3468da-0e6b-43a7-89d7-c623dca194f4 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:8e1bc242a1f48bda4cb59b8d03621bf1d66497a2244fd71e85d486bc02c28cd5

Observation a1e6ea0f-6d94-4544-b843-d8b360fb7079 · outbound

This paper cites What’s “up” with vision-language models? investigating their struggle with spatial reasoning.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video What’s “up” with vision-language models? investigating their struggle with spatial reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:62489eb690ac8f4a6cdac44f06e3490e8580b49ed8fd927916563a3f3ebc85c3

Observation a8a04b7b-c35f-4110-a7e2-7d9fd560e213 · outbound

This paper cites TVQA+: Spatio-temporal grounding for video question an- swering.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video TVQA+: Spatio-temporal grounding for video question an- swering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:15bd4e4c4bfeed97538388272a1573d0570552ce4fbac24cc6d6a708e48e56c1

Observation 40459037-d1c3-4952-b4ea-26fea0891468 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Evaluating object hallucination in large vision-language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:d0e9d5f42e89d37056abc6b581df0fc930c772e80c68bd0caed70cb82ae83dc5

Observation cfa54dc5-274f-465d-ad72-908d6a20d098 · outbound

This paper cites Note on the sampling error of the difference between correlated proportions or percentages.Psychome- trika, 12:153–157, 1947.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Note on the sampling error of the difference between correlated proportions or percentages.Psychome- trika, 12:153–157, 1947

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:04d88ab5fb8f7a3a42a4551e824e98379333f428e9fae2c1ee369dcb6804de70

Observation 8dab5a83-5af3-42ec-93f1-f2d889791963 · outbound

This paper cites Plot twist: Multimodal models don’t comprehend simple chart details.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Plot twist: Multimodal models don’t comprehend simple chart details

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:9e97573f49b73064143641f63bd41b824636c30585e062f753c707a898c1d156

Observation 387d0c24-268f-4ff0-b2bd-d84b93400a9e · outbound

This paper cites Qwen2.5-VL Technical Report.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Qwen2.5-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:0f315708fa7621e91b6fce9da4ab999285a4b5eeca4d330cd9d47d926acf783f

Observation 09b448a5-4725-47db-a3b2-1cd7321fc10a · outbound

This paper cites Probable inference, the law of succession, and statistical inference.Journal of the American Statistical Association, 22:209–212, 1927.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Probable inference, the law of succession, and statistical inference.Journal of the American Statistical Association, 22:209–212, 1927

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:309d0c896169061531badd4de81ebc5c5e4dddeb10a82a28c07b5e154d83c49a

Observation 1814fd4d-d517-4f64-8d66-cdfea2180a12 · outbound

This paper cites ChartInsights: Evaluating multimodal large language models for low-level chart question answering.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video ChartInsights: Evaluating multimodal large language models for low-level chart question answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:92aaec5034ea9143996f4d0fa346ad5f3f4d303e59fa0c59db5818462f515133

Observation cc7fce04-770a-4527-8d84-c237c0f8e3f0 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Can i trust your answer? visually grounded video question answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:16a3187ab6e905e702acca01e8d4f77239e89aadabe05490311add70042bb97d

Observation 5cd02eab-610d-404f-a04c-06e910111c30 · outbound

This paper cites LLaV A- NeXT: A strong zero-shot video understanding model.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video LLaV A- NeXT: A strong zero-shot video understanding model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:a249adf67cb538922a8c62f5c189ca928c42e4edb3a7c38492427ca7c80aa985

Observation e9bdedd6-e32a-439f-b16f-fa84fdaa75c1 · outbound

This paper cites Large language models are not robust multiple choice selectors.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Large language models are not robust multiple choice selectors

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:dac96231d19aff18bbe047e6fd316c62a97bcc313c760c87590404313a6dbc2d

Observation 7188285d-ddb4-48c1-b106-af8b91bd5e50 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Xing, Hao Zhang, Joseph E

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:ecbc268440d31aee831f64dadc7d61a9e456293ef0b2802a72afdf3e40dae293

Pith citing papers

No inbound Pith citation observations are available.