Pith. sign in

Paper Citation Record · LEDGER

How Important are Videos for Training Video LLMs?

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2506.06928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06928 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:51:07.923531Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:40:00.723218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65c748a2-cd0b-461a-97a9-c9cd92477f4f · outbound

This paper cites Qwen2.5-VL Technical Report.

How Important are Videos for Training Video LLMs? Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.773495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.773495Z digest=sha256:db03d08db550caf62aa13597a5134358066fcad4deb122eb3debaf2034b3f8b0

Observation 285cf941-cb86-40de-8a0d-f2cc99af891e · outbound

This paper cites ShareGPT4Video: Improving Video Under- standing and Generation with Better Captions.

How Important are Videos for Training Video LLMs? ShareGPT4Video: Improving Video Under- standing and Generation with Better Captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.459611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.779399Z digest=sha256:5bc3995a58e4c351f7687dd882848d2890068ac3b685fca447412d1369d2b3e9

Observation ea581482-7c3a-42a9-a25d-ee24df90f017 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

How Important are Videos for Training Video LLMs? VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.784472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.784472Z digest=sha256:6d9b0a711a726b7deb2381796124b7179312a70a360831c48c43a11c78785b96

Observation 03cff194-f5a0-4074-90c0-a2701db95f6b · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

How Important are Videos for Training Video LLMs? Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.790019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.790019Z digest=sha256:3b0e35c157e0e87710e8abb6560fa9f3c1c239c7611ee1c10e81114a6ca59917

Observation 4f6aa8c0-6d3b-4c63-9c20-f096c821bc1d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

How Important are Videos for Training Video LLMs? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.795669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.795669Z digest=sha256:ac44a9d0e43a29879a22e102c4b3b20a0e464b017984087ec4de24f04ca91696

Observation 1c2a4e8b-65ba-4337-a695-abd2ff200285 · outbound

This paper cites The Llama 3 Herd of Models.

How Important are Videos for Training Video LLMs? The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.800928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.800928Z digest=sha256:98a4685208338ff7abb22632587c5b2c04ceaa747150ece94d0981d83f015fae

Observation bf156721-0a82-4991-b65e-8aa7f00a3508 · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

How Important are Videos for Training Video LLMs? MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.443799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.806446Z digest=sha256:490b2efc4f00222920f82daf08d1c27f71ac4354b43981008ce84cfddda404d2

Observation 063e0f5c-e995-4454-a493-83357a8972d4 · outbound

This paper cites LoRA: Low- Rank Adaptation of Large Language Models.

How Important are Videos for Training Video LLMs? LoRA: Low- Rank Adaptation of Large Language Models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.427508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.811592Z digest=sha256:62e20d2af7882cf3e0882ae4d7b6cca14224508d6b21bdba5ea0685220f029dd

Observation ed31497f-5f07-422d-869b-cc8a4a3fd827 · outbound

This paper cites TGIF-QA: Toward Spatio-Temporal Reason- ing in Visual Question Answering.

How Important are Videos for Training Video LLMs? TGIF-QA: Toward Spatio-Temporal Reason- ing in Visual Question Answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.408349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.817057Z digest=sha256:c023d5f5b933feb99ffa36b48d20aa378fb22f4ed7d4ed76343a6414b47a5b1a

Observation 09ad63b2-12d0-452e-b6c1-6bcae633622a · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

How Important are Videos for Training Video LLMs? CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.389004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.822018Z digest=sha256:a92cbe1d8b547f8d4802be4f79a0835315f74d6c0c3f60d76cbfdb318ed67d70

Observation 053fb2fb-be20-41d5-ad0e-0dc963146e8b · outbound

This paper cites JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre.

How Important are Videos for Training Video LLMs? JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.373009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.827043Z digest=sha256:0497765540ddb9de176340df1b19fe250047886a7c13aad5163bfcfd8f044a8e

Observation a8b5ff99-15ec-41f0-8177-5e10f582f342 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

How Important are Videos for Training Video LLMs? LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.831670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.831670Z digest=sha256:3903077d86d8484c2687bf955966f680aedfc3f4889adbbc40734c8a27915062

Observation 39381d6c-50ee-468e-a4a9-55417ca9e93c · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark.

How Important are Videos for Training Video LLMs? MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.357663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.836882Z digest=sha256:8dcf2bae376ead1a48419f97c4cc0612bc8b1f3ca7f20db5ff5d4cf60d77cb4a

Observation ad71e2fe-5895-4c17-814c-8f8992cfa555 · outbound

This paper cites Temporal Preference Optimization for Long-Form Video Understanding.

How Important are Videos for Training Video LLMs? Temporal Preference Optimization for Long-Form Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.841588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.841588Z digest=sha256:e277ab56479d50b3b255debca94767a6544263548165310455569cbb13686b3e

Observation 6e1e2824-7908-4c1d-9206-0cf2f47da698 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

How Important are Videos for Training Video LLMs? LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.342865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.846524Z digest=sha256:1bddc494c9def44c086b2ec04dccf34f5dc244c6d0620a8c6b52e5541b5a9930

Observation 9f8edbfe-10aa-4d2c-ae75-c677098d699e · outbound

This paper cites Video-LLaV A: Learning United Visual Representation by Alignment Before Projection.

How Important are Videos for Training Video LLMs? Video-LLaV A: Learning United Visual Representation by Alignment Before Projection

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.327932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.851065Z digest=sha256:34a514472f3058b1733bf7333948ed26383bedafb6e94c4bab300285981da9b4

Observation 5b7ce0b0-5171-461d-a9b3-8f1314059ba4 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

How Important are Videos for Training Video LLMs? Microsoft COCO: Common Objects in Context

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.312225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.855756Z digest=sha256:251427f1240d3dfd244ae0ede8866148dbadbe22340c0120b1a29e41c1684923

Observation 9ebd11f4-cb7d-4fc7-9870-6582a89258f8 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

How Important are Videos for Training Video LLMs? Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.860377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.860377Z digest=sha256:9956ba1c9639712c4b111d42c431846158329cee94e93407a4bd96cb5fd4ea44

Observation bf28ca4d-4cfe-4ed5-8dd5-db851bd7a74d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Under- standing via Large Vision and Language Models.

How Important are Videos for Training Video LLMs? Video-ChatGPT: Towards Detailed Video Under- standing via Large Vision and Language Models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.296144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.865389Z digest=sha256:0e380dec2da1f31c634fa7033fbc10708f4f8148f33a2be3813a43c4c989e23a

Observation d9d66231-4270-478c-bf57-9fa8221d2497 · outbound

This paper cites an unresolved cited work.

How Important are Videos for Training Video LLMs? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:51:08.279071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.869984Z digest=sha256:9cb0e9918f29149f517b5e11735066392e65510a88c7d7b608c620a4b4db0316

Observation f86bf564-bf9b-4076-b424-f45354ce05a1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

How Important are Videos for Training Video LLMs? Learning Transferable Visual Models From Natural Language Super- vision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.262254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.874688Z digest=sha256:7b9777c697a02ec4abd09196c436d80e53aab0f0ce6acaae6cabd862bc7ccd1c

Observation 6dce3466-df81-459c-a18d-87100f0d2f40 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

How Important are Videos for Training Video LLMs? LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.879150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.879150Z digest=sha256:bb5933d93f71c5fec5a3ecc9ea65fa1996f3720d87f53bca8d14b22c8bdda002

Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.884760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.884760Z digest=sha256:5529ae46b54fc51f699503793b337a49286709f809864ca773bb9dda2fce149a

Observation c3ee944a-08d6-479e-a3b3-22b187f7925e · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Inter- leaved Video-Language Understanding.

How Important are Videos for Training Video LLMs? LongVideoBench: A Benchmark for Long-context Inter- leaved Video-Language Understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.244690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.889776Z digest=sha256:40e24daf796c21c4161c7d067e256fa035b5547cc4f5f0d9c3f78c82e8ae4151

Observation 4e0dc231-9de6-4f15-8f1e-0112554d8688 · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Language.

How Important are Videos for Training Video LLMs? MSR-VTT: A Large Video Description Dataset for Bridging Video and Language

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.228268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.895050Z digest=sha256:8831aaea5d476b4995c7681912552bbf066c015183470cd47a8a5f6fe6f82bfb

Observation 897dc92b-73ff-4532-9744-3ca2e42c1f23 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

How Important are Videos for Training Video LLMs? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.899664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.899664Z digest=sha256:870da6794c5d497f8eb3ed2d927a6a24e600f9ecb41ddc15fbca3e0e53ac2bfa

Observation 87ec161b-a15b-4e8c-afd8-4be53ef31414 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

How Important are Videos for Training Video LLMs? CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.211642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.904430Z digest=sha256:254ac19f0d443dfc0788df82e0003e83980cd2e5a17125458a69a55e3002c240

Observation ff108d37-dfb5-4f80-8fa6-b5a81e09e326 · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answer- ing.

How Important are Videos for Training Video LLMs? ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answer- ing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.195305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.909084Z digest=sha256:9a8fe3e87fb7b889f5a42dc792e2f476c1576cdc6a9d56c7ae8123523493af2d

Observation c96d5b97-b387-450b-a5f1-13dc333d2a9f · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

How Important are Videos for Training Video LLMs? Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.913992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.913992Z digest=sha256:a75524227d9a7b01e33313737ac6485d48ac968f17edaba39a9016ff81ed3c5e

Observation 5faf4ef4-d8e6-4340-b5db-458784a5186f · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

How Important are Videos for Training Video LLMs? Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.177820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:51:07.918987Z digest=sha256:f72af90237d96d3a41b751c757a5809db42a532fb4f1aed238372a43d50090f3

Observation b6f15c9d-18d7-4e73-9bcc-4a2501bf2bc6 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

How Important are Videos for Training Video LLMs? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.923531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.923531Z digest=sha256:3faf144c623bf909d42e1e19d0cbaa7e734b840f42d96c43f75a6db5a8c0a309

Pith citing papers

Observation 8c7647d5-6f3e-4110-8c64-34e90ee76aa9 · inbound

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks cites this paper.

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks How Important are Videos for Training Video LLMs?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:00.723218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:00.723218Z digest=sha256:89ed2e6291ff0c7f0ca6027db017766896a193a5916918549d4955b541880f6b