Pith. sign in

Paper Citation Record · LEDGER

How Important are Videos for Training Video LLMs?

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2506.06928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06928 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:51:07.923531Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:40:00.723218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65c748a2-cd0b-461a-97a9-c9cd92477f4f · outbound

This paper cites Qwen2.5-VL Technical Report.

How Important are Videos for Training Video LLMs? Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.773495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.773495Z digest=sha256:e651e93a7aee712782b6e0bad270fa5114531a74bca752813f9882df05a9ab8d

Observation 285cf941-cb86-40de-8a0d-f2cc99af891e · outbound

This paper cites ShareGPT4Video: Improving Video Under- standing and Generation with Better Captions.

How Important are Videos for Training Video LLMs? ShareGPT4Video: Improving Video Under- standing and Generation with Better Captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.459611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.779399Z digest=sha256:2d36af797d6f65aca92d33bce680efb795b212e92660a9437749b170ca72aca4

Observation ea581482-7c3a-42a9-a25d-ee24df90f017 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

How Important are Videos for Training Video LLMs? VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.784472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.784472Z digest=sha256:008a458e9b1bb043fa0d030193eef3ef31969d20bdc0566dccdf8e7edcbd32df

Observation 03cff194-f5a0-4074-90c0-a2701db95f6b · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

How Important are Videos for Training Video LLMs? Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.790019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.790019Z digest=sha256:4ba97b0c0395bc7982dd64aa197169915a56b5436fc2df18762324f6eea0ad64

Observation 4f6aa8c0-6d3b-4c63-9c20-f096c821bc1d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

How Important are Videos for Training Video LLMs? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.795669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.795669Z digest=sha256:75ab6e57d777d49dbdcd6f95fc2c47124fcf81b877255feaf321e1179d7a94bb

Observation 1c2a4e8b-65ba-4337-a695-abd2ff200285 · outbound

This paper cites The Llama 3 Herd of Models.

How Important are Videos for Training Video LLMs? The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.800928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.800928Z digest=sha256:fa9708a6be49a9dd381619239bc18c07a9f886f49451c50edce01b22a331f1ad

Observation bf156721-0a82-4991-b65e-8aa7f00a3508 · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

How Important are Videos for Training Video LLMs? MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.443799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.806446Z digest=sha256:1eb4721154d95bd5a5d132754233e31238fbc40692b9d1e88d7abbb46876cb93

Observation 063e0f5c-e995-4454-a493-83357a8972d4 · outbound

This paper cites LoRA: Low- Rank Adaptation of Large Language Models.

How Important are Videos for Training Video LLMs? LoRA: Low- Rank Adaptation of Large Language Models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.427508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.811592Z digest=sha256:8b1f64f143c4c9a6ff950fa9e05c120f78abb0c5d772547488043ab19de3db37

Observation ed31497f-5f07-422d-869b-cc8a4a3fd827 · outbound

This paper cites TGIF-QA: Toward Spatio-Temporal Reason- ing in Visual Question Answering.

How Important are Videos for Training Video LLMs? TGIF-QA: Toward Spatio-Temporal Reason- ing in Visual Question Answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.408349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.817057Z digest=sha256:5fed8c4dd4569474ffd6c65201ecaf6084f9361fe698c4c4415d69ab4eee0d6e

Observation 09ad63b2-12d0-452e-b6c1-6bcae633622a · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

How Important are Videos for Training Video LLMs? CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.389004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.822018Z digest=sha256:e20535e8fff55996741b42877bd77d95a28c6ad460878bd556edc9b5a9ece5dd

Observation 053fb2fb-be20-41d5-ad0e-0dc963146e8b · outbound

This paper cites JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre.

How Important are Videos for Training Video LLMs? JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.373009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.827043Z digest=sha256:b7cf1caee187bbafb4cb16ad91e505143736e60f41cbd84ec3ec70149e5cf78d

Observation a8b5ff99-15ec-41f0-8177-5e10f582f342 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

How Important are Videos for Training Video LLMs? LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.831670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.831670Z digest=sha256:8509debdd1360dc909b7cca02171e571fbae1d5cc91fe7efb637517328b47992

Observation 39381d6c-50ee-468e-a4a9-55417ca9e93c · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark.

How Important are Videos for Training Video LLMs? MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.357663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.836882Z digest=sha256:afa84876dcba1b552f07d21a823bac0ca861ce52f3bb3cb0e4df75b7e44e803c

Observation ad71e2fe-5895-4c17-814c-8f8992cfa555 · outbound

This paper cites Temporal Preference Optimization for Long-Form Video Understanding.

How Important are Videos for Training Video LLMs? Temporal Preference Optimization for Long-Form Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.841588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.841588Z digest=sha256:ce7dff9317053f9932ecd70665014fb527839e02bc040fcbfdbc7dad248c99c5

Observation 6e1e2824-7908-4c1d-9206-0cf2f47da698 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

How Important are Videos for Training Video LLMs? LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.342865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.846524Z digest=sha256:eb236e08537c64ed89c1956ca821cdc7e54a1643c7b35a2342c33ab4430aab4a

Observation 9f8edbfe-10aa-4d2c-ae75-c677098d699e · outbound

This paper cites Video-LLaV A: Learning United Visual Representation by Alignment Before Projection.

How Important are Videos for Training Video LLMs? Video-LLaV A: Learning United Visual Representation by Alignment Before Projection

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.327932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.851065Z digest=sha256:f956db9334131f11d9e81a62f5ac184d9b52a149c2b3636b05a4c2358d0b91e7

Observation 5b7ce0b0-5171-461d-a9b3-8f1314059ba4 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

How Important are Videos for Training Video LLMs? Microsoft COCO: Common Objects in Context

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.312225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.855756Z digest=sha256:780e2c4f1f90ea12a91a99399ac1f8f10eb63594f19247e3819c3bdd295ecdc5

Observation 9ebd11f4-cb7d-4fc7-9870-6582a89258f8 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

How Important are Videos for Training Video LLMs? Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.860377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.860377Z digest=sha256:d5e89fb899d5ef3a5d24ba07bcc017e929c0bea8f8b84669c10fa0bc5fca1dca

Observation bf28ca4d-4cfe-4ed5-8dd5-db851bd7a74d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Under- standing via Large Vision and Language Models.

How Important are Videos for Training Video LLMs? Video-ChatGPT: Towards Detailed Video Under- standing via Large Vision and Language Models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.296144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.865389Z digest=sha256:c16a96d78ba20af0d44db157dc213e575e15ef182322e520a799d8db74d20c74

Observation d9d66231-4270-478c-bf57-9fa8221d2497 · outbound

This paper cites an unresolved cited work.

How Important are Videos for Training Video LLMs? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:51:08.279071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.869984Z digest=sha256:c0c37a034c919cc8def1a28702e08aec5063bd37770216d46f88e191c9a4ef7d

Observation f86bf564-bf9b-4076-b424-f45354ce05a1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

How Important are Videos for Training Video LLMs? Learning Transferable Visual Models From Natural Language Super- vision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.262254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.874688Z digest=sha256:b3f731060b9738a714d249db8800353d84cc72d3dbaa7e4c0d48b17ce5b2a0fb

Observation 6dce3466-df81-459c-a18d-87100f0d2f40 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

How Important are Videos for Training Video LLMs? LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.879150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.879150Z digest=sha256:22fb822df1fc7bb9309054593e585a94434b8893b8d99f89378e5826febc27ca

Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.884760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.884760Z digest=sha256:6d9c3d722312e00297d2d03b1e8285d13baee77d945d99b6fe9cf85a66e36fe5

Observation c3ee944a-08d6-479e-a3b3-22b187f7925e · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Inter- leaved Video-Language Understanding.

How Important are Videos for Training Video LLMs? LongVideoBench: A Benchmark for Long-context Inter- leaved Video-Language Understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.244690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.889776Z digest=sha256:2329317d9d8b6f52c160b2875dfa91335c0326f7195e20feae11ed031bd0fea3

Observation 4e0dc231-9de6-4f15-8f1e-0112554d8688 · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Language.

How Important are Videos for Training Video LLMs? MSR-VTT: A Large Video Description Dataset for Bridging Video and Language

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.228268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.895050Z digest=sha256:867a1d48b01cfd6f3254ec80d616156cc0b3c5164e887b3bbeeabd62e2cf38c8

Observation 897dc92b-73ff-4532-9744-3ca2e42c1f23 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

How Important are Videos for Training Video LLMs? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.899664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.899664Z digest=sha256:3fa25b908e66c04a1e6fad268d561d6daf3389d9e182cb54bfc5d278c3db46eb

Observation 87ec161b-a15b-4e8c-afd8-4be53ef31414 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

How Important are Videos for Training Video LLMs? CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.211642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.904430Z digest=sha256:a7b18f63cdb39da46b83a7f5b77bba142f7bc0cedeef2c4f035cd32ec27fc367

Observation ff108d37-dfb5-4f80-8fa6-b5a81e09e326 · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answer- ing.

How Important are Videos for Training Video LLMs? ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answer- ing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.195305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.909084Z digest=sha256:9f3e5be1b58b43df557bb1b84c4ad5f8d04b1c6aa40ee2043b544e1fe0078f7a

Observation c96d5b97-b387-450b-a5f1-13dc333d2a9f · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

How Important are Videos for Training Video LLMs? Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.913992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.913992Z digest=sha256:a04faeaffca4a64fdd9aa99f8911ac1cae87ce241984e893b66f04bf70f161ec

Observation 5faf4ef4-d8e6-4340-b5db-458784a5186f · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

How Important are Videos for Training Video LLMs? Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.177820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:51:07.918987Z digest=sha256:714d0937d21453d840045d514b9fe4c78238a164c48a916c947958588a400504

Observation b6f15c9d-18d7-4e73-9bcc-4a2501bf2bc6 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

How Important are Videos for Training Video LLMs? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.923531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.923531Z digest=sha256:38b00f30d53b8801db4d48140fe01f65d78b991c8f3123a31fe607b824fdb9ea

Pith citing papers

Observation 8c7647d5-6f3e-4110-8c64-34e90ee76aa9 · inbound

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks cites this paper.

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks How Important are Videos for Training Video LLMs?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:00.723218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:00.723218Z digest=sha256:fb3d6d95b4624b38865c24cedf259b0f626b47f6239be6e54bed7ce251f0baf7