Pith. sign in

Paper Citation Record · LEDGER

CoS: Chain-of-Shot Prompting for Long Video Understanding

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 13 inbound Pith citation observations for arXiv:2502.06428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06428 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:31:17.004986Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.322337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:54:22.417988Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3cfbcff-8b4d-487f-a123-e2ae1da5721e · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.907077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.907077Z digest=sha256:6c93c68eb4456a3d74c78ebf1348a2a71cdde266bbd2c5d7ac15bef90df933b9

Observation 219f3a82-976a-4de4-a9b6-24ae8631d8be · outbound

This paper cites Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought.

CoS: Chain-of-Shot Prompting for Long Video Understanding Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.918337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.918337Z digest=sha256:0d079f2e8bc38506b400e465fd0e91d485b06d4792ac01465b455e1a3256f5c9

Observation 1d5c33d3-2e3b-476c-bf38-ed8bb295810e · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

CoS: Chain-of-Shot Prompting for Long Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.929272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.929272Z digest=sha256:79deb1544ab626c0408c86d2ce6f44a480e3c7e2184bb49bfac1796b233d43b5

Observation fdb0e794-5148-4563-b00f-95334841fdcb · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.935585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.935585Z digest=sha256:24a6294b5d405a711e3b262a2f46fea3ecb87233f339354741f2c313f75b7f83

Observation 3ab614dd-274e-44cb-9262-13cb022a3a00 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

CoS: Chain-of-Shot Prompting for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.940927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.940927Z digest=sha256:e909057cc80f812d122a63e07478a0800067d4703f3b732e0ac2ae845a05e014

Observation d94be632-f5e8-4134-a3d9-aaf768b0069f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.945981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.945981Z digest=sha256:5c153ca477d3a4b3722c6d3e687e9e2094319fd540bd1447d7e554041224388a

Observation ffb0da68-ea18-4f26-a790-2bc945b493ce · outbound

This paper cites Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation.

CoS: Chain-of-Shot Prompting for Long Video Understanding Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.951472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.951472Z digest=sha256:f444f19029db1d6e7b794d08cc349f79753ec9d9ffdced6cfbc9ee2ae6f36a0e

Observation 70421979-ef9e-42f6-a482-5d1d7dcc3c9e · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

CoS: Chain-of-Shot Prompting for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.956964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.956964Z digest=sha256:f5c1144ba8c8ee5acfd77b5b776fcf765e79ce724a3fdf4e42d44c4b2f697c9b

Observation 617c1912-08a4-43c1-bf56-54b0c49fdebe · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CoS: Chain-of-Shot Prompting for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.962333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.962333Z digest=sha256:04021c0fecddc2020bc52e0eed33b9f3e26cd03f23bfcde9684c1e51314c8336

Observation cd22ef6a-5f86-4416-ad29-b5ba31469af1 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

CoS: Chain-of-Shot Prompting for Long Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.967715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.967715Z digest=sha256:13cbea3974f21d694635e20dcf96b9e5fd351e538ded586bf3a3aab73c2f36ce

Observation 51403870-c7c4-45e1-b315-addd99359310 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CoS: Chain-of-Shot Prompting for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.972964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.972964Z digest=sha256:d8da6fe66b3017b1f7d246ec75026c9363f4bb4799594fdeb3e098f525459351

Observation 01335963-1e98-4664-9da6-dd7359c035fd · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

CoS: Chain-of-Shot Prompting for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.983567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.983567Z digest=sha256:b7a676ba26e2bf42790f09ac04873f8003369010b8f512804e48db4c3d80f570

Observation 14b1aa0d-bb35-4671-b001-716569e42406 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

CoS: Chain-of-Shot Prompting for Long Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.989115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.989115Z digest=sha256:d9062f1fed9a4a355c20997cf5ac46db90299ca74404d996c78ac58bd5074072

Observation d38a661a-07d0-4337-94fe-949196180adb · outbound

This paper cites Long Context Transfer from Language to Vision.

CoS: Chain-of-Shot Prompting for Long Video Understanding Long Context Transfer from Language to Vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.994358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.994358Z digest=sha256:ad9aa5122266004cd28e247c5eff61eb3a4c5eacf1fa2ffb7789a068973b23aa

Observation 19f72c75-8c39-4b97-9a71-529552cd72b0 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

CoS: Chain-of-Shot Prompting for Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.999790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.999790Z digest=sha256:bbabcadc20fef2caace1df4ea2743f67be5ff377adb146643e24f2227f617157

Observation a48d4f16-66ce-4796-ad69-5d90b1a28171 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CoS: Chain-of-Shot Prompting for Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:17.004986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:17.004986Z digest=sha256:cb3c5318b0d28d8ced101f30b9b36fd4203f19278f6612d39c600ad910e71e40

Observation 2afeedc0-9db2-4457-ba4f-2a60d8bdbda1 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

CoS: Chain-of-Shot Prompting for Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.978260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.978260Z digest=sha256:89273c757008b044886dcb67020d8b7066c5e9b285d1ff220f0db933e489e7e4

Observation df2d0ef1-8548-4d87-b0bf-07d2d3616d7b · outbound

This paper cites Mixtral of Experts.

CoS: Chain-of-Shot Prompting for Long Video Understanding Mixtral of Experts

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.923977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.923977Z digest=sha256:b28355f6087f5f441d0fd12c0f7000222fd1bb7c69bbd7f52aa1fa9e19c3b73f

Observation 299b4d8f-4f58-4ea2-b784-d23ec64689c3 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

CoS: Chain-of-Shot Prompting for Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.900888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.900888Z digest=sha256:e387cc859bc91956635487d7f435369c59b77135373ccd654adb8d0fe8f94865

Observation c2d0ea23-9410-47c9-b020-8e6ac73192ec · outbound

This paper cites VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection.

CoS: Chain-of-Shot Prompting for Long Video Understanding VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.912764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.912764Z digest=sha256:8c41e9289897759a97e313ee66c654ac56bd648f289407c5a49bbbabdcd95465

Pith citing papers

Observation 3beebf5b-6499-41ad-95b5-4d720e775111 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.141676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:dadd3680b4985aad0a50db2b3bd8ba54f8cf126e88ffad0fde7ae97c1c101031

Observation db2ab293-089b-4639-b534-f9246a05f5a8 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.091933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.091933Z digest=sha256:30f9466b42e6153d64bf66a925df42cb50907c4dbad4674c88b57ac45d91af08

Observation 86481275-9056-4f32-9893-441e6d675a85 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.111365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.111365Z digest=sha256:4303ef6ac73cf07ea83a216822441c5d75b03e331d7a752978bdfad9ef628e82

Observation 2608cca2-ea4e-443e-af2f-9c56d48093d5 · inbound

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models cites this paper.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.796119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.796119Z digest=sha256:59a0f27e7feb71ccfab4b3949e57396bacdd43f7b0f2072c13b46f8a69d54446

Observation 9419f8c6-e79c-4b8a-87e9-0244b78e9076 · inbound

Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding cites this paper.

Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:32.335081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:54:32.335081Z digest=sha256:a7663cdbce6468c259b74b063d860f2367b978aed36524d41818a2e82469c8fe

Observation 43fd0233-742b-426d-9870-781e792e941d · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:58.207978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:58.207978Z digest=sha256:157b2b96df2abf9b26e7efb8b8a53671ffe47073eba9d866963392d65a87a2a7

Observation 514acf4c-28ec-443f-9c54-f96477d49094 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.731548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.731548Z digest=sha256:349c40c214ab8013fe17a9d74ea0decbd5eeb15afb62024a1a9956ca0c2a15be

Observation 5985101b-e687-4c51-ad81-83573ef7fe84 · inbound

Act2See: Emergent Active Visual Perception for Video Reasoning cites this paper.

Act2See: Emergent Active Visual Perception for Video Reasoning CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:23.124557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:34:53.683729Z digest=sha256:c90710f36a7d56cf950b65675ec1677de38f6265a9a8709d979dbb02b2ce61f7

Observation 5ed05f66-4101-4cf6-910a-c0112a66b507 · inbound

PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought cites this paper.

PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:31:14.211557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T07:26:17.850942Z digest=sha256:f769829a398a7d43b5b5d213fe530162d967386b2a1c5aa5668786e3a059aaf2

Observation 27bcad07-30e0-45fa-b29d-75c5862742ff · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T05:56:07.910455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:f9b83126fcf4890ca86b4251ea5122be47f8f58a06097ee67465d0c1364c4903

Observation a3bf1934-716b-4c9b-a41b-98b311482ab2 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.419385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:0c2593a7c01b7a31eacb6a87348255f77e2afc38a0164f780c90baa1a27e0f69

Observation 023e1e81-9023-4cee-9360-1e5ca636e1a5 · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.688411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.688411Z digest=sha256:9f6d5a1c96a39b306aec223e491a7b33e2d6c64dab29d2772368cfdb85e9c7c2

Observation 89644a61-823b-477b-9050-d893594e292f · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.322337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.322337Z digest=sha256:93feccd43e3edd6f8d2008097b97abb697c10112946880e546645fdb97109798