Pith. sign in

Paper Citation Record · LEDGER

VidCtx: Context-aware Video Question Answering with Image Models

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2412.17415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17415 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.391770Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.808179Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36b59a45-b65a-4919-a5fd-3d4d50acf58b · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.251138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.251138Z digest=sha256:d6ae9c2a19ff17e2588e2da5c53055d341686bbfd6e216584d5f8f57f37bf3e7

Observation a0b15ddd-9712-425a-8d9b-836f8593e592 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.256797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.256797Z digest=sha256:1202c5c801e107027bc8e3b11a8cb755d19acbb50220fc0f534774c630d26758

Observation d862ea53-59ad-4fec-a7e4-5741daf06e7c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VidCtx: Context-aware Video Question Answering with Image Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.261821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.261821Z digest=sha256:fe36992474dab1c88d6a0ebfbaf2638b20420a9f87a1dafe4dbab3686722a225

Observation 976db303-747f-4c80-bb8d-17d02eaadee7 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VidCtx: Context-aware Video Question Answering with Image Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.267504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.267504Z digest=sha256:3eab1f6865e5d4744464dc1a37932d354d7c95b399cc87b87ec1e4ee8da6fc4f

Observation 2cec20e0-c2b2-46ca-926a-2cd54c61a4e0 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding bench- mark,.

VidCtx: Context-aware Video Question Answering with Image Models Mvbench: A comprehensive multi-modal video understanding bench- mark,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.860581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.272473Z digest=sha256:c054baaeb3e33cb8d80e677072160b52dced4f7e02ac45d1a4a03e3608a9f369

Observation 34d2f7df-a516-4953-bf2c-2d7e264b4e14 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VidCtx: Context-aware Video Question Answering with Image Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.276954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.276954Z digest=sha256:27227d536fc9c8be68cf9687b867acd70f14f923511245efbd2c49e6739e19ae

Observation 7f5ca850-f005-485f-bcec-d2d08ccae1fd · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.282682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.282682Z digest=sha256:462da3c28ae40f6e0aee59435d11d499b5e9625057a96959fdc0166946ec8fdd

Observation b73ffffa-6723-4754-8183-1b97d3d9654e · outbound

This paper cites Self- chained image-language model for video localization and question answering,.

VidCtx: Context-aware Video Question Answering with Image Models Self- chained image-language model for video localization and question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.846020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.287475Z digest=sha256:d554593461e91b1d87080592bf75d3db6089fb3b3ded648a1a6c78040888f57e

Observation e782af15-4a88-435d-b6b3-b7598ae1a837 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

VidCtx: Context-aware Video Question Answering with Image Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.291907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.291907Z digest=sha256:45deeacce2cd09c92955b03336ad8ebbc5e56e0dd527a93074a90b4c76bf693b

Observation 92a4f946-5951-4242-9af4-05643838b515 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent,.

VidCtx: Context-aware Video Question Answering with Image Models Videoagent: Long-form video understanding with large language model as agent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.832421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.296639Z digest=sha256:95a36ef7d08beecf5f3b2b3f5977c842e17c63ce0fbcef038dcf5d7e51a1d71c

Observation 779c2b7a-300b-454d-b289-4758696ef7eb · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

VidCtx: Context-aware Video Question Answering with Image Models A Simple LLM Framework for Long-Range Video Question-Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.302286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.302286Z digest=sha256:57e49c2e298d02f40cb88667a26093890dfb9157273cc8199f56cdcf97c7219c

Observation fb9a8cb9-4666-4f03-af54-b49098b2ee9b · outbound

This paper cites Language Repository for Long Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Language Repository for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.307324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.307324Z digest=sha256:6b521daec6ddce7734d52d5329b88d9e9ec2405180d2d64acefd52396d96ef84

Observation daf1d424-5f21-4a3c-8b3d-d98dfa29eb43 · outbound

This paper cites Question-instructed visual de- scriptions for zero-shot video answering,.

VidCtx: Context-aware Video Question Answering with Image Models Question-instructed visual de- scriptions for zero-shot video answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.816975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.313074Z digest=sha256:d0d2efc9dd29a023b906a55e690ad3af02d17f22121c9c80358655465dbdc1ac

Observation 84ac54d0-55df-4757-a43e-2305ddb96686 · outbound

This paper cites Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.317436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.317436Z digest=sha256:00a7bd47e9c693fe617255b623c6b4e0f335ae3b9e7cfb1a1178246692a7f40a

Observation 6736c950-611b-4dab-ad9e-ced58d2443c2 · outbound

This paper cites Large language models can be easily distracted by irrelevant context,.

VidCtx: Context-aware Video Question Answering with Image Models Large language models can be easily distracted by irrelevant context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.802738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.322333Z digest=sha256:417f3885c4fee3a531a41b283f8838fbbf4194afac54cf9d5706ff1e4b813619

Observation db2d6455-f26a-4ea6-bc2a-23a4cd9236ee · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

VidCtx: Context-aware Video Question Answering with Image Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.787037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.326968Z digest=sha256:5d6367d8a56a521d5d86b7333534e453f51c4226c24864562989eb0e68f0b298

Observation 5e862572-2f5e-4d59-b473-7d4514489060 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.331502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.331502Z digest=sha256:9176d272be987db809abc640a178bff22be6030378fb55495914ed0fc4a45285

Observation f1290a7d-a931-433b-a740-3d6bc81295ca · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,.

VidCtx: Context-aware Video Question Answering with Image Models Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.771872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.337272Z digest=sha256:8f943f5b5210843f45e5a255eb6347762470ea372dd40831c8f0b495348d1f30

Observation c3947363-3c7c-4dd3-b60e-3dc3edf60c6b · outbound

This paper cites Enhancing multimodal sentiment analysis via learning from large language model,.

VidCtx: Context-aware Video Question Answering with Image Models Enhancing multimodal sentiment analysis via learning from large language model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.756953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.341609Z digest=sha256:00e776c2bd224f01ffd38d8574e7dace6bdd3aa774f7ea1b102f3960590d7ad8

Observation bd9a91da-8403-4e3e-90cb-d3d9c73af29a · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition,.

VidCtx: Context-aware Video Question Answering with Image Models Video-of-thought: Step-by-step video reasoning from perception to cognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.742345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.346135Z digest=sha256:c5817dd1365c461c6e72be1d15b13d64541cc2f26d79ee098f04d1a3bedf0966

Observation 3e5aa05f-549f-4f40-9081-d470ee0498b7 · outbound

This paper cites Vamos: Versatile Action Models for Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Vamos: Versatile Action Models for Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.350899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.350899Z digest=sha256:4e75a6b50283fc2239b5a71aea6ff3991062016c884d75c40651b2681081f5ad

Observation 81f5fac8-7b5a-489e-bab4-c6a35ee3a06c · outbound

This paper cites Large language models are zero-shot reasoners,.

VidCtx: Context-aware Video Question Answering with Image Models Large language models are zero-shot reasoners,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.727217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.356308Z digest=sha256:4ba5451a5bf0dc5ce2ffd9e0f51c0396944080319aa60b5730e777f57ccebb68

Observation 7e21f624-5f9d-4657-9b4b-40809b13690e · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

VidCtx: Context-aware Video Question Answering with Image Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.710596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.360643Z digest=sha256:55858942c9904bf655c0452ecbba6eb432be8312d40a3134fc03cd49e8fb951a

Observation 09818ffb-dafd-4aeb-a383-1fd0d3bc1f0d · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal mod- els,.

VidCtx: Context-aware Video Question Answering with Image Models Compositional chain-of-thought prompting for large multimodal mod- els,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.691893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.364961Z digest=sha256:29859cfd036be8d68cddf3ac3ee152ececf3592018c232de28edadc286e83080

Observation fa23c586-feed-418c-91dd-c826e7f0678e · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

VidCtx: Context-aware Video Question Answering with Image Models Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.676088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.369044Z digest=sha256:ab2fe762470d8f2ce5a70c73c20753b34373db3c68543c974d8cfb04729239cb

Observation ed5e1194-f9b4-4603-a465-a88004e6276a · outbound

This paper cites Intentqa: Context- aware video intent reasoning,.

VidCtx: Context-aware Video Question Answering with Image Models Intentqa: Context- aware video intent reasoning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.660083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.374146Z digest=sha256:c4540568468e9a09142e4e8786809b4e101c7e1657a1447b35e22b2d2d3998ee

Observation 1dc03d09-aa3f-4fc7-b7d6-a78b825ce3d0 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

VidCtx: Context-aware Video Question Answering with Image Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.378423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.378423Z digest=sha256:6b94c100bcf9a44232e1a8d7e60f59b1f8deb31aecc60499c6947df95fc490a2

Observation 29a74aa0-e440-4351-acaa-3a00735895d7 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models,.

VidCtx: Context-aware Video Question Answering with Image Models Verbs in action: Improving verb understanding in video-language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.644900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.382519Z digest=sha256:66fdfc0564a06bf12c3628e7a9376efb9e94f59cca44fe0cad5f40abb1ce42eb

Observation 527a5034-1726-4fac-984e-db07468774e6 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.387257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.387257Z digest=sha256:ebea50fe99dbabd4a73614952d0a7c3f145c0998b5e1ffc0ba469c9e03f9af4e

Observation 199033f3-a787-4436-aa41-9da7e36fc172 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

VidCtx: Context-aware Video Question Answering with Image Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.628488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:29:51.391770Z digest=sha256:533e8d5731dd24214a148497359e95bbe77648372745c2b198e1c047d03046f0

Pith citing papers

Observation 4dccfb87-3f24-4aae-bd55-8b45679640d5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning VidCtx: Context-aware Video Question Answering with Image Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.809467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:ce244310a35c747fc3990f3d1d72988257e0190c965bf6268465d5946b336f41