Pith. sign in

Paper Citation Record · LEDGER

VidCtx: Context-aware Video Question Answering with Image Models

As of 21 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2412.17415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17415 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.391770Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.808179Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36b59a45-b65a-4919-a5fd-3d4d50acf58b · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.251138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.251138Z digest=sha256:d6ae9c2a19ff17e2588e2da5c53055d341686bbfd6e216584d5f8f57f37bf3e7

Observation a0b15ddd-9712-425a-8d9b-836f8593e592 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.256797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.256797Z digest=sha256:1202c5c801e107027bc8e3b11a8cb755d19acbb50220fc0f534774c630d26758

Observation d862ea53-59ad-4fec-a7e4-5741daf06e7c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VidCtx: Context-aware Video Question Answering with Image Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.261821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.261821Z digest=sha256:fe36992474dab1c88d6a0ebfbaf2638b20420a9f87a1dafe4dbab3686722a225

Observation 976db303-747f-4c80-bb8d-17d02eaadee7 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VidCtx: Context-aware Video Question Answering with Image Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.267504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.267504Z digest=sha256:3eab1f6865e5d4744464dc1a37932d354d7c95b399cc87b87ec1e4ee8da6fc4f

Observation 2cec20e0-c2b2-46ca-926a-2cd54c61a4e0 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding bench- mark,.

VidCtx: Context-aware Video Question Answering with Image Models Mvbench: A comprehensive multi-modal video understanding bench- mark,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.860581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.272473Z digest=sha256:c29055e5cdb31f19c96c4a1628fef9f93d8a5304bf35caa35b2c51256e0730a1

Observation 34d2f7df-a516-4953-bf2c-2d7e264b4e14 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VidCtx: Context-aware Video Question Answering with Image Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.276954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.276954Z digest=sha256:27227d536fc9c8be68cf9687b867acd70f14f923511245efbd2c49e6739e19ae

Observation 7f5ca850-f005-485f-bcec-d2d08ccae1fd · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.282682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.282682Z digest=sha256:462da3c28ae40f6e0aee59435d11d499b5e9625057a96959fdc0166946ec8fdd

Observation b73ffffa-6723-4754-8183-1b97d3d9654e · outbound

This paper cites Self- chained image-language model for video localization and question answering,.

VidCtx: Context-aware Video Question Answering with Image Models Self- chained image-language model for video localization and question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.846020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.287475Z digest=sha256:1ab47555460752a767c9822f8040247e4256b9f3ee96fda365725644673974d8

Observation e782af15-4a88-435d-b6b3-b7598ae1a837 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

VidCtx: Context-aware Video Question Answering with Image Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.291907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.291907Z digest=sha256:45deeacce2cd09c92955b03336ad8ebbc5e56e0dd527a93074a90b4c76bf693b

Observation 92a4f946-5951-4242-9af4-05643838b515 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent,.

VidCtx: Context-aware Video Question Answering with Image Models Videoagent: Long-form video understanding with large language model as agent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.832421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.296639Z digest=sha256:8c2b31a3c80e9c60a8f86da5ec8eb49bac07b92e88b9d694ea365e3223367c84

Observation 779c2b7a-300b-454d-b289-4758696ef7eb · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

VidCtx: Context-aware Video Question Answering with Image Models A Simple LLM Framework for Long-Range Video Question-Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.302286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.302286Z digest=sha256:57e49c2e298d02f40cb88667a26093890dfb9157273cc8199f56cdcf97c7219c

Observation fb9a8cb9-4666-4f03-af54-b49098b2ee9b · outbound

This paper cites Language Repository for Long Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Language Repository for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.307324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.307324Z digest=sha256:6de6c984bd71903aad1721ba3f8229f0f41ea520077cd12421e6fae416c1e04c

Observation daf1d424-5f21-4a3c-8b3d-d98dfa29eb43 · outbound

This paper cites Question-instructed visual de- scriptions for zero-shot video answering,.

VidCtx: Context-aware Video Question Answering with Image Models Question-instructed visual de- scriptions for zero-shot video answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.816975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.313074Z digest=sha256:37242c77def5ec35b638b4348870ff37228273b89bc06b7db2b407ccfb45ea4f

Observation 84ac54d0-55df-4757-a43e-2305ddb96686 · outbound

This paper cites Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.317436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.317436Z digest=sha256:00a7bd47e9c693fe617255b623c6b4e0f335ae3b9e7cfb1a1178246692a7f40a

Observation 6736c950-611b-4dab-ad9e-ced58d2443c2 · outbound

This paper cites Large language models can be easily distracted by irrelevant context,.

VidCtx: Context-aware Video Question Answering with Image Models Large language models can be easily distracted by irrelevant context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.802738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.322333Z digest=sha256:acc3249b9f06b6c9af5a2ff041c664f4478b46496ad1d64179c252b0c8c037ba

Observation db2d6455-f26a-4ea6-bc2a-23a4cd9236ee · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

VidCtx: Context-aware Video Question Answering with Image Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.787037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.326968Z digest=sha256:d24ba83da20445a76b4a98d1b4acd123370e75f148c9c38c8456bd30fe04c841

Observation 5e862572-2f5e-4d59-b473-7d4514489060 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

VidCtx: Context-aware Video Question Answering with Image Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.331502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.331502Z digest=sha256:9176d272be987db809abc640a178bff22be6030378fb55495914ed0fc4a45285

Observation f1290a7d-a931-433b-a740-3d6bc81295ca · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,.

VidCtx: Context-aware Video Question Answering with Image Models Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.771872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.337272Z digest=sha256:50f410db12d22f1559d978aa69a59569335ce97f97c39976a4539a0230990459

Observation c3947363-3c7c-4dd3-b60e-3dc3edf60c6b · outbound

This paper cites Enhancing multimodal sentiment analysis via learning from large language model,.

VidCtx: Context-aware Video Question Answering with Image Models Enhancing multimodal sentiment analysis via learning from large language model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.756953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.341609Z digest=sha256:4f5542ce8b233447a5e906495cc18184b95c87d6bf5d80e8e13bce8e7b97d7f0

Observation bd9a91da-8403-4e3e-90cb-d3d9c73af29a · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition,.

VidCtx: Context-aware Video Question Answering with Image Models Video-of-thought: Step-by-step video reasoning from perception to cognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.742345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.346135Z digest=sha256:4c4fc2898af2d0304376c165106502bad65952258b8a8ecb0c5dfd3beda895c1

Observation 3e5aa05f-549f-4f40-9081-d470ee0498b7 · outbound

This paper cites Vamos: Versatile Action Models for Video Understanding.

VidCtx: Context-aware Video Question Answering with Image Models Vamos: Versatile Action Models for Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.350899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.350899Z digest=sha256:4e75a6b50283fc2239b5a71aea6ff3991062016c884d75c40651b2681081f5ad

Observation 81f5fac8-7b5a-489e-bab4-c6a35ee3a06c · outbound

This paper cites Large language models are zero-shot reasoners,.

VidCtx: Context-aware Video Question Answering with Image Models Large language models are zero-shot reasoners,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.727217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.356308Z digest=sha256:8bb393b8288a21c0a077f15c482cb7887d10be06b8e4fc436c64f8db3a70c5e3

Observation 7e21f624-5f9d-4657-9b4b-40809b13690e · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

VidCtx: Context-aware Video Question Answering with Image Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.710596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.360643Z digest=sha256:3752e6eb50561595bb3c16c6c80f26c9c39f5e1b0c48e4dab7a6e1e33f02161a

Observation 09818ffb-dafd-4aeb-a383-1fd0d3bc1f0d · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal mod- els,.

VidCtx: Context-aware Video Question Answering with Image Models Compositional chain-of-thought prompting for large multimodal mod- els,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.691893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.364961Z digest=sha256:0b995c91a5f2deffbf665e9bec6f8ca0f7544ed2069a07c8e15cb2184a35cc93

Observation fa23c586-feed-418c-91dd-c826e7f0678e · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

VidCtx: Context-aware Video Question Answering with Image Models Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.676088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.369044Z digest=sha256:136c05189ff3b0d9d628ca46ee7dc05b3d96bfd3af1a359185e58099121a5376

Observation ed5e1194-f9b4-4603-a465-a88004e6276a · outbound

This paper cites Intentqa: Context- aware video intent reasoning,.

VidCtx: Context-aware Video Question Answering with Image Models Intentqa: Context- aware video intent reasoning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.660083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.374146Z digest=sha256:0e490f24f2484ad22b1587dd59b2c4a38bdd58f01e32855b1e854da0d169206c

Observation 1dc03d09-aa3f-4fc7-b7d6-a78b825ce3d0 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

VidCtx: Context-aware Video Question Answering with Image Models STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.378423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.378423Z digest=sha256:6b94c100bcf9a44232e1a8d7e60f59b1f8deb31aecc60499c6947df95fc490a2

Observation 29a74aa0-e440-4351-acaa-3a00735895d7 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models,.

VidCtx: Context-aware Video Question Answering with Image Models Verbs in action: Improving verb understanding in video-language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.644900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.382519Z digest=sha256:34c113b4533b6d219df9ae44aff16d7e4435b68c3c8589eafcfdcf76423cf2af

Observation 527a5034-1726-4fac-984e-db07468774e6 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VidCtx: Context-aware Video Question Answering with Image Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:51.387257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:29:51.387257Z digest=sha256:ebea50fe99dbabd4a73614952d0a7c3f145c0998b5e1ffc0ba469c9e03f9af4e

Observation 199033f3-a787-4436-aa41-9da7e36fc172 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

VidCtx: Context-aware Video Question Answering with Image Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:29:51.628488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T05:29:51.391770Z digest=sha256:e3bc60c8d6f7ec6b623d89341baef8abc7d5969983c331d2be0a18e10260ccad

Pith citing papers

Observation 4dccfb87-3f24-4aae-bd55-8b45679640d5 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning VidCtx: Context-aware Video Question Answering with Image Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.809467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:f4ef962fd87a6e057c99946d89621a9fb699af607ba58ee4974b4f7b96974dd6