Pith. sign in

Paper Citation Record · LEDGER

A Simple LLM Framework for Long-Range Video Question-Answering

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2312.17235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17235 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:56:37.715901Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:03.011665Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df1358d4-cea1-4665-8cf1-34ee1e4e3d41 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.044985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:d93a0eb1b16efacd6bd552a5ee376010e26b54dcaeac5fbd2b8959e2df47ce5d

Observation 5fbec168-da9c-4c0f-8ff8-5d4de93b152a · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output A Simple LLM Framework for Long-Range Video Question-Answering

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.857614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:0144896b710d4e19de5e3e7bf169320ee7475588b1b39cb32a1c689060595b28

Observation 1cb44e77-97b8-4fcd-b2cd-d5a96f89ddfb · inbound

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process cites this paper.

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process A Simple LLM Framework for Long-Range Video Question-Answering

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:56:37.715901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:56:37.715901Z digest=sha256:8bdd95308e05a3dcaeb343bf6b57f21faac123dfe358ee5d1ba3b15a8a712cf6

Observation bd39f4d2-65c3-45ea-a8ca-545b572d313c · inbound

Understanding Long Videos via LLM-Powered Entity Relation Graphs cites this paper.

Understanding Long Videos via LLM-Powered Entity Relation Graphs A Simple LLM Framework for Long-Range Video Question-Answering

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:27.471264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:53:27.471264Z digest=sha256:c13d26197d7b5e2fcc73fc9be2f36a64ec2884e19f75ff9760e2dae1a96dd749

Observation d88c3ee7-3d28-4528-8fc2-6a4f70996126 · inbound

Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles cites this paper.

Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles A Simple LLM Framework for Long-Range Video Question-Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:53.167691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:53.167691Z digest=sha256:9f6dcd0748c0d1405c3eb3d5bd165c7d3f56883958d835c61764de94707c9278

Observation 89b6721f-3457-49b3-88e1-24ab7fba52b6 · inbound

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 cites this paper.

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 A Simple LLM Framework for Long-Range Video Question-Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:47.594214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:47.594214Z digest=sha256:b330ac90bfe03d91ab2915160adecb05d35e6ab4aea1fa9bb230397120f59914

Observation 0cd89e7c-11c0-45cd-8bc6-4f53a320ddf1 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:26.732647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:26.732647Z digest=sha256:13456b11ca9c0cfef4ff6a7db6b30e7194aa39178d3cc3dea637dbdc5ee632c0

Observation d6452b5a-89f2-4f23-9ca3-fe3a857482c8 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding A Simple LLM Framework for Long-Range Video Question-Answering

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.959356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.959356Z digest=sha256:91ed48a498b17aef1704b2b684f3c829745716d553362dded4cbb0514ff5aed7

Observation cd904f4d-844d-430f-a3a9-23449a5f5372 · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:36.153446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:36.153446Z digest=sha256:5bf44d03dd67cbb0a9ff2f9c5c5c548d4bfcd6f0882c45d886e525447dd43ff9

Observation 34e11361-29a3-4833-90e7-58ddcf13c67e · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.312571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.312571Z digest=sha256:dd5cb7df95348a85be590f38e8a0a61d37bf13e3b916b38537aea625cdc08d81

Observation 4c6df0f6-e239-4eee-b50c-4bf0e9b1d78f · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.767446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.767446Z digest=sha256:efac15fa8e75226d2c1c4b38e1913e132db9c7bd8ff095f4e020d649f02e9373

Observation cea25f6f-a7db-49e6-bb95-189a99b560d1 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents A Simple LLM Framework for Long-Range Video Question-Answering

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.621623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:8e1589423afe4614bbfdba69143b428adeb6d36e11a88f1b33fdf1ee44b974b3

Observation b693d7da-96f3-4119-aeba-636ba2d90386 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.682531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:d4fd976b38c66c3490a89583afa921516070883a331bbf20ad313c01ee3cd5f4

Observation 39f82adb-aefa-44aa-b2f5-9b44fd964678 · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:16.311655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:16.311655Z digest=sha256:bbcbd088c92fc65a214d999acfa0f60e218021222b68228f6b17be8e7e22d621

Observation 290ab08c-93d4-4de8-85c2-b2ab3dc664c7 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding A Simple LLM Framework for Long-Range Video Question-Answering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.889680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:dffb70f371168ddf64494bf1018e58957723c8cc923fd93f0952e75dad9a3b50

Observation 488288cd-eb57-4fc0-a317-92bbea7afcb2 · inbound

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks cites this paper.

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks A Simple LLM Framework for Long-Range Video Question-Answering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:13.224993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:41:42.061219Z digest=sha256:0c685dfeea01d1ea75898d088c97e7a6716da0dff9fcf7c8568e1114a3ae3d11

Observation 058c212c-0e18-4f04-9a59-49c33874108f · inbound

UNIVID: Unified Vision-Language Model for Video Moderation cites this paper.

UNIVID: Unified Vision-Language Model for Video Moderation A Simple LLM Framework for Long-Range Video Question-Answering

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:07:09.283746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:56:26.674841Z digest=sha256:1891dc3b617e517bb04db39b9d3ae3e9bf76172072c6f7422b4c46b3b64a99e9

Observation 4760a439-14da-41c5-b5d7-935380a77d67 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.012988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:3ebd2d49446d3e856371e9914a0624262f8b228309a3bfceae1e9d1695045261

Observation 5edf0bcf-564c-4d65-9551-ccf929071349 · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use A Simple LLM Framework for Long-Range Video Question-Answering

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.233670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:1161dc8dea0dd2053a8f15406d20cedb27430bbec168804dfd70fde9fada8c3a