Pith. sign in

Paper Citation Record · LEDGER

Self-Chained Image-Language Model for Video Localization and Question Answering

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2305.06988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.06988 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:16:02.785135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.467463Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01864bbb-665c-4f1f-bc48-dbb15c4be93d · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.509676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:9493b9bb088554dfbdf65f067d7f5989e14a9dd1e68dd759ee30400ece80a124

Observation c685a856-19e8-443f-a769-536199029640 · inbound

Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios cites this paper.

Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T16:16:02.785135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:16:02.785135Z digest=sha256:1fff8e6097385f6c3ca245fd26d3440d51dd1bb53b26985487b9ca4052d08197

Observation 31b2cf2f-876d-408e-9f13-a40eed83544e · inbound

Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark cites this paper.

Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:46.444396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:46.444396Z digest=sha256:0d0a391c969d5b4a8a01ab624430e11c5fbc55ec486c377d428d766f730d6c4a

Observation 03d7fa82-70c5-4433-a111-b76be3f778ac · inbound

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection cites this paper.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.204282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.204282Z digest=sha256:30254698b6936bdc74a5a8ffda158349728333afe49c0455adffbe5fd813f837

Observation f97690b8-ccdb-4431-9aae-c11b6571bf89 · inbound

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process cites this paper.

ReasVQA: Advancing VideoQA with Imperfect Reasoning Process Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T15:56:37.706498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:56:37.706498Z digest=sha256:96b298e23c12ed66132ead4aa4f421a9b1554a13fad6486fcdf00fe5b896a910

Observation 6843ffb7-f852-4166-b2ce-1aee311a5959 · inbound

Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering cites this paper.

Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:35.742811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:40:35.742811Z digest=sha256:45ebbd2b9e17932e305c616ca706f7ee25e19b080d04ebff968b63bfca910ffd

Observation 5b099f7e-35c1-4349-a697-ba25452e77cf · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.748814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.748814Z digest=sha256:bcea23224b71f8a4e22d0f3d08c724ceb03eb08157ab015cd8d9372952c9b4d1

Observation 0a16f30d-b100-4137-849c-bc3eb44ebdfa · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.185509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.185509Z digest=sha256:16599a0cfe5af4a73abac97070e7ed0da2f02cadc44e2ba573fb6519b231acc2

Observation 1d09069a-466c-42fb-bde3-01b85e1e5ce5 · inbound

Rethinking Video-Language Model from the Language Input Perspective cites this paper.

Rethinking Video-Language Model from the Language Input Perspective Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.468853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T13:24:46.360149Z digest=sha256:6c4736416b3e56380f7860e402a43d237ef7218486bd27b522e0902a184e71ca

Observation cfa677f9-9df0-448e-8108-7469cf311057 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:135d4e8f7a1550fc264d3050537e019f3baf8500f5c447a914421e608624d07f