Pith. sign in

Paper Citation Record · LEDGER

Self-Chained Image-Language Model for Video Localization and Question Answering

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2305.06988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.06988 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:40:35.742811Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.467463Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01864bbb-665c-4f1f-bc48-dbb15c4be93d · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.509676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:a3046790d11b83a979b9d923427bc4d424cc283bf8470d02179aa3dd15e29bb1

Observation 6843ffb7-f852-4166-b2ce-1aee311a5959 · inbound

Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering cites this paper.

Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:35.742811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:40:35.742811Z digest=sha256:eb7e0d5c737048107cab79553381b7139ed2882d23cf5c776d94cec67cba8ff5

Observation 5b099f7e-35c1-4349-a697-ba25452e77cf · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.748814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.748814Z digest=sha256:a793def5b4e68fc17b436f2604d126c7a83737e32bce28d0bc30a3eee46b1ac2

Observation 0a16f30d-b100-4137-849c-bc3eb44ebdfa · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.185509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.185509Z digest=sha256:cb3b0418e7a5016c00b88b5aee8c23a3b46ac9b0609c95d87005886c075a8aee

Observation 1d09069a-466c-42fb-bde3-01b85e1e5ce5 · inbound

Rethinking Video-Language Model from the Language Input Perspective cites this paper.

Rethinking Video-Language Model from the Language Input Perspective Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.468853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T13:24:46.360149Z digest=sha256:33031b37c4367522971c171a4596c881a65af959daf0d49311b67433a95fe95c

Observation cfa677f9-9df0-448e-8108-7469cf311057 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:0b3812910e300d2857dc51ab82eb81c9d16f2463e00fb8eb812ee648f316e437