Pith. sign in

Paper Citation Record · LEDGER

CAViAR: Critic-Augmented Video Agentic Reasoning

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2509.07680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07680 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:27:36.485905Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T12:10:52.100410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 049fe0d7-f676-4ee8-a644-fb49a454e9c4 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

CAViAR: Critic-Augmented Video Agentic Reasoning Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.487697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.487697Z digest=sha256:611298b4192d77161426003a97aa43b45194f91aeb7a921e80df7ac8dcfde8dd

Observation b25f5195-6d7e-4ad1-828a-26e7aa2bef37 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

CAViAR: Critic-Augmented Video Agentic Reasoning Videoagent: Long-form video understanding with large language model as agent

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:39.555502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:33.532164Z digest=sha256:5b1a87b528702a354c9d69ab8b7255000e6ab4a171e2298fc83088eb5558f50b

Observation eb4beec8-12f7-4875-8dcd-ff4ba7c6fc2d · outbound

This paper cites Self-chained image-language model for video localization and question answering.

CAViAR: Critic-Augmented Video Agentic Reasoning Self-chained image-language model for video localization and question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:39.282949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:33.685027Z digest=sha256:a3c376952ac105bc3a88d3dc2bf41083537d4587f6b930496dcea4888b4fc23b

Observation 4c6df0f6-e239-4eee-b50c-4bf0e9b1d78f · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

CAViAR: Critic-Augmented Video Agentic Reasoning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.767446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.767446Z digest=sha256:6cbbedcf2196e006a4587e2edd40a89747a6958ed7661f0e8f9107b7a1487a85

Observation cbf5aef3-9f65-4bac-b025-27198498e1e4 · outbound

This paper cites Qwen2.5-VL Technical Report.

CAViAR: Critic-Augmented Video Agentic Reasoning Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.848649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.848649Z digest=sha256:a2c896aa34f3dbcdfbebd9e4ea5987d6458eb08ce2a6fa9387e9c03cd3b5b26d

Observation 23865ffd-44b1-412e-9dd3-e5d1537515b8 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

CAViAR: Critic-Augmented Video Agentic Reasoning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.916746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.916746Z digest=sha256:76f80f73b367da87be1334f9b790d30008502025668cbb100682e22f8f2bcdab

Observation f014ec7f-807a-4c89-9dce-1e96c8687105 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

CAViAR: Critic-Augmented Video Agentic Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.985149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.985149Z digest=sha256:2df7fbfeca13c06b25fe974f65d7fc2f7ea457182ffb5fc687b0a9d9e926a077

Observation 3e73f29b-c5a7-439b-93ef-64ed5b7c8f48 · outbound

This paper cites Deep reinforcement learning from human preferences.

CAViAR: Critic-Augmented Video Agentic Reasoning Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.023884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.023884Z digest=sha256:5101c2eb90b4586b9d92a66f8ee14bad1e96d9840f64d03a426fef2184c1b4d4

Observation d3df2e3e-85dc-41b4-b9d9-24229e70ff67 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CAViAR: Critic-Augmented Video Agentic Reasoning Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.093889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.093889Z digest=sha256:928f8a7959ebbae02d3020e56aba8a8a069247131c02419bebbb64b5ee2734ab

Observation c6da55a7-b82c-4eb5-8146-49450b656c1a · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

CAViAR: Critic-Augmented Video Agentic Reasoning Visual programming: Compositional visual reasoning without training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:39.076824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.162507Z digest=sha256:32df33a5cb0cf7cde08524f31d870b7966995e7d05725d89c4229eb5eee2e115

Observation 716ce33e-8ff0-40c2-9706-7aa442423d74 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

CAViAR: Critic-Augmented Video Agentic Reasoning V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.250664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.250664Z digest=sha256:ea801ac7d52941ddf2a793c8bc72780d9ef4eeac979cacbf03123080131760ae

Observation 71085a8d-143a-4a89-a7f0-7ae194f3080f · outbound

This paper cites Avis: Autonomous visual information seeking with large language model agent.

CAViAR: Critic-Augmented Video Agentic Reasoning Avis: Autonomous visual information seeking with large language model agent

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:38.815122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.344479Z digest=sha256:f5d130c4482b2b901f2b57836511b28aaec6bbf37fb0ad70cd33d5ac3385ad4a

Observation 95a9f6aa-4510-4836-bd13-cebc99fc83d3 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

CAViAR: Critic-Augmented Video Agentic Reasoning Lita: Language instructed temporal-localization assistant

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:38.562817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.432709Z digest=sha256:4437154ca6fa5397d0f5fc56b39a58b5c7710e35dac606086e83ef6d939e615f

Observation 3eec420b-cac0-4f66-8a0f-3c9e2c41339a · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

CAViAR: Critic-Augmented Video Agentic Reasoning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.505356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.505356Z digest=sha256:7a7bf91d0af7da962ca15b3d8ce4b42ef66018166139a872e8895a4b635701bc

Observation 95e4685d-f480-447b-a7b7-9fa7fc7b395f · outbound

This paper cites GPT-4o System Card.

CAViAR: Critic-Augmented Video Agentic Reasoning GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.584209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.584209Z digest=sha256:3bce4d9195df38c6151a732e76a2fcc8f77618cfa563296cee9f804a42eaa85a

Observation 50ef2bbd-3407-46a4-9341-3bbdc84bbcb1 · outbound

This paper cites Inferring and executing programs for visual reasoning.

CAViAR: Critic-Augmented Video Agentic Reasoning Inferring and executing programs for visual reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:38.217645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.663781Z digest=sha256:9e1b48379700105205ac2998d30de66e57081dff7bebca848aef4c9a7c4a1749

Observation 023c8a3d-9bce-4b6a-a381-2ccb8e96d831 · outbound

This paper cites When can llms actually correct their own mistakes? a critical survey of self-correction of llms.

CAViAR: Critic-Augmented Video Agentic Reasoning When can llms actually correct their own mistakes? a critical survey of self-correction of llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.720565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.720565Z digest=sha256:d97d2e4ceb07022d31c4edf174ead43ac4b1591f691a4d363b4da65e1efaf375

Observation 5541fcc5-e5a6-4d80-9f2f-79a137d62e86 · outbound

This paper cites Analyzing Modular Approaches for Visual Question Decomposition.

CAViAR: Critic-Augmented Video Agentic Reasoning Analyzing Modular Approaches for Visual Question Decomposition

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:27:36.771441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.772715Z digest=sha256:c0c25c3bddee4cfbb704de87f5371fd497ac274eb4117782375ee649c520822b

Observation 99d14adb-1b33-4738-b1ac-f9cad089bc02 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CAViAR: Critic-Augmented Video Agentic Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.850328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.850328Z digest=sha256:53864628592b3f8bebcd4bb6046d2f168bf1920a54029794be4ac2721d8e74c1

Observation b74080bd-de62-4cc1-a0af-725f2ad4bb48 · outbound

This paper cites Visual instruction tuning.

CAViAR: Critic-Augmented Video Agentic Reasoning Visual instruction tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.906290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.906290Z digest=sha256:45d4d5da0bb120c46e00e329b3b961b1b47dfaba00175b44bd7ad35ccb4aea7e

Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.988625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.988625Z digest=sha256:527b069523c021bba0d42b7ca00d8b6ad3e81e444104ea4ff8e73ffc30cd284c

Observation 880be053-394e-41c9-ab1e-f23542e56d5c · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

CAViAR: Critic-Augmented Video Agentic Reasoning Morevqa: Exploring modular reasoning models for video question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.924087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.082929Z digest=sha256:dc71b65f7f9a710c41ea8c120a7fe31653951bee1dd9717576baaaf2f979fe8e

Observation 63d47165-15d3-4cd5-91aa-91e43c1180a2 · outbound

This paper cites Neptune: The Long Orbit to Benchmarking Long Video Understanding.

CAViAR: Critic-Augmented Video Agentic Reasoning Neptune: The Long Orbit to Benchmarking Long Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.156049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.156049Z digest=sha256:16492bbdb01bb89a18d555c9332b451841d316dd8cfa77270ed110025d045b56

Observation 265e53e3-d135-4834-b8b0-50bf400f7b8d · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CAViAR: Critic-Augmented Video Agentic Reasoning Learning Transferable Visual Models From Natural Language Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.240767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.240767Z digest=sha256:a7b043417ed1a43a3e83ef905a594492ea12c2d20a1174af74d026b30e924b2c

Observation 69a3f923-2596-4a81-87e0-50deecf14655 · outbound

This paper cites Towards truly zero-shot compositional visual reasoning with llms as programmers.

CAViAR: Critic-Augmented Video Agentic Reasoning Towards truly zero-shot compositional visual reasoning with llms as programmers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.693313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.300020Z digest=sha256:bbfe9ab1313dc47c736bdff1c2dba5ce2a1b3a7856866ce86c5686d25e3880c9

Observation 8a3db425-0d2f-4e69-8bff-727904329d93 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

CAViAR: Critic-Augmented Video Agentic Reasoning Vipergpt: Visual inference via python execution for reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.461547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.383420Z digest=sha256:6c1a9b47225385adf8efaacc5effd2cb4fd07aa43f38f104dd80d385189c53a7

Observation db927cdf-c630-4c20-bc8a-a8182517c755 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CAViAR: Critic-Augmented Video Agentic Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.470462Z digest=sha256:3a6db043fc6bb67a6c3d479c32ffe122ef96e60c641e4ec198c60caa2ada4ff9

Observation a690c6b8-d0b8-41e9-96a8-8ebb40ec52a0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CAViAR: Critic-Augmented Video Agentic Reasoning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.557978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.557978Z digest=sha256:4a0f032be1ab866d3c4d8a1c731f5a02d630e96a6eef9a4fa6bebd58c6a5e422

Observation 586c35f5-5129-415a-8c77-c46ab18dac95 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

CAViAR: Critic-Augmented Video Agentic Reasoning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.258238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.656988Z digest=sha256:2a6762fc2a1990ae7ac24f083b6058e4c543dca5abaa326908efecc338eb39c8

Observation bd1bf8a3-3a4b-460c-9550-4b478d18e79d · outbound

This paper cites Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?.

CAViAR: Critic-Augmented Video Agentic Reasoning Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.715349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.715349Z digest=sha256:df311319cd9e2798530ef16f959ce5ad9add638230f8564b9fad7d5767d5f8f6

Observation 528a6cee-34cb-4c96-8bef-8f2c879b270d · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2023.

CAViAR: Critic-Augmented Video Agentic Reasoning Cogvlm: Visual expert for pretrained language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.764777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.764777Z digest=sha256:6b3e1201700b800832bdf807cd8a08fab6dfa508c46f090e67141c43bdf03167

Observation 299ce0ad-3cbb-4ac5-853d-24923866c794 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

CAViAR: Critic-Augmented Video Agentic Reasoning LVBench: An Extreme Long Video Understanding Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.879398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.879398Z digest=sha256:5f646fbb463d20c8fb7a9cc4cbb24788c958f887751616b14539c6f637f22f12

Observation 504c5fee-dfba-40a0-a62d-89d70448a760 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

CAViAR: Critic-Augmented Video Agentic Reasoning Videoagent: Long-form video understanding with large language model as agent

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.058131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.890132Z digest=sha256:971c557ba92a62d01b43341c3fce51a501d0a1a2e936b01746e21281b88864bc

Observation d7f41027-5e6f-4fd2-b83f-dcac4abe59d1 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

CAViAR: Critic-Augmented Video Agentic Reasoning VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.899881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.899881Z digest=sha256:85fe8231e9b850505e94d4090c5d983e278d0f2bbb985592e7ce4904f28cc49f

Observation 4feab88e-c7aa-4056-a943-b782fd904b82 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CAViAR: Critic-Augmented Video Agentic Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.993623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.993623Z digest=sha256:a36a5c90093c244a4902cdc77d1e6022f9134a458a8c103a177d9f593451470a

Observation ee04e35f-2e5d-4c03-a0ba-e946184a1057 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

CAViAR: Critic-Augmented Video Agentic Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.085495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.085495Z digest=sha256:4630fcfabb5a9520e55386beaca8aeb8f071b907aceb7d6dad52b223f9e1db7e

Observation 1d399245-0291-4f27-94ce-e11d4405cf26 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

CAViAR: Critic-Augmented Video Agentic Reasoning Star: Bootstrapping reasoning with reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.221921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.221921Z digest=sha256:2bbe9e355e7c459c5f45cf8f288a55eb998d4a882c7afcdb5ac77a10bdcb8af3

Observation 277939ff-dd2b-4604-af9b-c3836914d89b · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

CAViAR: Critic-Augmented Video Agentic Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.323801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.323801Z digest=sha256:d0a8c5e4777b0dad192bb0fe4a353c50abfab3152fba9cb350e41a9b272876ec

Observation 61073f3f-8553-4916-b6cb-ec7067afd79e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

CAViAR: Critic-Augmented Video Agentic Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.408216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.408216Z digest=sha256:6032a78d9c94a784ab504c9ab443e3099d529684057e8fbd0bcb77b42840397d

Observation f0aa01ba-00b7-4a5d-83f6-8b028f9903ae · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

CAViAR: Critic-Augmented Video Agentic Reasoning Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.485905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.485905Z digest=sha256:02dcf5245e03c2633276c9387561d281626d7e13750494f76990c00f9c97f9a8

Pith citing papers

Observation 8c9e2c2c-f4c5-4319-9144-045b73a1fd8c · inbound

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization cites this paper.

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization CAViAR: Critic-Augmented Video Agentic Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T12:10:52.100410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T12:10:52.100410Z digest=sha256:9a0c07072b41ac7c099e3db190e552479dd14ca930473aee2c305464ed23b9b0