Pith. sign in

Paper Citation Record · LEDGER

CAViAR: Critic-Augmented Video Agentic Reasoning

As of 14 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2509.07680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07680 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:27:36.485905Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:43:38.267412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:43:41.244884Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 049fe0d7-f676-4ee8-a644-fb49a454e9c4 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

CAViAR: Critic-Augmented Video Agentic Reasoning Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.487697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.487697Z digest=sha256:9ffb627f8691bfc86990c52ed3fcf8a49be21ced498f02e2edca7f671cb13a99

Observation b25f5195-6d7e-4ad1-828a-26e7aa2bef37 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

CAViAR: Critic-Augmented Video Agentic Reasoning Videoagent: Long-form video understanding with large language model as agent

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:39.555502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:33.532164Z digest=sha256:00f063fa84a9d8a26c8712f1d8463a379969f888b6bc26f5af719b4e3ab489e5

Observation eb4beec8-12f7-4875-8dcd-ff4ba7c6fc2d · outbound

This paper cites Self-chained image-language model for video localization and question answering.

CAViAR: Critic-Augmented Video Agentic Reasoning Self-chained image-language model for video localization and question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:39.282949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:33.685027Z digest=sha256:b0fdad22721a250ceb342ce1cee2453a4d36ac4df79a49d26935de5431cead0a

Observation 4c6df0f6-e239-4eee-b50c-4bf0e9b1d78f · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

CAViAR: Critic-Augmented Video Agentic Reasoning A Simple LLM Framework for Long-Range Video Question-Answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.767446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.767446Z digest=sha256:1825d559ca5738aff1e7b3dd8f1eeecfe97b12ad1e717fbf668cb4daff53f961

Observation cbf5aef3-9f65-4bac-b025-27198498e1e4 · outbound

This paper cites Qwen2.5-VL Technical Report.

CAViAR: Critic-Augmented Video Agentic Reasoning Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.848649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.848649Z digest=sha256:d6733164ebd386c615ece4effe7713ea3fd25494c632518018625422666d2fe6

Observation 23865ffd-44b1-412e-9dd3-e5d1537515b8 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

CAViAR: Critic-Augmented Video Agentic Reasoning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.916746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.916746Z digest=sha256:e4fc481cf834ce4f839b0664c18ab98f3937dfc62a8a6ab124fa9378f31d93e6

Observation f014ec7f-807a-4c89-9dce-1e96c8687105 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

CAViAR: Critic-Augmented Video Agentic Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:33.985149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:33.985149Z digest=sha256:9377ce2220e786cd25ef6f51a13afafc3e5bb39aaaac28bd64d4be1ef112ac49

Observation 3e73f29b-c5a7-439b-93ef-64ed5b7c8f48 · outbound

This paper cites Deep reinforcement learning from human preferences.

CAViAR: Critic-Augmented Video Agentic Reasoning Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.023884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.023884Z digest=sha256:9ec86dafc16a2702e85784b0582782ce07aad428a5395feaf6dafe4bfbf26385

Observation d3df2e3e-85dc-41b4-b9d9-24229e70ff67 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CAViAR: Critic-Augmented Video Agentic Reasoning Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.093889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.093889Z digest=sha256:054710fc64fcccafcaaaa50877bc221747add56a3924a7b2b4a7969a89a237cf

Observation c6da55a7-b82c-4eb5-8146-49450b656c1a · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

CAViAR: Critic-Augmented Video Agentic Reasoning Visual programming: Compositional visual reasoning without training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:39.076824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.162507Z digest=sha256:7aa5a472861466823946c9f7b722fb6d84405508daf25644a85a29b2f3cc2bd2

Observation 716ce33e-8ff0-40c2-9706-7aa442423d74 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

CAViAR: Critic-Augmented Video Agentic Reasoning V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.250664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.250664Z digest=sha256:33467b0f9698fd2cf5447fe9523566c418ab02e6d55e29d499c8522b1cafa643

Observation 71085a8d-143a-4a89-a7f0-7ae194f3080f · outbound

This paper cites Avis: Autonomous visual information seeking with large language model agent.

CAViAR: Critic-Augmented Video Agentic Reasoning Avis: Autonomous visual information seeking with large language model agent

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:38.815122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.344479Z digest=sha256:e57d62601c3ce6495a45faa85e531941cbaa814498e46790529837d9f8d684f3

Observation 95a9f6aa-4510-4836-bd13-cebc99fc83d3 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

CAViAR: Critic-Augmented Video Agentic Reasoning Lita: Language instructed temporal-localization assistant

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:38.562817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.432709Z digest=sha256:f3dabb89e14ca0a5994cfe07943246a103969dabc58ddd390518a01ab2f3ddb1

Observation 3eec420b-cac0-4f66-8a0f-3c9e2c41339a · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

CAViAR: Critic-Augmented Video Agentic Reasoning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.505356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.505356Z digest=sha256:0ec481f23cfca342fca6cd4486a62f9df867af9c8aedfa66688b77fbd256e98b

Observation 95e4685d-f480-447b-a7b7-9fa7fc7b395f · outbound

This paper cites GPT-4o System Card.

CAViAR: Critic-Augmented Video Agentic Reasoning GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.584209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.584209Z digest=sha256:2a56d4188a07ab2351256ffa9e79d7012ed1a1f7a108ab87a3893618c865cfda

Observation 50ef2bbd-3407-46a4-9341-3bbdc84bbcb1 · outbound

This paper cites Inferring and executing programs for visual reasoning.

CAViAR: Critic-Augmented Video Agentic Reasoning Inferring and executing programs for visual reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:38.217645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.663781Z digest=sha256:235fac1f362a53de67b694be54c82d9c6db1420a8c5c0f90f91ebd9e89f307aa

Observation 023c8a3d-9bce-4b6a-a381-2ccb8e96d831 · outbound

This paper cites When can llms actually correct their own mistakes? a critical survey of self-correction of llms.

CAViAR: Critic-Augmented Video Agentic Reasoning When can llms actually correct their own mistakes? a critical survey of self-correction of llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.720565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.720565Z digest=sha256:5d1901ff77cc77ac09203394fa1e2ed9bb051a9cdb83f51d757041a63c8a0792

Observation 5541fcc5-e5a6-4d80-9f2f-79a137d62e86 · outbound

This paper cites Analyzing Modular Approaches for Visual Question Decomposition.

CAViAR: Critic-Augmented Video Agentic Reasoning Analyzing Modular Approaches for Visual Question Decomposition

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:27:36.771441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:34.772715Z digest=sha256:5982c9b18a6ef1e0d5b19d258a0010f308e7c8bc5e10b725fb76c544fd46c8f3

Observation 99d14adb-1b33-4738-b1ac-f9cad089bc02 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CAViAR: Critic-Augmented Video Agentic Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.850328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.850328Z digest=sha256:8598506ae88c1b55624d622ee696b53c4c81b365d2c923f286e31fe9e081d9d4

Observation b74080bd-de62-4cc1-a0af-725f2ad4bb48 · outbound

This paper cites Visual instruction tuning.

CAViAR: Critic-Augmented Video Agentic Reasoning Visual instruction tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.906290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.906290Z digest=sha256:99f14f3687df8f0e1699608203f8365f5fc4bf798e7ff75c225315e6d7c5ddfd

Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.988625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.988625Z digest=sha256:92c13fdd28a2a41cdbccd83af6734eeb7eb7462b5827c354e18ad6cff19e96dd

Observation 880be053-394e-41c9-ab1e-f23542e56d5c · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

CAViAR: Critic-Augmented Video Agentic Reasoning Morevqa: Exploring modular reasoning models for video question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.924087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.082929Z digest=sha256:995c8592918d5591013ae3bebb3b1bae54524fe9c5990d1c85a5facce0ea920d

Observation 63d47165-15d3-4cd5-91aa-91e43c1180a2 · outbound

This paper cites Neptune: The Long Orbit to Benchmarking Long Video Understanding.

CAViAR: Critic-Augmented Video Agentic Reasoning Neptune: The Long Orbit to Benchmarking Long Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.156049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.156049Z digest=sha256:4317c1b06c37f59b14bf53fe713d992910f80686a9182a1e9c69e8e7f53d909d

Observation 265e53e3-d135-4834-b8b0-50bf400f7b8d · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CAViAR: Critic-Augmented Video Agentic Reasoning Learning Transferable Visual Models From Natural Language Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.240767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.240767Z digest=sha256:541f4396b648072774fe24b07b9146a11e76a7103135feb371477e07f3ebf1c8

Observation 69a3f923-2596-4a81-87e0-50deecf14655 · outbound

This paper cites Towards truly zero-shot compositional visual reasoning with llms as programmers.

CAViAR: Critic-Augmented Video Agentic Reasoning Towards truly zero-shot compositional visual reasoning with llms as programmers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.693313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.300020Z digest=sha256:f4f0dfa342ff23fc99ac75173d5d46ec5466afc8dc775da9c31d3d022a6e228f

Observation 8a3db425-0d2f-4e69-8bff-727904329d93 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

CAViAR: Critic-Augmented Video Agentic Reasoning Vipergpt: Visual inference via python execution for reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.461547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.383420Z digest=sha256:872565a7556b494bf3f255e59e839e6ba15c2d5845ff215d869b8d564d5ad1dd

Observation db927cdf-c630-4c20-bc8a-a8182517c755 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CAViAR: Critic-Augmented Video Agentic Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.470462Z digest=sha256:5e61cf3b0382721d4d46b723355fb2497d397a29df2f4c89ec09f38e29740eb9

Observation a690c6b8-d0b8-41e9-96a8-8ebb40ec52a0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CAViAR: Critic-Augmented Video Agentic Reasoning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.557978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.557978Z digest=sha256:15471e217476340893246e66903b8440acd36d7cb287b7b4ec93ca073569f588

Observation 586c35f5-5129-415a-8c77-c46ab18dac95 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

CAViAR: Critic-Augmented Video Agentic Reasoning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.258238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.656988Z digest=sha256:73465a6cc91a90dea1381423e883e34863ae33907366cf287fcb61bfaa0c5224

Observation bd1bf8a3-3a4b-460c-9550-4b478d18e79d · outbound

This paper cites Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?.

CAViAR: Critic-Augmented Video Agentic Reasoning Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.715349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.715349Z digest=sha256:ccc9e0aa13e336e0dd49fff9ce679754cff42722180a9c8f29878b994ad0b095

Observation 528a6cee-34cb-4c96-8bef-8f2c879b270d · outbound

This paper cites Cogvlm: Visual expert for pretrained language models, 2023.

CAViAR: Critic-Augmented Video Agentic Reasoning Cogvlm: Visual expert for pretrained language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.764777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.764777Z digest=sha256:1ce6c268b088402d4609853c92f5f9a28d1cc3a2bbfc264681f7687d9f57a0a2

Observation 299ce0ad-3cbb-4ac5-853d-24923866c794 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

CAViAR: Critic-Augmented Video Agentic Reasoning LVBench: An Extreme Long Video Understanding Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.879398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.879398Z digest=sha256:00932f78f0173085a1f7eaf30e40df1bef6cf50fe1dee0448dee510fc3c2b4e0

Observation 504c5fee-dfba-40a0-a62d-89d70448a760 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

CAViAR: Critic-Augmented Video Agentic Reasoning Videoagent: Long-form video understanding with large language model as agent

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:27:37.058131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T21:27:35.890132Z digest=sha256:53392032a2a71f14b2e71184137a3de6ce4e1b486555f1d0a99e9d3fac03bf07

Observation d7f41027-5e6f-4fd2-b83f-dcac4abe59d1 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

CAViAR: Critic-Augmented Video Agentic Reasoning VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.899881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.899881Z digest=sha256:4eeb1f1411f6dcf31ab6ec7e294cff3bcc5a85667e6ae824a85f002c9ee514ed

Observation 4feab88e-c7aa-4056-a943-b782fd904b82 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CAViAR: Critic-Augmented Video Agentic Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.993623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.993623Z digest=sha256:e77711104d2fdb36756ec1005b3bce5919b4bab4a9f4aebaaa4b82b60cde889f

Observation ee04e35f-2e5d-4c03-a0ba-e946184a1057 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

CAViAR: Critic-Augmented Video Agentic Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.085495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.085495Z digest=sha256:55b9a5a6cf75b86d84175f5100a36128e6bd889bd7ff964593bd3612a973fc86

Observation 1d399245-0291-4f27-94ce-e11d4405cf26 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

CAViAR: Critic-Augmented Video Agentic Reasoning Star: Bootstrapping reasoning with reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.221921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.221921Z digest=sha256:db9952d95c839120b44089a676f288f102ebdaff6d36b5cf4c2a8a22ba6cf7f6

Observation 277939ff-dd2b-4604-af9b-c3836914d89b · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

CAViAR: Critic-Augmented Video Agentic Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.323801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.323801Z digest=sha256:a6dc731f256792a4b8a9ce48533972c9367e0f298377e6bcda3f575d944f5738

Observation 61073f3f-8553-4916-b6cb-ec7067afd79e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

CAViAR: Critic-Augmented Video Agentic Reasoning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.408216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.408216Z digest=sha256:9cfd830a16baae4b77cc30261e26ede84da3e2f375e28491141575a73632ff22

Observation f0aa01ba-00b7-4a5d-83f6-8b028f9903ae · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

CAViAR: Critic-Augmented Video Agentic Reasoning Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.485905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.485905Z digest=sha256:8a5e06a115e09342a5f4fdd2ab610629277bc6f33e2d050b0a20594d07716635

Pith citing papers

Observation 8c9e2c2c-f4c5-4319-9144-045b73a1fd8c · inbound

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization cites this paper.

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization CAViAR: Critic-Augmented Video Agentic Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T12:10:52.100410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T12:10:52.100410Z digest=sha256:f610aa6f08694c2ff5f4c216c1b876c659053324bb3a78c2ff995d0b1f4ecf3b

Observation 30134ecf-8401-4bd6-bcd7-22d96c8cddba · inbound

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning cites this paper.

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning CAViAR: Critic-Augmented Video Agentic Reasoning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:43:41.250664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:43:38.267412Z digest=sha256:8423c6af0a5916713050c39482ba8338b763c22349d30f480a08df0f22b6fdb4