Pith. sign in

Paper Citation Record · LEDGER

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

As of 14 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 6 inbound Pith citation observations for arXiv:2411.18203.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18203 v5

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:27:33.494574Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:56:01.975427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:16.802423Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved55
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 034a4a1f-25b6-421c-ac16-9da194944231 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.082475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.082475Z digest=sha256:cf77a6b5a96c135d606383e1e37f89a038b881b5ce064477579ae7b3ec28d76c

Observation 8d605576-d546-41e7-a892-58bb37b3040a · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.088452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.088452Z digest=sha256:e5c3f76cdee4e4280c799063218191d20957a4afca73c7611be9f7e51458586b

Observation 1f148c67-6662-4cdc-bf32-6d7fac83682a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.094064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.094064Z digest=sha256:7156c5d024c4d0c411947e715b4add60de1d03154cdb931d7d229e4c5b44e677

Observation 2167781a-0df1-4004-8d67-5b91be868b77 · outbound

This paper cites A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.101376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.101376Z digest=sha256:62dd24d3dc638b9c0d1057cc818865835559520ade44ecfecab3c229681a3fa7

Observation d8221068-bef4-43bd-b7a2-42b5c0ed0d0f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.106768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.106768Z digest=sha256:9219d71ebd29cbcadfe901e7722e5adfabc6e1c315bd56dd22e43a90858f39dd

Observation 9d0e1993-dd2c-499c-a8e4-8e3114a8f8e8 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.112655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.112655Z digest=sha256:c2a143d6566b6f3a567e34749f0d7bb7d0cfd3d2fe12e027c91c25d8dde089e3

Observation be18d887-8108-4785-aca9-1caeec776a58 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.843995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.118009Z digest=sha256:99a1c907a45c966a312bd69fec2610cba39be932300d2d32a6a78f181f564c1a

Observation 92ea7bac-5f81-481e-a256-205d9dfafc9d · outbound

This paper cites CAST: Cross-modal Alignment Similarity Test for Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning CAST: Cross-modal Alignment Similarity Test for Vision Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:27:34.137559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.122826Z digest=sha256:7ff5115739a1952d7e0e96465b707a1477bef0901a35a5e965aaccac0a091d1b

Observation 979b5fa9-23a3-4333-ae07-42be3cf973d2 · outbound

This paper cites Gemini-1.5-pro, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gemini-1.5-pro, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.826294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.128478Z digest=sha256:5381a1a280f4ae78f29e29197a308c7465c33a50a2b12db76d9b38c1f9dea18f

Observation 95a58b8a-9adb-4645-9eb8-49aa8962c5b5 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.133335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.133335Z digest=sha256:4a0c278889c1bbe7321d357d249318278d87c9eeaec4bcf5da274c7cb68a84d9

Observation 0b954e3b-be66-4f36-a2ed-e11590f3c993 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.138054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.138054Z digest=sha256:638904bec80a4bb5dadd078fd6681bdb707cb275e0ac3860730b39fd2e248585

Observation 6be69ab7-8e1b-4659-ac89-bf5ecb436ff1 · outbound

This paper cites Chatglm: A family of large language mod- els from glm-130b to glm-4 all tools, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Chatglm: A family of large language mod- els from glm-130b to glm-4 all tools, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.809098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.143614Z digest=sha256:9cd2ffd4231fea62c044d9d818b827d372613f4ef7ee1ddd6541f9ca87a57f33

Observation c9079af9-cc7f-4116-917a-db2bbdc4aea0 · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.148992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.148992Z digest=sha256:e7e0d6452314aca6f92f65c425ef628c01b694f32999177f63cdb94a1f4e8bbb

Observation 3731d655-9dc0-467b-a652-6bc6913c5e71 · outbound

This paper cites Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.153845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.153845Z digest=sha256:18bbd35680a06f196b5b3d3f6cdc9831d010436f048d61d7e766eeefc9a24d9b

Observation 5a8f6af5-373d-4af3-b4b6-b2e5fd8a749e · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.158423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.158423Z digest=sha256:98da2303c2a14906391aa966824c0234e552b97477b1ead156724dfb58af4e56

Observation 6860837d-de31-4a5d-947d-ed08abc4ad61 · outbound

This paper cites St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.163296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.163296Z digest=sha256:c1cbc81b48e0d062de3eb5651a33baf7d2ede84759b0a80bc3582b1483ccecc9

Observation 954f6cb4-c8c7-4c1f-9733-fc183994a71e · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.167657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.167657Z digest=sha256:b15d0db1c7af9590176a7ab1c85e9dc6ae5fc9736161ae16dec40c57e0e191e9

Observation d5c2c2d9-f8ab-423f-b1b2-7c8658dd7850 · outbound

This paper cites GPT-4o System Card.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.173463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.173463Z digest=sha256:570a4a22ab3d36b50c298fac87c6410ec085ed79ecc7a56f4fdf8b489ccab05a

Observation efff575c-602d-41a1-8d0a-52067462f1e5 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.178940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.178940Z digest=sha256:3fae3f297c1d60ae926058003b9b7b5b2ad49666147cd4d6be36d239b9de08bf

Observation 31478b72-4aa0-4807-8de0-ab17a6049e56 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.183592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.183592Z digest=sha256:8af4b367a7310e0751d23aef9b49157c8e840f987ef47d9a5147e5c9b483771d

Observation 8b38f95b-ee23-47cf-ba57-cae2dcaf62b8 · outbound

This paper cites Large language models are zero-shot reasoners.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Large language models are zero-shot reasoners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.756313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.189253Z digest=sha256:2ebd5cf0c4442df63c9944e1a225030e54aa3ec629c5b46a5fb2048f7ec0eedd

Observation 6bb5eb8b-6966-45ff-9ac3-634d3d9238f2 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning In-context Reinforcement Learning with Algorithm Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.194094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.194094Z digest=sha256:071365b96213812528056ec4b76e9f025a2081acdf074fa9e9835c75d3996d70

Observation 1e646ab3-d26a-4239-9865-26d51064c042 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.199026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.199026Z digest=sha256:0de4456e45cf87041477ed7ba8fb37df92cde9a16604e1b9e74cda4c2575bac7

Observation a01a592e-44bc-406f-a3c8-a5ab89e09ed0 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Silkie: Preference Distillation for Large Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.204022Z digest=sha256:f485ea71b70641a835b7aac8dddb05df3feee0939f0d40feade78a3da152842a

Observation 9f764ecb-cd88-4e22-9e05-d9beeb962bcb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.208990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.208990Z digest=sha256:64713dbe359c3c12ed4e811d41cf4e77926ba7dff9a0ecc4091c8bd3b4223f57

Observation 52f9c951-d830-489f-9ce0-fb42c12adb18 · outbound

This paper cites Let's Verify Step by Step.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Let's Verify Step by Step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.214333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.214333Z digest=sha256:6327631b7c3d92412aa66368a1641ea17b599704f826438d420d95a7d86d8aef

Observation 048b293e-b7d8-4666-bcab-cd50cb634d64 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.740575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.219132Z digest=sha256:b30fd6f39eda1e6d03338269f47c956fd1dba7c93b3f53223ac278b3fd21d9ce

Observation 28aa8fb3-447d-4c19-9214-0325a9f0d193 · outbound

This paper cites Improved baselines with visual instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Improved baselines with visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.223939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.223939Z digest=sha256:ff2e2dbf7940314df7645123000313e0d34b9697fead91e497c3191496bba665

Observation 8bf43c82-feae-46e3-949b-2f7e1a3902c6 · outbound

This paper cites Visual instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.228945Z digest=sha256:babfadd55d4a05cfe7566753832129c5efb0f84d2ccb932e8445f31790a1154e

Observation 422cb8f1-5374-49ab-8131-14a6d80f8491 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.702876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.233647Z digest=sha256:a2c2fc8648a38fca6781f1c7f011a7561782bb45bdcfb1205e47f06f490a08eb

Observation a0b62837-e951-479b-bf63-bf0766baa20d · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.238871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.238871Z digest=sha256:107a9175b43723a4887a53fd3894a2e3d3bee094f4e0558da45099e07bdffb23

Observation 06df21f4-afa2-4141-b501-5b7627c19648 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.685242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.244301Z digest=sha256:841725517efb722096d1d28dc9d70dec007f202068b14574155ab36794b22daa

Observation 87ede84d-8f10-41bb-a46d-ac44467b0c45 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.248747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.248747Z digest=sha256:851f5912fd821b19be8b8c6f17947169f13ddb655f4a7fc58a32383da2739535

Observation 6495c531-cf02-4d26-9b78-8d16ed8d249f · outbound

This paper cites Self-refine: It- erative refinement with self-feedback.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Self-refine: It- erative refinement with self-feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.652549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.255007Z digest=sha256:550b598d2c3da2a5006ede960fd29c546409e15aedeb55b92eab8550d1d87aaa

Observation 8d63ecfa-947c-4889-b4fe-80eec2bde86e · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.259690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.259690Z digest=sha256:cd44b38de2502f23dfa271f179d32af6bd595ce972d4642bf4d175997f7bc8b6

Observation c2bc84c7-b20a-4dd5-bcf5-f641e5a9f4e6 · outbound

This paper cites Llama-3.2-11b-vision, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Llama-3.2-11b-vision, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.633589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.264863Z digest=sha256:5aecd1e31234d107d1e7715b1a6ea9ba930ca69a2658354ba3ba20a2decc2bec

Observation e3d8be13-5f55-4c14-9033-84d6af8a8726 · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.269702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.269702Z digest=sha256:c1684e0c87474072ea5e1bb8e4ca588bd3b65b2053591ba83a96dbb07242b029

Observation d1ecdada-f367-49fc-b211-fecd9b6778f8 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gpt-4v(ision) system card, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.615390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.274830Z digest=sha256:45863f6fcd39b9ae382f8f1738e5e9c440d17fc10bc97907b1c04ed196bed40f

Observation ba6bca92-4df6-40ab-8c03-ec2b0de626d0 · outbound

This paper cites Hello GPT-4o.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Hello GPT-4o

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.597694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.280224Z digest=sha256:9763db8ff943ba3efece42947101bf83c90303aefae62149284f56e433100a69

Observation 4ab1b5ae-9ce7-4de7-bf41-ed108ee4c08f · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence,.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gpt-4o mini: advancing cost-efficient intelligence,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.285300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.285300Z digest=sha256:a94ca47dbae0bffb068bce771f1310b19655a71cf751234c75999bd1886466c9

Observation 58cc6361-bb9f-4fd6-a1d4-e57357a9388f · outbound

This paper cites REFINER: Reasoning Feedback on Intermediate Representations.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning REFINER: Reasoning Feedback on Intermediate Representations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.295813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.295813Z digest=sha256:c4471cdfb624463c73d2ed91bafe9dddf22fd51c46e7fc273f5a7f048cbe3ffe

Observation d29ed5ec-50d7-43f7-9572-3f24f6753b9d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.552174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.301581Z digest=sha256:3eab6b805e17cdb9a0f7e0d8afce02a0c40e69fae3a0fdfc5343adda7423b6bd

Observation 78046100-131f-421f-9de1-8f0bbfb6e49c · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.533861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.306734Z digest=sha256:f5caa37ad9f99d31fcb2499c7af6241be5ab676115b8a225f74a73baaf92582b

Observation e0cf320e-7c19-4f4f-bc7b-83ce7c498fbf · outbound

This paper cites Learning to summarize with human feed- back.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learning to summarize with human feed- back

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.517214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.311992Z digest=sha256:3b5eaa5aa0810eb2591e43a14302488c1c8cdf086ac34a805e6dd8884954dd86

Observation 3c94c04f-de9e-4a45-b91e-7bfca6409f05 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.316975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.316975Z digest=sha256:2990f8eac8d87b53efa60db4a220c154c53e0951017541882ffa9ddb4366fcca

Observation 266575a5-ea79-44cc-91ad-8395ae060719 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Policy gradient methods for reinforcement learning with function approximation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.498472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.321810Z digest=sha256:7d1a81825523aa2c1092dba4b16f803952ac5726d6339846a054f293b248e75f

Observation f02a630f-3b4c-44b0-83de-099917c43c33 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.326629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.326629Z digest=sha256:72de4180f276195c9fd7003f8978f7acf5751424821170611658c05d65bb3de9

Observation 2bbdabe7-3460-4516-b7d9-ab3ac4b6feea · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLMs cannot find reasoning errors, but can correct them given the error location

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.331621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.331621Z digest=sha256:99ff452454d0af74166073172e371f264a2e9821a31811b88258cda6fc0b8628

Observation db7a87f6-8cf6-4d55-818b-9e8a057a47ef · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.337313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.337313Z digest=sha256:b9bb544538d1628d90b8a6cbf10e6d500dae335c4cc0066366b35bacacb87e48

Observation 17cfad8e-0140-4bac-9f76-730c358064b0 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.342823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.342823Z digest=sha256:05b7743fc73b88d37403efad3863812506110e84a87f571ba661310d4f2ec517

Observation 5b9e669a-383e-4016-9fcc-099afb58f0ed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.348467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.348467Z digest=sha256:c361c3afec72fd9b3b8ebe03c82b0361eb6d211f077585eb41bca923d5076a56

Observation 917ed366-babe-4e96-b341-d5776162cc89 · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.353133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.353133Z digest=sha256:58680cf8e1c60638fc746f027396c2e19e2cd09085edf3251ccf7772ab993505

Observation bfaeb753-cdcd-4070-aee9-6197022b293e · outbound

This paper cites Grok-1.5 vision preview, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Grok-1.5 vision preview, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.464735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.358203Z digest=sha256:d2df238c754d72a96530c6f47389f1931ab60a8b5ed8ca9fe1d9baa284fc5070

Observation 710f667f-063a-44ce-8c8e-4465ef4a39c5 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.363118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.363118Z digest=sha256:31595ceb8c29bc27424c1dd4ff3ce154a5429136d842ef3103e980788be487b6

Observation 1087fa5c-e965-4048-83aa-9e4506dd8fe0 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.446738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.368550Z digest=sha256:85524344c99c7f703a7ce9ada4334f98f71705ec0c260b5b36fc82544917763f

Observation 0625d111-4cbe-4aba-8ce7-1f3acee4675c · outbound

This paper cites Learning From Correctness Without Prompting Makes LLM Efficient Reasoner.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learning From Correctness Without Prompting Makes LLM Efficient Reasoner

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.373868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.373868Z digest=sha256:6fa4a74a99d66d00758e9f8083e887d162b147d2ea0047070fff7e119f630414

Observation 10c014a5-3ca3-4a3a-b5d6-c94b48c9adc5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.379680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.379680Z digest=sha256:068516e3b2abd96df61e496c49d557a33d47c5108da355d4f6737b7f626791d8

Observation aa5b0aad-32cf-48c4-b16c-0316c66d5e4d · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.428259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.385690Z digest=sha256:432fb15f1578c13db1e5d8b5ca7870134d269a9fc9ce5627a8f3d2caa672d33f

Observation b9a1de13-d079-48e7-ad42-73c40f21d99b · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning TextGrad: Automatic "Differentiation" via Text

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.390355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.390355Z digest=sha256:005e723610225e4960c95eb5ead91795f8e7d2baf096cfbf7661911cab790318

Observation bde391a1-1fb1-4797-91be-ec23c524300c · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.397485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.397485Z digest=sha256:23fb4a4f7104d45acb47b89fb685a6e69bcd76f30839676c445025e95ec76eba

Observation 4c064ee6-a811-4e7e-87f3-006990c4a0b1 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.403459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.403459Z digest=sha256:aac286bf42688582403fc8ed1cf19d3e9a9413c56c800c7b0522ad88091d05d6

Observation a8d72bef-65e9-4fce-9748-e725d2b3ee7a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.408895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.408895Z digest=sha256:29e34258b5019b8826cb84ecb99c8536a9ea086ed9dc0e6049b69998c28e7eb8

Observation b0448ac7-50ff-4f4e-92e1-a6e1e4d13b76 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.414817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.414817Z digest=sha256:582a62f9876e4c6998aa6d00853e4abce9358814c4b22c862bd0c48c883eb198

Observation 191d3506-af80-45c3-a24a-ad877de06601 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.419523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.419523Z digest=sha256:36c493ad2cba3b075504bdf0ea8bf3d3b7699afdadfad033abeb0cb2086d2e1b

Observation 763db4c7-4943-45db-9dd3-d365522f91a6 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.424887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.424887Z digest=sha256:4b97ebea3077b581f4af0d726b02e0de71cbbead3bbd5feb32300adde05fcb9c

Observation 2cec8827-945a-465a-b57d-ed01b74ed74b · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.430241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.430241Z digest=sha256:ea430e4682983c2fc7e59febcbb782ab624957f8f534210d4100e794fbd30b5e

Observation 14067b64-e3d4-4320-8f78-261170d301f2 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Calibrated Self-Rewarding Vision Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.435515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.435515Z digest=sha256:16ec4a416cdb51f32282b7b52e6272f1017e0ba60912e5f8e8bbbf6be5aeff2c

Observation 221d1c84-03bf-489e-a144-4c5434667161 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.441121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.441121Z digest=sha256:181ff56f2ba7515fb89eea37b3b498f8a64a7c86a889ec10e3945d029eae5c7f

Observation 6a56b08c-c0b5-4c9d-9001-1f9b8121916e · outbound

This paper cites VGA: Vision GUI assis- tant - minimizing hallucinations through image-centric fine- tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning VGA: Vision GUI assis- tant - minimizing hallucinations through image-centric fine- tuning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.410929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.446544Z digest=sha256:c78776990cdedda88bdd9a2c0c2bf44c08c298d73a7369f5da85dd17ecab9952

Observation e3b08c04-0c7e-447f-baba-6f68ca9d0e44 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.390545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.451714Z digest=sha256:a43a286b962be9c1c7fbc9d2f6ca363490ba24286ee18d016be2c8b7b2ded089

Observation 99993b3a-981e-4019-bf0f-125bb30f6618 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.374596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.458625Z digest=sha256:7873729cb5725c36ed7531c68eccc4103a02b0d4b51d4e23454e225a9c1a5e8b

Observation f0b1f474-4668-44e5-b06c-db83442fcaa3 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.356708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.465299Z digest=sha256:30eddb92a7e1c4d41b47b7e406e17b028872f3ca813ca510d095d05a0e849785

Observation 48cc2e8b-a86b-46e1-a840-d6ab3448261e · outbound

This paper cites For preference-aligned fine-tuning, we utilize Direct Preference Optimization (DPO) on 29,012 samples from the critique- VQA dataset.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning For preference-aligned fine-tuning, we utilize Direct Preference Optimization (DPO) on 29,012 samples from the critique- VQA dataset

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:27:34.332142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.471122Z digest=sha256:65a03b2becc812edf7dcd809913900c5f1f98e5a1137cdccc5583110cc49371f

Observation 19a48ded-338b-493e-a1e1-4d160a218afe · outbound

This paper cites In this section, we will list out the hyperparameters we choose for evaluation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning In this section, we will list out the hyperparameters we choose for evaluation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.311708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.475936Z digest=sha256:b47757de427aee1322240efa889378061c8b982ff51ea091b06c7aa43eba64e6

Observation 470ff0d4-673e-42dc-94da-807a7963785c · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.295105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.480557Z digest=sha256:9134038fa3a29329b46de9e7b7823ffbd63d4a484bf21d752cf07ad1dfb4c207

Observation 1e61a02c-5122-4f37-a31c-9e03ba72aab7 · outbound

This paper cites You can find them in Figure 6, Figure 7 and Figure 8.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning You can find them in Figure 6, Figure 7 and Figure 8

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.274023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.484764Z digest=sha256:4444f993bc7c8f83399dbac7cef4bdf8635f78144652560bbbc7a307acb283f5

Observation 1a395061-503d-4360-aaf8-4d1ae46dc235 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.253077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.489216Z digest=sha256:5e9e7d923ac5099e3aef2045ca672c49dcb3daa8bdce45f162cc74a85c7ec812

Observation e39a1e81-85d4-441c-b558-f7f657f6c09f · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.236407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.494574Z digest=sha256:88f97f16c0a701472d47c135880592d274b9c81ffaf070eae42a850d89b00072

Observation 84196f46-7d29-4e07-98c6-98ff17b6ef59 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T11:27:34.569322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:27:33.290232Z digest=sha256:04dabb32b245bec150a10e6fdc51808a34d426b9c2ad6c5c03ff9ce3cbf3b1fa

Pith citing papers

Observation f7ccdbab-a1c1-489f-8523-1d22700a6232 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.308563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.308563Z digest=sha256:310d223e8c0cd89106f640fb961205d6152bc8cce636a06a7a8ce4a27d796ea2

Observation 3dd5afbd-cc1f-4a06-883e-04f55644e95a · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.875185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1d48100ee5905dcdc2d67f1331fc0fae5a3084ad9bcd5b937a6628494a684249

Observation f0639fcb-0b98-48d9-b284-46d6eb8d4c83 · inbound

Test-Time Hinting for Black-Box Vision-Language Models cites this paper.

Test-Time Hinting for Black-Box Vision-Language Models Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.061496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T21:16:45.748116Z digest=sha256:a5d3176fbd5e1cb58aa2a0736b13738c41803ce33c3efee389c240826ae6d264

Observation c33ebc55-f171-4fb7-a660-b1458662dde5 · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.803674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:033cab30f55296e24695b468f387cc758025fb0cd5e97b3e5c10724ce2c075dd

Observation 534359d2-0f71-4412-8ef1-5ec5e2c6d94e · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:16.664589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:16.664589Z digest=sha256:a98be7a2ef422bb9b9dbf10d432c274da8a89a1dc78ebde229b906feadb48856

Observation afc8d2b1-b606-46cd-840f-99b58920453f · inbound

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus cites this paper.

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:56:01.975427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:56:01.975427Z digest=sha256:38779db92ed65057e0f069a7453be73310c3be3df11cb1bc6c948d918ea19198