Pith. sign in

Paper Citation Record · LEDGER

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

As of 13 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 6 inbound Pith citation observations for arXiv:2411.18203.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18203 v5

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:27:33.494574Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:56:01.975427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:16.802423Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved55
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 034a4a1f-25b6-421c-ac16-9da194944231 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.082475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.082475Z digest=sha256:ee25c0199a92dfc940eefd7746392afb3584fa8cd8251c62e13ab824006489b3

Observation 8d605576-d546-41e7-a892-58bb37b3040a · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.088452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.088452Z digest=sha256:12b995bb5168057e32cae9696c7937151b3742753e5e192e237022a6f4b8abc3

Observation 1f148c67-6662-4cdc-bf32-6d7fac83682a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.094064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.094064Z digest=sha256:580f65d747de2eaa6925e9f4385d8afbb210c59b0c0c9c3252c094633c9250ad

Observation 2167781a-0df1-4004-8d67-5b91be868b77 · outbound

This paper cites A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.101376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.101376Z digest=sha256:d3a7c110625c3b0a47c13e77302d6a64e1b0ad45b09246b1c9428cb92c33b42a

Observation d8221068-bef4-43bd-b7a2-42b5c0ed0d0f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.106768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.106768Z digest=sha256:eb70caea1072ded28da0dcafe92bae1bd69cac3eaddf78843e6d86b0a3f5f01b

Observation 9d0e1993-dd2c-499c-a8e4-8e3114a8f8e8 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.112655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.112655Z digest=sha256:b86f13c6b34e3b02cb2a09c28e0857d25d4da4f14e636fa5d97db219e5f803ec

Observation be18d887-8108-4785-aca9-1caeec776a58 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.843995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.118009Z digest=sha256:47babe4a6975b7763358a5b3085d98d03aab7153f93ad8e2ed6e7a6ff12299a8

Observation 92ea7bac-5f81-481e-a256-205d9dfafc9d · outbound

This paper cites CAST: Cross-modal Alignment Similarity Test for Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning CAST: Cross-modal Alignment Similarity Test for Vision Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:27:34.137559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.122826Z digest=sha256:8b1b01c5b7e8fde5b8f2b951359faa27111cd1a087cf5b2dbdf484c1dd24b038

Observation 979b5fa9-23a3-4333-ae07-42be3cf973d2 · outbound

This paper cites Gemini-1.5-pro, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gemini-1.5-pro, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.826294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.128478Z digest=sha256:a523aa6344a30740b19293ea106f325cadeabb50fc8a2e8b7fa162e82a72a219

Observation 95a58b8a-9adb-4645-9eb8-49aa8962c5b5 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.133335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.133335Z digest=sha256:67ec1f8629a066044bfdbca784a88b49b2557f72aa46ae0c33cc4b89c9b09aa4

Observation 0b954e3b-be66-4f36-a2ed-e11590f3c993 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.138054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.138054Z digest=sha256:fefb9eb1abb57eb234d3f6909f5d6f62d1d7e1a649d9aead1eb04c3fd80d5388

Observation 6be69ab7-8e1b-4659-ac89-bf5ecb436ff1 · outbound

This paper cites Chatglm: A family of large language mod- els from glm-130b to glm-4 all tools, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Chatglm: A family of large language mod- els from glm-130b to glm-4 all tools, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.809098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.143614Z digest=sha256:1c7ac0f4163f336ec6a52ba55da72d3aabd2b3e5db32ac15ffa59338bcd672c2

Observation c9079af9-cc7f-4116-917a-db2bbdc4aea0 · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.148992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.148992Z digest=sha256:688f8449f3ccce4469aa0f54517da5f7076ede30a6dfa9d5dd3722d8fc73893d

Observation 3731d655-9dc0-467b-a652-6bc6913c5e71 · outbound

This paper cites Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.153845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.153845Z digest=sha256:52d50ed05279d4af4e52a337442fd72d4ac7614a1deb26f36d990071deac7a96

Observation 5a8f6af5-373d-4af3-b4b6-b2e5fd8a749e · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.158423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.158423Z digest=sha256:e386bdc5318236864054380e162b39fdd727d0fc527e00026a255f7251661a1b

Observation 6860837d-de31-4a5d-947d-ed08abc4ad61 · outbound

This paper cites St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.163296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.163296Z digest=sha256:26004437ff1650c477ccae4f75b2fbe088f2810d70422dce0a44eec714b82be4

Observation 954f6cb4-c8c7-4c1f-9733-fc183994a71e · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.167657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.167657Z digest=sha256:76c5dbb1461939194f5f288da476b28ae7ce0959a9ed27be0875cfc5017cbd25

Observation d5c2c2d9-f8ab-423f-b1b2-7c8658dd7850 · outbound

This paper cites GPT-4o System Card.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.173463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.173463Z digest=sha256:97db62616cb549f23f86a0d7ef9beb88f98180b7201acdd24c6dbd71be89c3db

Observation efff575c-602d-41a1-8d0a-52067462f1e5 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.178940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.178940Z digest=sha256:a8dd4cf7ef18913e8cc712eb74c79beaa9e59b9a05d373c4b6a3039a960089a5

Observation 31478b72-4aa0-4807-8de0-ab17a6049e56 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.183592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.183592Z digest=sha256:578a95ea2099b4de7808bcef1956108056208e9529d0611e0c772698d5d37df5

Observation 8b38f95b-ee23-47cf-ba57-cae2dcaf62b8 · outbound

This paper cites Large language models are zero-shot reasoners.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Large language models are zero-shot reasoners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.756313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.189253Z digest=sha256:e49cdfc8410f31439f710aa17f1f9fe5b178be902dd5b66b9f2c05f916d1eea9

Observation 6bb5eb8b-6966-45ff-9ac3-634d3d9238f2 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning In-context Reinforcement Learning with Algorithm Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.194094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.194094Z digest=sha256:24f778e5fab51b42a6394e49c078a48d31f1dc399814e75a4fd8e96cdb462e81

Observation 1e646ab3-d26a-4239-9865-26d51064c042 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.199026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.199026Z digest=sha256:76ee1b3f40fb8a7063e615ac38e7f6f1912035db23c592c7a1f76475c3733f5b

Observation a01a592e-44bc-406f-a3c8-a5ab89e09ed0 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Silkie: Preference Distillation for Large Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.204022Z digest=sha256:a7493d6295c0a1ff078de54bf033e89566e008b26e3aa29acd5c9808dea00198

Observation 9f764ecb-cd88-4e22-9e05-d9beeb962bcb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.208990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.208990Z digest=sha256:212051d2090a2a3327b0fc96af2262e1dda1d8843773c65abed0d5267bc0de5c

Observation 52f9c951-d830-489f-9ce0-fb42c12adb18 · outbound

This paper cites Let's Verify Step by Step.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Let's Verify Step by Step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.214333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.214333Z digest=sha256:f2a7bf4b009efd1185e85ebad61a4e10da2f5536ae8489cec1160d44cb2a6ce0

Observation 048b293e-b7d8-4666-bcab-cd50cb634d64 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.740575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.219132Z digest=sha256:3d7d5054bc235421a7787cc66fb953579d17324e5e2f60b448c251d5abb30b30

Observation 28aa8fb3-447d-4c19-9214-0325a9f0d193 · outbound

This paper cites Improved baselines with visual instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Improved baselines with visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.223939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.223939Z digest=sha256:a85b6c9a04ddf5374eb5808f2468e9c5c08229506a814accd453f4e3eb893db9

Observation 8bf43c82-feae-46e3-949b-2f7e1a3902c6 · outbound

This paper cites Visual instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.228945Z digest=sha256:3141763626fea7821c16c04ff8b281bccd5ba875cc660665c9ed5d8679b81602

Observation 422cb8f1-5374-49ab-8131-14a6d80f8491 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.702876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.233647Z digest=sha256:6de4e04831603528cf9875ead372a58e43c64b90510526b37b786539a0e94a5a

Observation a0b62837-e951-479b-bf63-bf0766baa20d · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.238871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.238871Z digest=sha256:e28da36d8d3b7321c085d58f9f394b376b9dda89bf26c4fb89083f9b836a3278

Observation 06df21f4-afa2-4141-b501-5b7627c19648 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.685242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.244301Z digest=sha256:f7a9110971d65f66f83a9f359b13f9abc1d4317e0ca5ef37b59d111b798f0a1f

Observation 87ede84d-8f10-41bb-a46d-ac44467b0c45 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.248747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.248747Z digest=sha256:cc0271b62b6f06aec033b03cbb4799bedab7196ccff6493ea51008b22dd1d238

Observation 6495c531-cf02-4d26-9b78-8d16ed8d249f · outbound

This paper cites Self-refine: It- erative refinement with self-feedback.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Self-refine: It- erative refinement with self-feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.652549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.255007Z digest=sha256:50a5cb11b6d9374782c968d0c256a2eff035de3e55660f68e0d887828e3b8122

Observation 8d63ecfa-947c-4889-b4fe-80eec2bde86e · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.259690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.259690Z digest=sha256:dfe48bb04a355084270565c8c18dce0ff75355445e403a3d433727ed822e0c4b

Observation c2bc84c7-b20a-4dd5-bcf5-f641e5a9f4e6 · outbound

This paper cites Llama-3.2-11b-vision, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Llama-3.2-11b-vision, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.633589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.264863Z digest=sha256:bc40f63a89e6cbc90582cca31ee1779645ab0e55121de2a18a1c09f7feccddbd

Observation e3d8be13-5f55-4c14-9033-84d6af8a8726 · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.269702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.269702Z digest=sha256:7fef683b5f7d6f34dce2816f0b7127d722a9ac0693d8194d0c0e54a7e22d4c75

Observation d1ecdada-f367-49fc-b211-fecd9b6778f8 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gpt-4v(ision) system card, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.615390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.274830Z digest=sha256:bf7d463f53c690399ae7151556e21a5ab694d3e15a8f54e0d1d7042c6fb7a880

Observation ba6bca92-4df6-40ab-8c03-ec2b0de626d0 · outbound

This paper cites Hello GPT-4o.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Hello GPT-4o

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.597694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.280224Z digest=sha256:7320d0f154558b26dbc793bcf991599e777b097f51b04fdd5054df7b7986e8b6

Observation 4ab1b5ae-9ce7-4de7-bf41-ed108ee4c08f · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence,.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gpt-4o mini: advancing cost-efficient intelligence,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.285300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.285300Z digest=sha256:a471fb758979d41713577519d075453f5bfd5a06aa214e47fe2aee64ba62cee7

Observation 58cc6361-bb9f-4fd6-a1d4-e57357a9388f · outbound

This paper cites REFINER: Reasoning Feedback on Intermediate Representations.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning REFINER: Reasoning Feedback on Intermediate Representations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.295813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.295813Z digest=sha256:751bea38b104498d63b78dd7713421cb081bcdc3942c101f20adbf8bfce55848

Observation d29ed5ec-50d7-43f7-9572-3f24f6753b9d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.552174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.301581Z digest=sha256:b983f38e852a9217d0d1832917e917187d2f10c0dc946bf6dd3dd20c07a2ac6d

Observation 78046100-131f-421f-9de1-8f0bbfb6e49c · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.533861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.306734Z digest=sha256:7b405d539d929f2deefbd06999db665dbd1bd8109fae2975c2e85691f9788bb3

Observation e0cf320e-7c19-4f4f-bc7b-83ce7c498fbf · outbound

This paper cites Learning to summarize with human feed- back.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learning to summarize with human feed- back

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.517214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.311992Z digest=sha256:38d404321569c62e34a5246b9e93e90b2022c778548f2e63e6b04a53b6c6dd76

Observation 3c94c04f-de9e-4a45-b91e-7bfca6409f05 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.316975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.316975Z digest=sha256:a7d648c71f6e8d8c3b9611490f8afec681ab1d53b46dfa994d34e231b9b602bb

Observation 266575a5-ea79-44cc-91ad-8395ae060719 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Policy gradient methods for reinforcement learning with function approximation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.498472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.321810Z digest=sha256:fd341f59509e06c7d452137325375f21e451999ece6c62f9d9c9d2e86101822a

Observation f02a630f-3b4c-44b0-83de-099917c43c33 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.326629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.326629Z digest=sha256:ad091e6862b7422e11c8c5c555fefb2ab8df722b1ffd5e52cfcdf8236ddc0bb7

Observation 2bbdabe7-3460-4516-b7d9-ab3ac4b6feea · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLMs cannot find reasoning errors, but can correct them given the error location

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.331621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.331621Z digest=sha256:a98bfa1f247e6618f35a1079a93fb19985b642396c06de63ed892185c0517c39

Observation db7a87f6-8cf6-4d55-818b-9e8a057a47ef · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.337313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.337313Z digest=sha256:5fe8619ea136cbf094c770e05b66861e5c8672e0974d553a125ee561320ebf7f

Observation 17cfad8e-0140-4bac-9f76-730c358064b0 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.342823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.342823Z digest=sha256:7e70de0241640f179112ac8d4aebdb4d635800529d9b0238b5099054a75fb3e1

Observation 5b9e669a-383e-4016-9fcc-099afb58f0ed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.348467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.348467Z digest=sha256:5fe1372457277cd83ea07af7c67d7f785d0f6db3c85e36f94038194e9887e565

Observation 917ed366-babe-4e96-b341-d5776162cc89 · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.353133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.353133Z digest=sha256:b798fe098ea4ad83f2583796bcf627ec83b7667c46e60725142f6504ce2979cc

Observation bfaeb753-cdcd-4070-aee9-6197022b293e · outbound

This paper cites Grok-1.5 vision preview, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Grok-1.5 vision preview, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.464735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.358203Z digest=sha256:0e7912709d853a14870639f5fe24276211b3517089aba7b9a453ae64fef20633

Observation 710f667f-063a-44ce-8c8e-4465ef4a39c5 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.363118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.363118Z digest=sha256:51fb3a38dbfcc91876edad639d13b88707b4c2f6ff355b37bd7f9d58a1e00e34

Observation 1087fa5c-e965-4048-83aa-9e4506dd8fe0 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.446738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.368550Z digest=sha256:fe5d9207231658d5544fc55a67ab4fd6c75da260459685ac22d8905e30f4b254

Observation 0625d111-4cbe-4aba-8ce7-1f3acee4675c · outbound

This paper cites Learning From Correctness Without Prompting Makes LLM Efficient Reasoner.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learning From Correctness Without Prompting Makes LLM Efficient Reasoner

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.373868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.373868Z digest=sha256:1405412038f407c634ecbf37e601985c2491f49f53472d34a5624a6206f9a824

Observation 10c014a5-3ca3-4a3a-b5d6-c94b48c9adc5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.379680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.379680Z digest=sha256:21ccad3b1a438ec4e3631d0192762b146895fd74b25e9915af45fb80901bc5b3

Observation aa5b0aad-32cf-48c4-b16c-0316c66d5e4d · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.428259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.385690Z digest=sha256:67a6e65cae0440ad1ebe598c866084330e10f15ab769abff643b20a38f885083

Observation b9a1de13-d079-48e7-ad42-73c40f21d99b · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning TextGrad: Automatic "Differentiation" via Text

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.390355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.390355Z digest=sha256:74e6e8717c6c23bf3b486c10d03e6edb658b9b5cff2a9868bfb5e4000cbee1fd

Observation bde391a1-1fb1-4797-91be-ec23c524300c · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.397485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.397485Z digest=sha256:161ab929c238f74f100ef6213188103942e1ca40be7cf7728e210a957a234da6

Observation 4c064ee6-a811-4e7e-87f3-006990c4a0b1 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.403459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.403459Z digest=sha256:e8f61137acb35bc98c1ff005ccd6d2edb810adf61c223f739c532e414440d4aa

Observation a8d72bef-65e9-4fce-9748-e725d2b3ee7a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.408895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.408895Z digest=sha256:000905aad7a9c39a7a0da021bb9cdd7c5399f2c6c9cb6c8394f9ef8afd0c23c2

Observation b0448ac7-50ff-4f4e-92e1-a6e1e4d13b76 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.414817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.414817Z digest=sha256:beb95f2f71a1e54aeb0fa700b15e70b42ccf76cb7d1a2190750bd256fdb73b83

Observation 191d3506-af80-45c3-a24a-ad877de06601 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.419523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.419523Z digest=sha256:aaf0300f4d8531c6f2ce94d69aa3a9b71658f1109157e5c257df081bb18f3ebf

Observation 763db4c7-4943-45db-9dd3-d365522f91a6 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.424887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.424887Z digest=sha256:ec01f66c605aad95a5822b4ae6bf0b150bfb7cbc106f909e662fc147419d3b0e

Observation 2cec8827-945a-465a-b57d-ed01b74ed74b · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.430241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.430241Z digest=sha256:9e3a99e436d473a3bd48b50e3ba4eee4ea8e7cc7ff020d023ff1006b9fa80792

Observation 14067b64-e3d4-4320-8f78-261170d301f2 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Calibrated Self-Rewarding Vision Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.435515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.435515Z digest=sha256:161a5e16fcecdb71b22706ddbd1b4840f8c8e72a7413b784579a36e8f422e654

Observation 221d1c84-03bf-489e-a144-4c5434667161 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.441121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.441121Z digest=sha256:39fa816c0266060a04628f664289ab2e85f78bb47b3a7937753d4c0cf70cd517

Observation 6a56b08c-c0b5-4c9d-9001-1f9b8121916e · outbound

This paper cites VGA: Vision GUI assis- tant - minimizing hallucinations through image-centric fine- tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning VGA: Vision GUI assis- tant - minimizing hallucinations through image-centric fine- tuning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.410929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.446544Z digest=sha256:ed509cc115d2d67686ea4f06fbc0dcbfc2d855affab5e35f234997fca6616c74

Observation e3b08c04-0c7e-447f-baba-6f68ca9d0e44 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.390545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.451714Z digest=sha256:e83dea1b38f059e4eb59037770e2c389701284dec7ee1b932f79e960455a40fb

Observation 99993b3a-981e-4019-bf0f-125bb30f6618 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.374596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.458625Z digest=sha256:c19e17d0778e74428dc614257fed95171bd351cc60c8652b676cc77140079272

Observation f0b1f474-4668-44e5-b06c-db83442fcaa3 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.356708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.465299Z digest=sha256:e9d8275e1071273a301b43ed94cf31f66178dcec79090211fa7ebee86f60610d

Observation 48cc2e8b-a86b-46e1-a840-d6ab3448261e · outbound

This paper cites For preference-aligned fine-tuning, we utilize Direct Preference Optimization (DPO) on 29,012 samples from the critique- VQA dataset.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning For preference-aligned fine-tuning, we utilize Direct Preference Optimization (DPO) on 29,012 samples from the critique- VQA dataset

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:27:34.332142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.471122Z digest=sha256:2fa6c80f41fc0cf29265004326fa3928acf52ef8dcf97757839693e0a61791f7

Observation 19a48ded-338b-493e-a1e1-4d160a218afe · outbound

This paper cites In this section, we will list out the hyperparameters we choose for evaluation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning In this section, we will list out the hyperparameters we choose for evaluation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.311708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.475936Z digest=sha256:15f913e332d4de387dcbbffda17e1686baa1e155be6d2b06621a35ef857e1e99

Observation 470ff0d4-673e-42dc-94da-807a7963785c · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.295105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.480557Z digest=sha256:6189ffba19994dd5ba1ecfb54ee83a299fb44123fdae225c8d1933e600c84d9a

Observation 1e61a02c-5122-4f37-a31c-9e03ba72aab7 · outbound

This paper cites You can find them in Figure 6, Figure 7 and Figure 8.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning You can find them in Figure 6, Figure 7 and Figure 8

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.274023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.484764Z digest=sha256:d5765df940125ea83377f1ae1334d955ffb50b2d06ddf2a7c6f07556779e3ce3

Observation 1a395061-503d-4360-aaf8-4d1ae46dc235 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.253077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.489216Z digest=sha256:0c80f8414dac928c961057355d0fe9d1ac910f5e63aac2501fb651f1e81e3f1a

Observation e39a1e81-85d4-441c-b558-f7f657f6c09f · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.236407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.494574Z digest=sha256:c8ca1d4dca5e1e12ce501a277191ed80eb0460a04a9f07ce488799ce27af8e0b

Observation 84196f46-7d29-4e07-98c6-98ff17b6ef59 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T11:27:34.569322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:27:33.290232Z digest=sha256:7930e9a62be145abeb53215c1d375aa6a27434ee229ed93383ea17c87a98c7fc

Pith citing papers

Observation f7ccdbab-a1c1-489f-8523-1d22700a6232 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.308563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.308563Z digest=sha256:dfbda918b0df73a441240669fe09d756753da2b286e28a5a3d0c26f95cdfe5b2

Observation 3dd5afbd-cc1f-4a06-883e-04f55644e95a · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.875185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ef33d3511716049e7164bd4d2796b703f42326e210d0aaec08d521230532b482

Observation f0639fcb-0b98-48d9-b284-46d6eb8d4c83 · inbound

Test-Time Hinting for Black-Box Vision-Language Models cites this paper.

Test-Time Hinting for Black-Box Vision-Language Models Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.061496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T21:16:45.748116Z digest=sha256:5190b0792261c29c0a257bcbbeda893f3605ae4a8b05c3acf93af8ac7bed4bae

Observation c33ebc55-f171-4fb7-a660-b1458662dde5 · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.803674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:da06ee00338638a54caad4abe0d9435adf16cc968d881eb21a937727a4acffdd

Observation 534359d2-0f71-4412-8ef1-5ec5e2c6d94e · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:16.664589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:16.664589Z digest=sha256:5e5d8d519edbf49de66baeb18fc0d7d9557a0b5e0b07b06dfdbe8de71c2b3f82

Observation afc8d2b1-b606-46cd-840f-99b58920453f · inbound

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus cites this paper.

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:56:01.975427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:56:01.975427Z digest=sha256:753d98b568e7ddc7e2d967d53f05c5f083366a9cca4ab1b0a40f728dc6745bf8