Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 13 inbound Pith citation observations for arXiv:2509.00676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00676 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:39.981501Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:05:25.733973Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.567326Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2fc229f3-c54a-4901-ad3e-1b51b820e4a1 · outbound

This paper cites Qwen2.5-VL Technical Report.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.590118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.590118Z digest=sha256:86f3b1a6cd593c7be769dbaca4f4685e51c9db84f1b7551a57ef5ecd1848bc8d

Observation b5727700-f78a-4c26-99df-12418ed65885 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.596363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.596363Z digest=sha256:6e1999408dac235aaf3a27b9d10a05a82f481fc2eb29787b39608a66cc70606c

Observation 5250fb7f-1788-47ef-9e57-2043d760ea48 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.601972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.601972Z digest=sha256:0102fd132dd720dadd76936d9b6a9099b89d95fa45fb3a14e89c83a038586433

Observation b0b34e68-9e88-4e54-865c-cba0aaa15173 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.607386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.607386Z digest=sha256:9c6ef3127df580d9ebf85ac8fad914bf278919a12ecbc83f737c9c80b7024d57

Observation 6cb47259-c3aa-4b8a-b303-e66ff9a16ecf · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.613659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.613659Z digest=sha256:2192228f656588254904d940be02cfdf2280a4f5388949d5a35ddeeadc1093cb

Observation e84394d7-224c-45a2-b0ab-6454bf1f3611 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.620300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.620300Z digest=sha256:b3c61b7ab7862dc19145d773dc088e41cd6278cbc6c30d8dfbb9460435ada3fe

Observation 7fffd557-bb35-4f9b-bed8-0e1784643ed6 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.626362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.626362Z digest=sha256:cb1ee92d5bf8037ab261b0f71538caa5d36e836fd54e0acdb9b876a672e3f675

Observation b3588781-8ee6-4566-963c-9e74225587a6 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.631783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.631783Z digest=sha256:269ca2c686561e48cc7002d2a706d657f35c460e9a7a267bbdaa2b4f6bab1c9d

Observation 71821fab-acc8-4ada-a812-97ad3ddfe4df · outbound

This paper cites Scaling laws for reward model overoptimization.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling laws for reward model overoptimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.637885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.637885Z digest=sha256:91c1d4c07376d80d553d831977eceb735977bc4db22749395da37bffe51ed48d

Observation 7f011b7d-e72a-4d58-9801-99b13bf437fa · outbound

This paper cites Interpretable Contrastive Monte Carlo Tree Search Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Interpretable Contrastive Monte Carlo Tree Search Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.642992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.642992Z digest=sha256:99e4678d0c563467aaadcff5c2358baf5cf696d38e9a261d8b7bae9bfc4b4159

Observation d3971631-1abe-431e-9f16-cf3e1610741b · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.648950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.648950Z digest=sha256:c7344668251cd71c21a6554353882e6a1808f034f6e037baad264ea39611f88d

Observation aa570e30-2544-4fe1-8f67-24262a221426 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.653975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.653975Z digest=sha256:715ed983026c49736dc3461fffd684f47ec1952cefa8ea1c783b9c712646f1ea

Observation bac3de43-4733-4cfc-8917-fae8087cf7ea · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.659106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.659106Z digest=sha256:4e522d473dd65effcc023d85ab719f628763d4abf3512d7dc7875b1277fa01ec

Observation b349b044-87c8-4a74-8c06-fc34d9966ed2 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.670700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.670700Z digest=sha256:d84da4205ec0998ed0fe645356f57417cc6b0f51aaeb800a032e8f1840272310

Observation a538029e-3803-4a07-aa4f-de78abda9025 · outbound

This paper cites OpenAI o1 System Card.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.675483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.675483Z digest=sha256:10a430a0645bf17ee83fc09e3c9df5f0dd5cc21ed6bb3a17180b0541ab2e5526

Observation b9ced648-191f-4b70-8392-a741a2529c46 · outbound

This paper cites A diagram is worth a dozen images, 2016.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model A diagram is worth a dozen images, 2016

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.680307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.680307Z digest=sha256:661e922ac4b66df10b6968b033fe843bf76d1ded1d07f3bf09569b65f3018a17

Observation ef2a548f-1c98-4a1a-830e-a71a6d157ba2 · outbound

This paper cites Process reward models that think.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Process reward models that think

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.685138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.685138Z digest=sha256:3b757d5025354cdf4bf53dd1a2ceb9820e8a9be6be1c854ecf08c37ab1c1b388

Observation 178119a2-420c-43e2-891c-510f2a3b9fbc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.690648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.690648Z digest=sha256:736ecedf4ae9d23b3fbab0b04540dc2eda32bdb10a2b89b9c9ed08ebc8b3fb30

Observation de18d49c-57ed-402f-a32f-3de92e04252b · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.696801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.696801Z digest=sha256:445cae5031e60fe71ea487fbbc736f4710f86748701a74a3e9c71ffbe3ddbc9c

Observation 4c40fdb1-95b6-49e4-821c-730d39f4ef74 · outbound

This paper cites Let's Verify Step by Step.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.701994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.701994Z digest=sha256:4b0a6e67b8a29afe26b38971f7609dc85e968e9536ae64f3be6d977632eaa667

Observation 57fd7618-4392-4cec-957b-96f152aab42e · outbound

This paper cites Let's verify step by step.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Let's verify step by step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.706968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.706968Z digest=sha256:e8823d6c2aca3ab227362153ab308f2c309da0b41c24bdc934296bfa20750fff

Observation 3305b9b3-48f4-453e-aaee-f65503ffafa2 · outbound

This paper cites Visual instruction tuning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Visual instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.711719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.711719Z digest=sha256:605e86bf4b2fd62496cf42f2a60c532dc90ae5ba535ea48ba5942089d2f0be29

Observation 3e770e54-9690-40ea-86c1-9bb263ec143a · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.716806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.716806Z digest=sha256:455ca9d740c07f1e344ebd5b91fe0b60da71ad0e752dc2d0752c7c78ba5d2393

Observation 07490fcf-1e7d-491e-92c7-5752b4acaefe · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.322013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.722147Z digest=sha256:e7a686f4f6352381bbd65827d84e6070888516175a10d1fac4ae7e9658d8cb65

Observation 31857890-368e-45ca-9821-07fb765e8b8a · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.304994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.727020Z digest=sha256:d474c0b70ab212e97402e43521c8d5b555509fd1c6d23dc820c965bb909b5962

Observation 723e0e19-0b3c-4c7d-80a3-904aee3178e1 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.732122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.732122Z digest=sha256:124e5e5b6ef4c1c1cd2b30eadb372aab3da07e4a96bc85026b172793c58b9567

Observation 74715f75-17e2-4199-8624-cfb7477650ad · outbound

This paper cites Generative Reward Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Generative Reward Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.737223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.737223Z digest=sha256:12325d9aa301b3c76ec8a3f25aae5f969f48a1c524ad43424ef1e6117ab684b3

Observation 5a250e68-03df-442c-92d8-f8ca24badbbf · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.742920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.742920Z digest=sha256:5fab8fdc7d317c351b7506304e162243c56612ef9a34bf250bb1e356f5210b41

Observation c55f1b8a-1eaa-499e-9d5f-ff3c7c2de962 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.747881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.747881Z digest=sha256:9935cec4ba433f4740c4d421a3560cc2ec49f19cd109657b05ab9e889fb71722

Observation 34765e4b-2911-4aa3-8025-95466ac147c8 · outbound

This paper cites Gpt-4v(ision) system card.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Gpt-4v(ision) system card

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.277440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.753145Z digest=sha256:20321439c84278abcba4e70e5ea7ab55574d739423656bb5891a8d4a2832c698

Observation 51c071c5-8514-4663-b398-99f275f20ad1 · outbound

This paper cites Learning to reason with llms, 2024.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Learning to reason with llms, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.260915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.758088Z digest=sha256:1710d44285bcfd8b404d50ee3b7195c2b079e738d3ddf4c9d4f616b10ff58e03

Observation 3fcc1a58-bc25-43b9-bb6a-09da4fb64832 · outbound

This paper cites Training language models to follow instructions with human feedback.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.763610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.763610Z digest=sha256:6a413829281661ca5e809781bf9b68f5f9187b0f985c482267aaa9a7ac92c3ec

Observation 946504b3-bcbd-45dc-8439-217a274ad464 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.768630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.768630Z digest=sha256:be2c36f06e9ff5287fb47d1a92ba96a572b011097ad66e7bf2027885ac2739d1

Observation c9be0652-2f35-4f3a-9ca3-c9eccf814e1a · outbound

This paper cites VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.774604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.774604Z digest=sha256:230b13df5ebd21908acfbf44e3700d13cd055d9534cdc7f603d63975bda78e8f

Observation 7dcf6623-c1af-4cbf-9a42-11066e888fad · outbound

This paper cites Vision language models are blind.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Vision language models are blind

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.232261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.779624Z digest=sha256:766889b0f6b45e82a18e99a09cbbaca8258795188f4ab95d966df46a51e51ad9

Observation 9bc0b383-70dc-41d8-b6f7-78728df2e511 · outbound

This paper cites ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.784349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.784349Z digest=sha256:b51edb7276ed2f59b92354b0c1daacb378a0cd5fb6bfd26f1150cb30883c395e

Observation 2b0c2e3f-79b5-4094-bab6-81b11772e5cc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.789396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.789396Z digest=sha256:adabfdaafef854d8d903ddc19a0c3b5554aca3efa365311780d11aa5f01b105b

Observation 92d408ac-a580-4b8e-afc5-0296b78e67ed · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.794093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.794093Z digest=sha256:e02cf46e9f49cc8c2c5f6506ba4cd9c8259a4b0274208b8f2056ad84df89dbab

Observation 43047e17-8353-43d1-9bfd-bdee1207effc · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.799101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.799101Z digest=sha256:043732662e207b79da8de3beead213c4d0772112c730eb49e174477d9f817529

Observation 1283b163-a25e-4989-99da-ea011bde4706 · outbound

This paper cites MiMo-VL Technical Report.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MiMo-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.805838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.805838Z digest=sha256:9c50bbe5ea75d0f8ed8fce7d5a21cc2541abbf204aab9b0f3edf58931f59e825

Observation 879f3a98-fe4c-4a9a-8316-1e03eaad7e5a · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.812161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.812161Z digest=sha256:44a51b98f973099889a2f56ea4e96112cfadf8a3c6701aacc3663ec2861f70d0

Observation fa75da25-bfe4-4e13-822e-b46c35599431 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.817974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.817974Z digest=sha256:d5ec72834e6e33b3b5907f97de60c3ed944e3a365e0ea25e38d14c46741a2f9b

Observation 6164018c-232a-4b4c-9561-770c7be1db9c · outbound

This paper cites Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.823412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.823412Z digest=sha256:a080a0aba6c81e1677cfe88f4ed302b4fcb053a68b22b7caa336127fa656f9f9

Observation 0805543d-08da-4fe1-8fcf-2341a6090190 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.829566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.829566Z digest=sha256:da6cd820f81410b5b00cc44851bef0d167c5308b408b06d35ff3eee2d0a93802

Observation 7370178e-d63b-4e1e-a69d-d9e8cceec03c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Measuring multimodal mathematical reasoning with math-vision dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.215347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.835197Z digest=sha256:42751546aa77d75b079e0f4eb38b3031de6bbbb8b010782c75713367673305ed

Observation 5e30e8a2-f9a7-480f-ac96-142f40427ac7 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.845687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.845687Z digest=sha256:bfc2f6c01e98fb3b206e44ade3028044d8f4eac36a6695f17af597fce8ae2ca5

Observation 05e5b6d7-ed0c-4f79-b7bb-65ec5234e508 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.852228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.852228Z digest=sha256:d85cce2008b19aa23235f2909dbce3e18615d7c8c958be0edebf5b65976cef90

Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:24:40.393917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.858355Z digest=sha256:5fd1f732db9680ec48d20d3d7c1b167a1e7bd03f5b8625621d297ce08d494523

Observation 9baea096-df66-4209-84a4-cc84c740d227 · outbound

This paper cites ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.864757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.864757Z digest=sha256:a6e80d1e3d14a4f9c908c8cea8584a8097d955ea4540b48f6545f78c7902e822

Observation 47e2937b-eda5-4503-9d59-4ba53a5bd6b7 · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.870088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.870088Z digest=sha256:47685c492386837d7a24f106a03715c3b13b57c6a21441687c45690e0acc5306

Observation 3674e452-b019-4857-80a8-1c47202b94e9 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Chain-of-thought prompting elicits reasoning in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.875987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.875987Z digest=sha256:4955f2a57f21ec4f0072131d5780b65515a1bf2d51bb95552cc2a70327b51464

Observation 5ad221a9-477a-4640-93cf-f9ec647b244d · outbound

This paper cites Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.880646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.880646Z digest=sha256:b7e0014e8397c6e792566f217c7126f0ee62ab6e197020ddc57297d83acdc8e3

Observation e3c446d9-6703-4203-825c-4baea7897aba · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.885634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.885634Z digest=sha256:dc88b7da7bd49084fff667f24765dc2c53dc6f2b74f9c587f2b2f8ad63c72d1d

Observation 768a6749-4d68-45da-9c1c-064e1b3fc2f9 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.892735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.892735Z digest=sha256:96bcad495488da111e5e88497fa6961e93b5e357d389bb3b493b83ef72614578

Observation 3ed8bf8c-fce5-4454-a622-c5bef9720374 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.898650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.898650Z digest=sha256:4cb055f846e10957e8a96cc8a1dd6fb3ffbb8a38bbcaef08f5883caf0bc26bb9

Observation d55eb9da-3eb0-4479-8494-2aca89b2e0a5 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.903930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.903930Z digest=sha256:580149b21b84b6ccf9800932fb14242a0f170aca0bce9fd3b469058c870c2df1

Observation 7f57d251-a640-4578-b521-e592117281b3 · outbound

This paper cites Llava-critic: Learning to evaluate multimodal models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Llava-critic: Learning to evaluate multimodal models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.183495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.910004Z digest=sha256:380c947d6fb06b6d12328141b2e7057d2517ca301fdf76c9e34a4fba3711f07c

Observation b840d0aa-88db-4ee4-8d7c-aea8368770aa · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.914942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.914942Z digest=sha256:674bf3b4d7fab17fd61bd221d1842b3d13f734c5f5154649fa79c0abc18c6b30

Observation eb46d933-46d4-4af1-a71d-31fb270055d6 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.920134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.920134Z digest=sha256:43651ac38e6d1edd649801d1fb19d763b1b7e99423137eb1232a83a812bc5b4f

Observation 9c80213e-e434-4764-b9a1-dd9b0fe19e70 · outbound

This paper cites VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.926180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.926180Z digest=sha256:2d2635a3e4a8c09571cebd05f95c981f2f9b43bb2d2da7f904d3284c22e14242

Observation aea25225-aca4-4e11-a217-6a2b89ddfc92 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.931710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.931710Z digest=sha256:654293bbd5bd27773c88b9d4cdfb168b1e3644427fd92cbe715bbf22aa363c02

Observation 2d5e40fb-ea29-4f77-b6e6-136cf38cf54c · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.936440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.936440Z digest=sha256:da8e3b1d0ef1d1f16bd3cfd2110e256ea8b8cf40aeb7d21d3bc1d82349dcbea4

Observation 3b2b7650-77ba-4a4c-8f8a-f9cdadc85895 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.942045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.942045Z digest=sha256:a20efa41562c79607c0a11718bbe8af8d2693d7360b4dcb11c57bc51563b507e

Observation a4ba5982-1ad0-4cc9-a6b8-e41474e70ca2 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.947616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.947616Z digest=sha256:99ef45660101530bd41c903da7ef83c9d308fed74f09c901f50147a85a0d3e01

Observation 8d24ae44-7827-42b1-bf90-4ea7efc56487 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.953253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.953253Z digest=sha256:17c04822a16030ff389c0b7ba9ee78d7ae20e1976d1f0e31cce7c600a4747eea

Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.958836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.958836Z digest=sha256:6e9367c816c501660e0ba53bc5043f94fa246467a4d08a12e1fe32160b5427b2

Observation 9d7afc8e-51e1-4898-9593-28a73edb2c14 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.964837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.964837Z digest=sha256:1ad10203c6afb60ef60d3d51a36b68a558cc49576e189bf6564be4f3248fd357

Observation b17e8c60-dd13-46b4-9bf3-6a46a0cd282b · outbound

This paper cites GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.969937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.969937Z digest=sha256:6b313035ec87febebf1df929d96ecbc6f793fbbb67fc6522390ea84d32974241

Observation 30cc9626-792b-4408-b9c4-d300a389e698 · outbound

This paper cites Mmvu: Measuring expert-level multi-discipline video understanding.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmvu: Measuring expert-level multi-discipline video understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.155334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.975792Z digest=sha256:06ad56f2ca5dd49fff9608154b3cc8c7b2cb125a8e9642ca4527886181a738e3

Observation 93c03e7c-9b72-4505-99e9-f8fe3882c242 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.981501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.981501Z digest=sha256:b643cc3a052c0a6ac9f7d669947993e2c7e4df368b5a6f9611a34ec577c0cb3c

Pith citing papers

Observation b27b417a-d61c-409b-9100-0b2d4f172ddb · inbound

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models cites this paper.

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:31:13.050244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T19:29:43.382392Z digest=sha256:7ac8ed32c5e3158ed3514e858bde99d88d4b776e20e011f7629f30026af63d34

Observation c68e6a67-1d50-4550-87ab-dd7ed69a86fe · inbound

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models cites this paper.

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T14:05:25.733973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:05:25.733973Z digest=sha256:e14308145b8f9046197f037ce7bf5a103699afdd3d46cb60d184bc0b21ceb445

Observation 9bdd290c-5381-40d6-bcbd-09890ecf0064 · inbound

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate cites this paper.

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:30:13.790750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:29:15.751497Z digest=sha256:5cb4e261bc54be0761516745491d6955cfe7980c5615d5f9c5c556a580b56cf2

Observation b1a125c1-20ee-4954-a735-d66748f1417d · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:51.701504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:9e8fefcb9e6d841840bbc39978a30f86eae4e945c35e693c2583fb7eed705c2f

Observation d21428b6-34ea-44c6-b988-a1c1f5f3883e · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.934946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:3bd01b574bc23e9c82abe81328bdfd0755d6f14e62fb6286a2e899f66ff41ec9

Observation 9a247ef6-a82d-46aa-b669-1071141cac14 · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.805123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:59d72210974fa2c5ec85ed3f905b93fb1fe91bf6ef02863f8bd71a9d2f30f20e

Observation f6e06dc6-c571-49a2-9719-de29a1412d78 · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.768132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:faaa2cd0e1ad35ed7ac6b3619757bab87f3a8ca61a4d12506536ced9f3c743c7

Observation eb90c0cd-3b97-4812-80b6-710292ac054a · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.496389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:0b2a974ecd8118e8183ca9aa48a56dd801d5b58edfd4b54527bdc92aee3b8798

Observation bc854a1f-8dce-4537-9616-c1aebfd183d1 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.651999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:1e68888ca3624b5155f538bbf4f732f929e3072b81e2ce1b2627339913281d1b

Observation b631b8b4-6629-4e9c-b17c-c41a0306dd70 · inbound

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety cites this paper.

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.569526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:03:27.948225Z digest=sha256:67378984fef11da14b1eb188b3195dc0ca4c37ae8ecd74e646514c857f126775

Observation 15ec6608-c26e-403a-8397-d6582a6b4fc3 · inbound

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety cites this paper.

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:48.808047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T05:16:19.502837Z digest=sha256:302ef9ce3835288422f9d8bea743add804509e55765a00b4864a1918ce8a12a3

Observation 98c9f657-6b68-4eb6-8288-294bb38a9e9a · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.208153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.208153Z digest=sha256:9c123738e92cc3a31cd578d1e340988eefc1c9161586ba15545ea92b18cbedff

Observation c5a3c139-f196-4cae-aa08-cd8e0f2cc00a · inbound

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift cites this paper.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.427244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.427244Z digest=sha256:2d165bc305d3fd95a620727caec99943bafbd773a1c2063164278ed61d744676