Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

As of 23 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 14 inbound Pith citation observations for arXiv:2509.00676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00676 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:39.981501Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:56:01.952589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.567326Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2fc229f3-c54a-4901-ad3e-1b51b820e4a1 · outbound

This paper cites Qwen2.5-VL Technical Report.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.590118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.590118Z digest=sha256:03b0464040786b2cd1787018ab4c81d4c259912bf92aef5a739f120c7c270cbe

Observation b5727700-f78a-4c26-99df-12418ed65885 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.596363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.596363Z digest=sha256:4ccc6fc06f724eb536de47892c534b93624031c99d815e40a1326a7cf624f0e1

Observation 5250fb7f-1788-47ef-9e57-2043d760ea48 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.601972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.601972Z digest=sha256:3565d376a16b2fb1ee5261fd1ed8c3d588b8db64cdddc8168c71fbb947738bfb

Observation b0b34e68-9e88-4e54-865c-cba0aaa15173 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.607386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.607386Z digest=sha256:0fcefacdb4ccffea2ff9ea8468045d7ba626247a0bf9ca2faec32fb55cc29f9f

Observation 6cb47259-c3aa-4b8a-b303-e66ff9a16ecf · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.613659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.613659Z digest=sha256:dac686c5ae4237f01265444582439cefdb5ee9efbab572d2467cdc557f10574c

Observation e84394d7-224c-45a2-b0ab-6454bf1f3611 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.620300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.620300Z digest=sha256:65cd5d8d43acfed06c9745eb4ecf4c4cde279bc800939ae8095cc34217b67a21

Observation 7fffd557-bb35-4f9b-bed8-0e1784643ed6 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.626362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.626362Z digest=sha256:41c3a70ae82e15ef9288c4e0472d9ffe915b7c597a292d5c6513412dcfe2d7d7

Observation b3588781-8ee6-4566-963c-9e74225587a6 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.631783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.631783Z digest=sha256:ac01ebf5d814f8dc04e54d8f93970f98026dea5b0e3f6e118f85101d6ebfb254

Observation 71821fab-acc8-4ada-a812-97ad3ddfe4df · outbound

This paper cites Scaling laws for reward model overoptimization.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling laws for reward model overoptimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.637885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.637885Z digest=sha256:a788a26e85332d3db439aedf097a44ddb20626e97283e17aba234de3ad2ee57c

Observation 7f011b7d-e72a-4d58-9801-99b13bf437fa · outbound

This paper cites Interpretable Contrastive Monte Carlo Tree Search Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Interpretable Contrastive Monte Carlo Tree Search Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.642992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.642992Z digest=sha256:1bba2c5f6e6d2f4c08464b002f208d8484faaf80f6245323dd04970d9600271c

Observation d3971631-1abe-431e-9f16-cf3e1610741b · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.648950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.648950Z digest=sha256:8637dcbf06e866b226b221dcdecadb4e5a7292d4f4fc9b5cc77afda012c48868

Observation aa570e30-2544-4fe1-8f67-24262a221426 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.653975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.653975Z digest=sha256:e7484e6dffe525a6c89a0fa318a389d17877f4e5cba0c02d3b185e692d73aa2c

Observation bac3de43-4733-4cfc-8917-fae8087cf7ea · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.659106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.659106Z digest=sha256:ae07c2e20979135d581e7c66646a5871449067ed8770d834df055cb414c8a2a2

Observation b349b044-87c8-4a74-8c06-fc34d9966ed2 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.670700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.670700Z digest=sha256:b6a77e98faa7a240f5cb8fe42f0de3ff34c8ace14be971f88d6cde640e4e6410

Observation a538029e-3803-4a07-aa4f-de78abda9025 · outbound

This paper cites OpenAI o1 System Card.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.675483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.675483Z digest=sha256:1ce17e2780c4e574070243910cab74015f9420a010430f378cdf612f3d82d92f

Observation b9ced648-191f-4b70-8392-a741a2529c46 · outbound

This paper cites A diagram is worth a dozen images, 2016.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model A diagram is worth a dozen images, 2016

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.680307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.680307Z digest=sha256:3e6c1031c64a987a0fea7ac9fcfeda5c33a3004ff031bbec1241bd3271bb5a45

Observation ef2a548f-1c98-4a1a-830e-a71a6d157ba2 · outbound

This paper cites Process reward models that think.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Process reward models that think

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.685138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.685138Z digest=sha256:bd721c8f354f56605a4deb22ee3539b8f1fcf8b188cc6ec89c975255b45ce8f6

Observation 178119a2-420c-43e2-891c-510f2a3b9fbc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.690648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.690648Z digest=sha256:429f0aebc160d8b773df5df348cc9bbc440b994791a3dc827298535fe2001790

Observation de18d49c-57ed-402f-a32f-3de92e04252b · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.696801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.696801Z digest=sha256:6be561e060090e2ac28698713a68ebc276476a6f1af8b40d0af734e6a481b039

Observation 4c40fdb1-95b6-49e4-821c-730d39f4ef74 · outbound

This paper cites Let's Verify Step by Step.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.701994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.701994Z digest=sha256:01e15951c8e73b78c81608eafec858e1ff4519f719333dd0d53c3201c2df1c13

Observation 57fd7618-4392-4cec-957b-96f152aab42e · outbound

This paper cites Let's verify step by step.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Let's verify step by step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.706968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.706968Z digest=sha256:06ed7126139d696ec1fbacb4b2d71a36515ac827df7c404f28c6cae8f83e71a3

Observation 3305b9b3-48f4-453e-aaee-f65503ffafa2 · outbound

This paper cites Visual instruction tuning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Visual instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.711719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.711719Z digest=sha256:2663a5730ff86008ef127e8574dd8de505a57ec07c12822a739e655e08fbc40c

Observation 3e770e54-9690-40ea-86c1-9bb263ec143a · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.716806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.716806Z digest=sha256:bc985948e8e080d6ec734da8b8f391bd03750b6ccbe332fb9b0fc79ea173f1be

Observation 07490fcf-1e7d-491e-92c7-5752b4acaefe · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216--233

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.322013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.722147Z digest=sha256:09dfea6e424e0644b83abf78316d6f79c575c9bac7bdd189019db47635e2a383

Observation 31857890-368e-45ca-9821-07fb765e8b8a · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.304994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.727020Z digest=sha256:5afcb9bba2ddfc5933f79f3fcbd50fe9ba3783124bb32898d72bbe9a989c88af

Observation 723e0e19-0b3c-4c7d-80a3-904aee3178e1 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.732122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.732122Z digest=sha256:ddfe545bc5b23c509884fd9fc9b7b0a852bbfffb0d9f06f7c036bff137da7309

Observation 74715f75-17e2-4199-8624-cfb7477650ad · outbound

This paper cites Generative Reward Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Generative Reward Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.737223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.737223Z digest=sha256:9a93e99789e6c10648b2077cbed12784127203f400aa05d392c3132834dab8f5

Observation 5a250e68-03df-442c-92d8-f8ca24badbbf · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.742920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.742920Z digest=sha256:a6ff0194627a96adcd83e0976486ada344bef94d6d6b40149b1659f14aa41b12

Observation c55f1b8a-1eaa-499e-9d5f-ff3c7c2de962 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.747881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.747881Z digest=sha256:e61875d347c6cdf018fe01013557c0c193bf395d69a41075c9bd3723137cb50f

Observation 34765e4b-2911-4aa3-8025-95466ac147c8 · outbound

This paper cites Gpt-4v(ision) system card.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Gpt-4v(ision) system card

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.277440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.753145Z digest=sha256:96bf5df11a1d87006e17a5e0db6234bb7ef25cd6d9b3e77c9b48b3638488e48f

Observation 51c071c5-8514-4663-b398-99f275f20ad1 · outbound

This paper cites Learning to reason with llms, 2024.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Learning to reason with llms, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.260915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.758088Z digest=sha256:a19af24be7a7547b3ce6bc0d2cbc4451790774592e56add4bb4d95b6776e5949

Observation 3fcc1a58-bc25-43b9-bb6a-09da4fb64832 · outbound

This paper cites Training language models to follow instructions with human feedback.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.763610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.763610Z digest=sha256:b4ad79e85acb238d6ddf4094495d95098c8ba2c18a2134a73962b2b7d814b405

Observation 946504b3-bcbd-45dc-8439-217a274ad464 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.768630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.768630Z digest=sha256:e581dc3d92e1bf54a04a5f26b1afedb9a485e16bfbd31d1dd5dc82d61cde6491

Observation c9be0652-2f35-4f3a-9ca3-c9eccf814e1a · outbound

This paper cites VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.774604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.774604Z digest=sha256:22e4d099a453cc152913c29d9934308e7ac7e52f08005aa246f0c25c313f3b15

Observation 7dcf6623-c1af-4cbf-9a42-11066e888fad · outbound

This paper cites Vision language models are blind.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Vision language models are blind

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.232261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.779624Z digest=sha256:4821bea700ac60d9c91509ea1f6eafaa5fb18b5ec15f4e7920e83103a4c3f9e1

Observation 9bc0b383-70dc-41d8-b6f7-78728df2e511 · outbound

This paper cites ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.784349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.784349Z digest=sha256:48ad3c4668ae2ab10b46de9fab7832d3983581e293743e0a45013b0d3cfe7d97

Observation 2b0c2e3f-79b5-4094-bab6-81b11772e5cc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.789396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.789396Z digest=sha256:e14b0f3b22c87eb9e75f0f1f593b0ad50b6e419014310eda35373b6301fb4b84

Observation 92d408ac-a580-4b8e-afc5-0296b78e67ed · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.794093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.794093Z digest=sha256:ba20a4d94ac445d23857da0bf3c087b8a6dc93d53b789fe64a99a007eafc21ed

Observation 43047e17-8353-43d1-9bfd-bdee1207effc · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.799101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.799101Z digest=sha256:04d9c732253a0a9861f5d803521f306225aa1e52d9909ccd12b5c800b734507e

Observation 1283b163-a25e-4989-99da-ea011bde4706 · outbound

This paper cites MiMo-VL Technical Report.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MiMo-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.805838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.805838Z digest=sha256:74448c11278735ae431528d70666e395a787a99b4a7fc581bd142ba3c4cfddde

Observation 879f3a98-fe4c-4a9a-8316-1e03eaad7e5a · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.812161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.812161Z digest=sha256:0fb131b74b1e81265f4989114426cf3ad8d119c13d4f1fa57a4c294550fed989

Observation fa75da25-bfe4-4e13-822e-b46c35599431 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.817974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.817974Z digest=sha256:c379c05dcea59a570d6ac6fb4f61607c61e930ed730c9def64284e0daa448db3

Observation 6164018c-232a-4b4c-9561-770c7be1db9c · outbound

This paper cites Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.823412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.823412Z digest=sha256:e6cab832c617fbe5814251918ddbbaefb3e06fe3e03142f1918ab13ea449c82f

Observation 0805543d-08da-4fe1-8fcf-2341a6090190 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.829566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.829566Z digest=sha256:e3e2bcc6d1606a0ca174fc4d673a8605da94f6ac77f7d92861405a26aaf50f5c

Observation 7370178e-d63b-4e1e-a69d-d9e8cceec03c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Measuring multimodal mathematical reasoning with math-vision dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.215347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.835197Z digest=sha256:f4065425ed5e8d9769973c37d284bd7ed48f28d32ee6083cb07a6da2619ac66f

Observation 5e30e8a2-f9a7-480f-ac96-142f40427ac7 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.845687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.845687Z digest=sha256:5741902338489bbd0f901eacec154895f8aecdf3a83f8a1309f2896c38439da4

Observation 05e5b6d7-ed0c-4f79-b7bb-65ec5234e508 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.852228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.852228Z digest=sha256:02c3245901ee652a17cf036792d29416190a4044de706c904b11da059ae1a205

Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:24:40.393917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.858355Z digest=sha256:b250caf4b9b9cfdd2c5fdb21651c619227764580735753b82f2daaabe8ea1eab

Observation 9baea096-df66-4209-84a4-cc84c740d227 · outbound

This paper cites ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.864757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.864757Z digest=sha256:408967954c79f9824e132670257f68700a7c477358418a294be4925f7047c8ce

Observation 47e2937b-eda5-4503-9d59-4ba53a5bd6b7 · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.870088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.870088Z digest=sha256:be02428586c59da85de9696a44f9e14272734724a2431a2bf6629a147d850f24

Observation 3674e452-b019-4857-80a8-1c47202b94e9 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Chain-of-thought prompting elicits reasoning in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.875987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.875987Z digest=sha256:45b327347fd7b9121a0ae0ca0021a4dbdc016973452081cbb186d9519c715f1f

Observation 5ad221a9-477a-4640-93cf-f9ec647b244d · outbound

This paper cites Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.880646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.880646Z digest=sha256:310edbfc65784b56cb430777961bbc8ef0d185735fd67b2ab615ecd15d966ae3

Observation e3c446d9-6703-4203-825c-4baea7897aba · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.885634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.885634Z digest=sha256:6470772a9ea78406f0db6ff38bc9ffc0cb1f85c6ca7aaf1377a66f1a2e3a2f6e

Observation 768a6749-4d68-45da-9c1c-064e1b3fc2f9 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.892735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.892735Z digest=sha256:40f1108d42ae6ac428b236cad9b862c35d096f0183a0fbf7bcb0e361fd59198c

Observation 3ed8bf8c-fce5-4454-a622-c5bef9720374 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.898650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.898650Z digest=sha256:5a08dccece6c1daa136278c6754315335d7035bc9b0a8a94614d9339742e9f14

Observation d55eb9da-3eb0-4479-8494-2aca89b2e0a5 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.903930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.903930Z digest=sha256:ef8795d3fb2a47d30cd5bbff520c183d18ba552a9fe15d3bf910235a8273538d

Observation 7f57d251-a640-4578-b521-e592117281b3 · outbound

This paper cites Llava-critic: Learning to evaluate multimodal models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Llava-critic: Learning to evaluate multimodal models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.183495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.910004Z digest=sha256:0c84ba3b29a8f0237694b935801d20dd293932e0ac1b0c57bf907ada781bcdcb

Observation b840d0aa-88db-4ee4-8d7c-aea8368770aa · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.914942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.914942Z digest=sha256:cede4300d2d4e455970843c838a7408bf4921fd8681cec0ae1ae568afbcda432

Observation eb46d933-46d4-4af1-a71d-31fb270055d6 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.920134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.920134Z digest=sha256:cad21e743ecedfcde8eb3adfbd4575272f7a97428b105663641af9d962c72054

Observation 9c80213e-e434-4764-b9a1-dd9b0fe19e70 · outbound

This paper cites VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.926180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.926180Z digest=sha256:eea26639d22287ff89e3d1fea08790bffa7dbc9f4044d7e43f6f378be04181dd

Observation aea25225-aca4-4e11-a217-6a2b89ddfc92 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.931710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.931710Z digest=sha256:339db7085d850d7d72086d7c3fe2d6e5ca1d1b231b0d4008fac71f7bd5a5db4d

Observation 2d5e40fb-ea29-4f77-b6e6-136cf38cf54c · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.936440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.936440Z digest=sha256:fbb8611bb107ced88d9cf543c4874df208d2aeb8e81a87d430c0a0058ef6cbf0

Observation 3b2b7650-77ba-4a4c-8f8a-f9cdadc85895 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.942045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.942045Z digest=sha256:6e190eaa97e24c6ba1998c5afb6744ada6d435d78b0ea983d63a0bb8245ea5a6

Observation a4ba5982-1ad0-4cc9-a6b8-e41474e70ca2 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.947616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.947616Z digest=sha256:692d3b07d69b321e24704ef88be75288bf5ed8dd2b707cba9ff490285ba96602

Observation 8d24ae44-7827-42b1-bf90-4ea7efc56487 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.953253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.953253Z digest=sha256:82f9c13ad9ca53a70d8d468a2f3b69a9e67c4d12a00e2d6d569511579908a400

Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.958836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.958836Z digest=sha256:822ddda943b9f31ed9965089493965759e9e63f444cfdd2d55e5944e18f6224f

Observation 9d7afc8e-51e1-4898-9593-28a73edb2c14 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.964837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.964837Z digest=sha256:c617bde5e0444d7777a295a2a5793b8cb50fac9725ec11d43aa853f91b3f6656

Observation b17e8c60-dd13-46b4-9bf3-6a46a0cd282b · outbound

This paper cites GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.969937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.969937Z digest=sha256:c53b4264cdf693e4ec0cfa9924c180094182f64e34e869a8e08ab713d375ce33

Observation 30cc9626-792b-4408-b9c4-d300a389e698 · outbound

This paper cites Mmvu: Measuring expert-level multi-discipline video understanding.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Mmvu: Measuring expert-level multi-discipline video understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:41.155334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.975792Z digest=sha256:0da67782af1a36b5449a7885ce81af842af5527936554ea405215aa0d5f05569

Observation 93c03e7c-9b72-4505-99e9-f8fe3882c242 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.981501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.981501Z digest=sha256:e71b21cf3cb7cd0b20e188d9561da98d3b6eeb8bb5d6ea9498f46763f3eeeae5

Pith citing papers

Observation b27b417a-d61c-409b-9100-0b2d4f172ddb · inbound

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models cites this paper.

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:31:13.050244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T19:29:43.382392Z digest=sha256:b9ca4ca0a53760c37570c77fee4844467bb0a58e0c9ded05eb2751732c43ac85

Observation c68e6a67-1d50-4550-87ab-dd7ed69a86fe · inbound

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models cites this paper.

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T14:05:25.733973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:05:25.733973Z digest=sha256:a5c6082198c9439556209a232bbb666e2a7caf73a57481957c5a8ef59c017ad8

Observation 9bdd290c-5381-40d6-bcbd-09890ecf0064 · inbound

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate cites this paper.

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:30:13.790750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T14:29:15.751497Z digest=sha256:c2f093a8ed0d8c35cead16b97d5d68094fde156355fbf15c10d842f185ea71a8

Observation b1a125c1-20ee-4954-a735-d66748f1417d · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:51.701504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:acce5bb3cce2808c98ec41ac90de7d81612824f2e55cbd5c756081e25721c642

Observation d21428b6-34ea-44c6-b988-a1c1f5f3883e · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.934946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:87f5fbe1981972feef8e69a34acb5b6450e2c16212210a0d4f16d2dc442e02b0

Observation 9a247ef6-a82d-46aa-b669-1071141cac14 · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.805123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:4cbfd17ae655fa46b46556382bc19ba9ba44e723d533e47e6ee10dde8a1aa9a1

Observation f6e06dc6-c571-49a2-9719-de29a1412d78 · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.768132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:df1c7977753efff3fa36362ba5a2299b7875aac3ceb8e86a763ffb905c7d0ef7

Observation eb90c0cd-3b97-4812-80b6-710292ac054a · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.496389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:240ca4f399caa26385a4612467ceffec2d9bf4e68ed882dc5b7a208cbe610bdc

Observation bc854a1f-8dce-4537-9616-c1aebfd183d1 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.651999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:b87371c547561923bc97cd91754ad2e2475d122f6c6637ade68584b3f79f46ed

Observation b631b8b4-6629-4e9c-b17c-c41a0306dd70 · inbound

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety cites this paper.

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.569526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T00:03:27.948225Z digest=sha256:85929aadfc3a0f31a83cc23b2c54c52ff0e44c9c410af4afc50a845e15c33876

Observation 15ec6608-c26e-403a-8397-d6582a6b4fc3 · inbound

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety cites this paper.

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:48.808047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T05:16:19.502837Z digest=sha256:02cc6e8ea04d08cc20513d5785029a7d86bf9ed3902d7fe5fa65c16932ddf67d

Observation 98c9f657-6b68-4eb6-8288-294bb38a9e9a · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.208153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.208153Z digest=sha256:5a1988df3d9340493a9ae027f87a4a68d0b5f49a4783d0fa88e17146df303912

Observation c5a3c139-f196-4cae-aa08-cd8e0f2cc00a · inbound

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift cites this paper.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:42.427244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:42.427244Z digest=sha256:30738a6adeae958be11824bd2a69bf207b39cabef0ce773384c3535ec365df61

Observation 456a6657-a62a-4bc1-8d2a-d16f12c85bb5 · inbound

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus cites this paper.

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:56:01.952589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:56:01.952589Z digest=sha256:a67880e5981efeb76abedfc795340bb87d17c4e67b573a0e11526f70c13c6b93