Pith. sign in

Paper Citation Record · LEDGER

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 5 inbound Pith citation observations for arXiv:2506.10128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10128 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:40:13.423620Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:39.864757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.446166Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 529e2bc9-2351-42ea-a4c2-476f671be824 · outbound

This paper cites Vqa: Visual question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:18.085418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:06.644920Z digest=sha256:a495e03b3e41a676cc17616c679023978f1cf0a0202506c956d2bfdb48a36319

Observation b187ba96-3bfb-41b1-907c-2d4870c92b3a · outbound

This paper cites BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.691389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.691389Z digest=sha256:3db4d230209807525916d57aa6306a0161652ec073c3fef921d6aac4ac53ebb9

Observation 295ab65c-4f3b-437d-a1cc-28902ee60831 · outbound

This paper cites Qwen2.5-VL Technical Report.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.769409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.769409Z digest=sha256:978d8bbec159e663d69d4dae555dc8aefd4811a53e8befb1804cfa9d2cccf8e7

Observation 0b6f18c6-6c7c-4c55-8a25-776297f6b43e · outbound

This paper cites Improving image generation with better captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.852093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.852093Z digest=sha256:1d26b71e3887b91c8447f4f56572182667e2c7c38ca8dc3af80e38e35185f2a1

Observation 61c4cd51-1079-4709-8a57-8f561729cab0 · outbound

This paper cites Language models are few-shot learners.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.950928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.950928Z digest=sha256:ea41b9d365450d880c9c40b19569947320ceb7608a346972f582d02d20ecc8bf

Observation e148fb5c-6010-4a9d-ac4b-2882ddc5bda8 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision language models with less than \ 3.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs R1-v: Reinforcing super generalization ability in vision language models with less than \ 3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.814290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.055773Z digest=sha256:2a07f8f59035429e6c2cfb51b6df2e1f537307211fd39ca40a84d23ac1b35800

Observation b165f18e-9ec9-4319-b4ae-01800a07b4a2 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Sharegpt4v: Improving large multi-modal models with better captions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.479967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.167181Z digest=sha256:2d552386e4be0e3e3bfbf5b4b578bc71e2921501e8bf4a54bf770c1c06e28527

Observation 16c87061-4b08-42b8-9650-ddbf4263dfbf · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.282039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.282039Z digest=sha256:c010274290665025183c044a58162871c2e93c4c5dfd0e67d1c3b3b79afe4c1a

Observation 4278674d-e3d4-47ce-97c4-9aa33a023f0a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.352289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.352289Z digest=sha256:51d5209c0506ac1b5802728e5e7919a08a63023fa0702ae9da2058aa55c3ff92

Observation 0f547de1-4259-4664-ae43-04fc56b9f4fd · outbound

This paper cites Palm: Scaling language modeling with pathways.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Palm: Scaling language modeling with pathways

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.423678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.423678Z digest=sha256:5424179f6d41ec0d9826618b13c126be6af5dac321093612b64931dad0e8cfa3

Observation ecfc56bd-d7c3-441d-89cc-0496c50d2888 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.530509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.530509Z digest=sha256:b3c466de48f0283450451db0ee1bd44b907f45129bec1d2124fba8e6aba4b09c

Observation dbfc9be6-f678-4ffe-b456-45023e87f0ef · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.610397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.610397Z digest=sha256:ce936e7e3d4b1f9bbd360030ad26872cb837118cf924cbb504daa47dd30cd2cd

Observation 6d3f5767-fb1f-4614-8730-bdcd5fdd560e · outbound

This paper cites Enhancing large vision language models with self-training on image comprehension.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing large vision language models with self-training on image comprehension

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.196649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.695543Z digest=sha256:bebbfb41f60430ef6fbc06c8d7c8eddd72801a44385ce575e4ff537dc76e2089

Observation c9ae7c9d-2336-4ce8-a60a-05d82a6a63b4 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.764980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.764980Z digest=sha256:d6b77552f20fe7ef21cff7a507819a7c6a1f6b4598fcfaabdda70ee91fcad512

Observation 0af995ae-2b43-48e2-804f-57696d41c56e · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Virtex: Learning visual representations from textual annotations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.922173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.899067Z digest=sha256:033ca0527424a2a85a691038944f6f8838fef1686f12ea08d75b8fdc1f704689

Observation 22bcc49e-abe1-4239-a509-8b6c2363d5ac · outbound

This paper cites ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.980626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.980626Z digest=sha256:bc2fa367ed35a88a21bc6b4985bfed2b575c615fd2ed64852f1da5b5cd630c2b

Observation 57b69c66-5c8e-451d-b509-23db3c4f7fb6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.094721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.094721Z digest=sha256:b5daf32f50c22da9b7dd62ba1602a4eb0f3cac6eb6933d683b7341ad745dd9bf

Observation 9024f7a6-09d5-44b8-a2c3-b8d1e33ac7af · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.211304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.211304Z digest=sha256:6e6c4276fcfe0aeb8810f8f7d68f29da9b2f06064ed66e16955ff34e38a2732c

Observation 4a3b41b7-8ec1-4758-af92-85adf205d4fd · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Measuring mathematical problem solving with the math dataset, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.696254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:08.309188Z digest=sha256:62507f4c3a10b561f871d4a78bc16fb06b5cca53353a5864df631ca8e92802aa

Observation a280af47-a3c9-45c7-aff0-7c06e02d297f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.388160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.388160Z digest=sha256:0ff25740118a342e41980707961336a236ffac95e1e3da9f4540a8accd0df019

Observation d93473b7-9c98-41ff-b8f5-0e2bc50c2d90 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.503288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.503288Z digest=sha256:16a7b1f7b88360df9cf490ded69806cf831c82a48e587cbe9ea9baef679c5dc5

Observation ff924b2c-8786-46bf-8965-ee35e28c2680 · outbound

This paper cites GPT-4o System Card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.599239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.599239Z digest=sha256:00d15a6e6de95828411e3effdf0058331d84c3b2db7353905b3099ea65d6e0c7

Observation 55ce0b4f-f7f3-49ef-b31a-191a7cb46f55 · outbound

This paper cites OpenAI o1 System Card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.696397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.696397Z digest=sha256:7dae8b8120e302ad7a1c49095d6374cb208b12359110501730497e7d43c3dbc8

Observation 01e4d8dd-532b-4c51-b697-43c1f4837be4 · outbound

This paper cites Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.808188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.808188Z digest=sha256:e6a971c17cc8680a8dd2c2172c50818f66c0b4b14de6d20cd6f6f5d8c97ccb48

Observation 4fcdadb3-48e7-4df2-a402-d1309441b60a · outbound

This paper cites Large language models are zero-shot reasoners.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Large language models are zero-shot reasoners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.912895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.912895Z digest=sha256:2931c9b02a192b69814be693fcecb6db9619ccf0a263248ec4083362394a8aa5

Observation e728da3d-5c57-40c5-aa2b-ff39f97320fe · outbound

This paper cites Mawps: A math word problem repository.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mawps: A math word problem repository

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.994633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.994633Z digest=sha256:e6841143c1dd3dabc4c13891112cbdb4be2857d4ca1f3092999565bdaf2edd01

Observation 30fd708f-6fe5-4f35-a70a-d1d0c37805a5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.073387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.073387Z digest=sha256:26cb71dafe5fde1cfc9ff406ec21500b707917f87affd4ce9ec8450f7d27d2db

Observation 38288d0a-69ec-4b66-b564-7c40f6460735 · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Multimodal foundation models: From specialists to general-purpose assistants

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.437363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.168271Z digest=sha256:3e98360d7882e7fd6cc37806f89f30b47b1ca3648f8e93f539626d43f41c10e7

Observation 9f99e18b-c7a3-42f5-8ecf-e7e4d906f080 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Describe Anything: Detailed Localized Image and Video Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.244875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.244875Z digest=sha256:02b23960a415026291bba28611840b6dbf54044745192a8f3a6ecaaa50618039

Observation 95ca5810-2016-4a10-94ba-bc9b988f8388 · outbound

This paper cites Let's Verify Step by Step.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Let's Verify Step by Step

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.313417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.313417Z digest=sha256:87b0e061231475562e6bf2919315f4ac35ca6fc12750f40630a604bdda5c0371

Observation 41609576-8914-4c4a-8c88-98c69927a580 · outbound

This paper cites Visual instruction tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Visual instruction tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.436777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.436777Z digest=sha256:a0478ec4849aa94965cfe3456b7a4c32ce705b5511886f4b3c96d41093e9adcd

Observation 290e02e6-bfd3-4282-b57a-f67afa3e8367 · outbound

This paper cites Improved baselines with visual instruction tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.179830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.523923Z digest=sha256:80d1783d4547fd092bda3d6d69f847da675bcb10bbf092a70da7e535d44d62e0

Observation f4f3460a-ac9d-45af-b995-8b5b62531ca1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.949313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.582883Z digest=sha256:8a1c6b9244558ea035c026386a444813724232866ee34497e0fb5e96af7309a6

Observation a16ab06f-44fb-40d8-819f-5b49a4350f77 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.661306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.661306Z digest=sha256:0355108490cd8fad4de1286c9acfe72c441592b73e9726c7777659f0d0cd4a1b

Observation f3d4d0d7-fb17-48da-bc52-bc06a535396e · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.721963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.721963Z digest=sha256:6be9a2e1ecdc0a168d4de950003710ff96b1ff21fcb3c109b9b85428485a2891

Observation 6bb0b974-fd24-4e7b-bed7-a382064bb313 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.920757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.920757Z digest=sha256:7049ea825a46b68e3424ce4f5c0e74d6610d99de886697fc484cf1d1174df399

Observation d39a454f-ee62-445b-b7b6-4257d893cb8e · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.021002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.021002Z digest=sha256:0af4e6da4ad41d0aa232da274a7668371760c9e68be956a3fab3356b8bab2ab8

Observation 722d67f4-93c7-460a-a6be-ff80c6b107ed · outbound

This paper cites s1: Simple test-time scaling.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs s1: Simple test-time scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.113399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.113399Z digest=sha256:531c7bf18e6f3dbdc42564661d099ed712be7d10797f15d023a3b05696001e23

Observation 25071772-0c3a-415c-822f-7b225daad8ec · outbound

This paper cites Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.179618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.179618Z digest=sha256:24fdfa1d7f2c299f6c4547bffbfeb87e07004dc6673376e7c20fdba197a58d64

Observation 2adb2fad-fb06-4f7e-8be5-9b1f66733eff · outbound

This paper cites American invitational mathematics examination (aime) 2024: Competition problems.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs American invitational mathematics examination (aime) 2024: Competition problems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.691239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.250543Z digest=sha256:9d2a07226d081bcf7a1c71a42bd015fbb2f9b724bb07c6e2e6300ddcd211d1c6

Observation 37a766c3-8881-48b5-811e-16d0fdcefe63 · outbound

This paper cites Gpt-4v(ision) system card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Gpt-4v(ision) system card

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.331609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.331609Z digest=sha256:7dac8077604a8626bbfd7492efe1fe6e1f098d21c54e454f8c49e8f1758116f6

Observation 1caae194-1900-484a-bbe1-021d5c4fa937 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Are NLP Models really able to Solve Simple Math Word Problems?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.409957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.409957Z digest=sha256:6ea98e1240b9f35b3d77d9bb685a4636a18a9b0d8b997ec963e206bcaff181f2

Observation daf7b12b-f92d-41f2-a761-5b29de60949c · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.484752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.484752Z digest=sha256:b9a6ad00e90888fba7e3ae92b9fc94b95a44a8f3499cb36320a3ff7935a7cc4b

Observation ff5c66e0-b5f7-4028-9bfd-dc9b592fb697 · outbound

This paper cites Connecting vision and language with localized narratives.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Connecting vision and language with localized narratives

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.460773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.569209Z digest=sha256:202982aba49266f1581629d396dd69962737cbea499428a31e4d9907be51170c

Observation b8f3a763-69aa-4b6a-a1fc-332c2d1c65c5 · outbound

This paper cites Vision language models are blind.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vision language models are blind

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.655345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.655345Z digest=sha256:c03a12b36e4bc84dd4461a9e451ac4a600b3e6ccf898722f007a81b0fc7d14a2

Observation 6e65327d-6dad-4ae9-a7e2-549795a65948 · outbound

This paper cites Object Hallucination in Image Captioning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Object Hallucination in Image Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.721223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.721223Z digest=sha256:fa16655bee53495b8b40988ec99e871367abc3a65df0ace0474ab9dc109546d2

Observation d6973e5d-3d08-4214-bbe4-a5764c3ead37 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.790536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.790536Z digest=sha256:29cc4787f01f2758f4cd9f338057195ab8c60a3fa37f9ecd1a21e2513a436818

Observation 5e318aa9-e996-4b3e-b39d-12ec4690f758 · outbound

This paper cites Learning visual representations with caption annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Learning visual representations with caption annotations

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.227903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.860273Z digest=sha256:e1c4558299af6579c398410d5d30c70ecd4ba7904a63d7aed42e60a2636f465a

Observation 5d4ac2c6-bece-4c01-bdd7-b0b8913054f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.915730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.915730Z digest=sha256:555863e84396827e0b10a684053aec602bdaccb00d7f41ddf0d70da428aec9db

Observation 7de2822a-dbea-45c7-869b-77e795584691 · outbound

This paper cites Towards vqa models that can read.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Towards vqa models that can read

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.020825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.020825Z digest=sha256:2463882dc0696b4c632cebd8f8241163907ae3f04424a32a465247eb0de10f73

Observation 269f59d6-2c04-4850-a77f-cac1edfdf181 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.110047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.110047Z digest=sha256:494a1dd5a201b623aea99ccff92c6af718494ae0e5f50aa90610e6084b3831eb

Observation 60076ad4-8eca-49d0-9755-2c0751f93e70 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.186415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.186415Z digest=sha256:80820c1aa2d9fde3776cb6c63d61ae0cdf0b4c78728f50c721f638bdc6604623

Observation 4d1d6408-ca4f-4847-91fc-9ff115899846 · outbound

This paper cites Image captioners are scalable vision learners too.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Image captioners are scalable vision learners too

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.018460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:11.255836Z digest=sha256:7f492c644a698562103be88b1ceea019748b34beea4383e2d96bf55baa783999

Observation b64bb195-56a1-4555-9596-aaea9b217669 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Solving math word problems with process- and outcome-based feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.367240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.367240Z digest=sha256:547673a05d6bf665df2eb59365ce2c003132b75c3c66b1d90cced551b4a0b8fb

Observation a3c4f7f1-599f-46c6-8676-c515b928c0cb · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.443100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.443100Z digest=sha256:a9da81af76660bb5a24f4e6d5153dbe9c291ef3dcecb1e1f506de0ca45179f26

Observation 8964877a-7d06-4be7-a438-7545a15f4629 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Measuring multimodal mathematical reasoning with math-vision dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.511675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.511675Z digest=sha256:839fb01115b6e6e2150a6a81e4074338f06ea2e13e35c7acf2e51b10d51538e1

Observation 6b811447-8545-46ff-98c8-3544f21aa077 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.584688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.584688Z digest=sha256:2d432559651c78563bbb4b87a281e25405fa7e4e09ccb5c18aa782b68b75ba1d

Observation 677d734b-80d1-4bf2-b4ce-31d3dbac3d35 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.660360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.660360Z digest=sha256:4ae12980394818f3dcb539d3f014a1339f726103f83aa21172c076d884d92a10

Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.741673Z digest=sha256:5f7e169f55a8c8e3e6c76212652b1c5110fe111d1c91031ec7ee25e763ee667b

Observation a9559091-c7ca-459c-8da7-146301ceae8d · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.821726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.821726Z digest=sha256:bcaf1796a0725429af0fac9ebbf496b5763e5ce2e3e6cc6a4dae7187c93c64d4

Observation 2110b634-0937-40a8-9cef-88c00df6502c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.940289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.940289Z digest=sha256:f9e0007e2921cd825409f707f908df704f90af06bb8fe3622ed1ef18b8b57cac

Observation 3b43b8ad-bc20-443f-a553-8b2bf2522a54 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.800291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:11.995421Z digest=sha256:5c0803c084e525433d5e05cd68d4e962c09c27a8502eb56e413cbd1352a88e2d

Observation 3b1d654c-9cee-49bc-a702-9e6a175e7a8f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.078191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.078191Z digest=sha256:ebf0ba42c987bf43a8a25352ccb7219d069c1ec2ed4f6806d80c05af1a059b3b

Observation 54962dd3-54f6-4f1b-b9d8-8513fe06f20b · outbound

This paper cites Grit: A generative region-to-text transformer for object understanding.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Grit: A generative region-to-text transformer for object understanding

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.545092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:12.170101Z digest=sha256:56f6e081ebccace5ba5474b6afc01be7d8c3043d4f003f9a48594e72bf5e45d9

Observation ecbdf432-dae1-4225-80a3-8e3fafb23c0e · outbound

This paper cites Mind's eye of llms: Visualization-of-thought elicits spatial reasoning in large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mind's eye of llms: Visualization-of-thought elicits spatial reasoning in large language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.273757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:12.229035Z digest=sha256:bf4698e8a11063b6c1dc681af6ea655b080ce975599c7daf1ae99862f53d15de

Observation 52378c17-6432-4034-b1cb-652c1758abaf · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.319293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.319293Z digest=sha256:6b0ceb9a31514a9ca18c0ff56523e0dc2851510ec9e94694bbf7321f10c484df

Observation b15cf5c1-0f16-4079-9e14-187d98653191 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.386906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.386906Z digest=sha256:74cc8310975389285ed2cdc8a7f4fb8f24999645f8fdec0828977e75310b554b

Observation 8006c40d-0ba6-455c-9cb8-4c07c34460d3 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.468494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.468494Z digest=sha256:a1a75fdbdd856ce4dab1b8782f4f1e393f14499531f3ef3f7e7c13cea6ecf9ed

Observation e67b9949-c7f6-4feb-aa0c-0f818da0e6db · outbound

This paper cites Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.565375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.565375Z digest=sha256:ef21bd68f5d8a9e631fcc3bdae23c356ba36a5229c3e7bb5e0c00c4d4dcc3f9c

Observation 8d3a5b4f-6365-49df-ad0d-e2fd06af8d16 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Tree of thoughts: Deliberate problem solving with large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.641376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.641376Z digest=sha256:923813d883291c0771d2561c9a1917556a75299498fa18e4e99ca87d77cddee1

Observation 8e004ac8-a322-4de4-a87a-5b0d3bacc28b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.722299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.722299Z digest=sha256:3606c77095c9e2b7c2f0b6e3737d3aca03888a5d37bda4477d82071fbd446e4e

Observation 1764ca9d-63de-47ee-a84e-1487f08812a7 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.782437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.782437Z digest=sha256:90f553d7907da9a3efa55f83c83eab9fd9afc42e6145adb78dee75f69b489b50

Observation 547fe278-652e-451d-879f-2af3e13b2ae0 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.843716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.843716Z digest=sha256:3e3b279c964a639803e68d44c5ed30ed4d17c513b1f6157b4a42ec8f761b087f

Observation 0e7a1d09-8551-47e3-bb06-3a9c18d20f65 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.908645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.908645Z digest=sha256:b1eeaff398592b4a69368a3b7458249a3f1fe67f6fd6bd3afa11403b93ddeadb

Observation da366a1c-b9f7-4218-9d91-626a390921c3 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improve Vision Language Model Chain-of-thought Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.982412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.982412Z digest=sha256:3d40266055b22d5acd62d3e240cff8c5c698bfa39d692f35062856b6b868370a

Observation 825be651-6293-4306-aa15-83976889b032 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Multimodal chain-of-thought reasoning in language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.087716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:40:13.071416Z digest=sha256:7dd2c9a9c1d3ddf6b9b4b0f3f8787a3d744d8cb2db7a1fc4c032c0204f49e75f

Observation 69ab859d-d281-4114-ba82-84a7bf3038d7 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.125288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.125288Z digest=sha256:778e3c80958703621cafae3328fe8487a2a9564836a8d785c7340be5e64a4da4

Observation 1c4c4e5f-7d75-452b-8a3a-e4fc03ffd881 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.207587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.207587Z digest=sha256:e36a6892e0b8c1d0efc8bf5004a1cfd0f73db7d1e6edb428c1073d08d4487754

Observation addd42a4-9d0e-4b7d-883a-740d7ad3fc5d · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.311339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.311339Z digest=sha256:2f4b296ce319dbfe61465a91daa0ed12f08e05cc0afa5a192e05dc8fd4923ac2

Observation 833b404f-7f94-40e2-aa27-34de5a8f59f4 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Calibrated Self-Rewarding Vision Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.423620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.423620Z digest=sha256:3344e11f8fcba377003ce3b14ca3c56a5e26a3ea6d88f4eacb66f54cd6afa7a3

Pith citing papers

Observation 9baea096-df66-4209-84a4-cc84c740d227 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.864757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.864757Z digest=sha256:a6e80d1e3d14a4f9c908c8cea8584a8097d955ea4540b48f6545f78c7902e822

Observation c1007e17-2e9d-42c4-b6b4-27490a6c1045 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.673501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.673501Z digest=sha256:0ca762f408bd6064e16d9ef9df694e07c183424a2035b5cb1e925d255dafc605

Observation 3997fe41-a744-4773-a272-8bbe15188270 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.447802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:306289e64860fda8cb8f93b4cefad339881dd8be2616c89d73f823be7b598f6e

Observation f3d4dad6-963e-43b0-a461-ffe4a4862636 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:ab128917f2c1bd2e65f23b9035f80918f3b981eff05b5bbb2465a34df390ffdf

Observation 5d28c6f6-f5b3-443a-99e4-67578dd105ef · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.425487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.425487Z digest=sha256:e2ee0e7fb4c3d1dffe13a14cc7e8d46f79e4bdc0adbdb7ed4b39fbee1e70d2af