Pith. sign in

Paper Citation Record · LEDGER

Learning Self-Correction in Vision-Language Models via Rollout Augmentation

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2602.08503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08503 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:21:13.972406Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T06:45:27.857034Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T13:56:19.149879Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70d0a3f6-bd21-4ac1-b7cd-7906b79348f9 · outbound

This paper cites write newline.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.828338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.828338Z digest=sha256:1528f3b72565566277c5d7c052510b044d1d1a33bf32a3eeebfaefd9824ee5a5

Observation 205c4f7f-7bd0-4aec-85c4-f305160be02d · outbound

This paper cites Claude 3.5 sonnet model card addendum, 2024.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Claude 3.5 sonnet model card addendum, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.832568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.832568Z digest=sha256:ac71d2e0bc008cf80f817effead80b27e9d162904c52303a2d98c371fe47cbc8

Observation 57b51591-3107-47f4-aa48-85ae14a2c6fa · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.836456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.836456Z digest=sha256:936ecaed10a4afba84de8a3d902d5f446bd9a611ac8bb12056a480dab160b3fa

Observation f67dda4e-51dc-4deb-8b03-fe056126c410 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37: 0 27056--27087, 2024.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37: 0 27056--27087, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.840209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.840209Z digest=sha256:aa2cf741503578d0f1eb04cd94918ccd23a125a0e4c5a926adf6df5f18ffe1bd

Observation 4edfbaf5-67a7-4414-a103-f4cb6245babe · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.843641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.843641Z digest=sha256:abe3183245d78e82e0b44f115b681845db00c88a186998599e9d8530ea568071

Observation 5600bfb2-b8fa-4f5f-94ec-feae0a016c92 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.847157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.847157Z digest=sha256:eb6b3cb3b902c4e62953da490c4bee7b1c2342901cd8ad0870dc8c050a64039f

Observation d32c93b9-8e3d-4f66-a9b7-221f9d1ef093 · outbound

This paper cites and Zhang, R.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation and Zhang, R

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.850757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.850757Z digest=sha256:0a6575daa2d2c68afbae031c34cb7f8c17faf2fc10cb89628bbacb22aa31e871

Observation 29de44e6-5691-4c03-9cd4-d874e2091a90 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.854074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.854074Z digest=sha256:b6c0f878ac77e47436f0c7b9269079dec8c8ca66d31aef514ac4cb6c4512d639

Observation 7c2899e4-2dc5-4a85-9214-117d522bbe4c · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.857385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.857385Z digest=sha256:48e22f452f4596fab355fd8ada0b445f2ab62afa9aecbdb59f2bbfa05f4160c0

Observation 34583ee4-ee0c-444b-9d70-25bfed191d29 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.860519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.860519Z digest=sha256:33c3765b5d1f79f89050c8b0759edebbaa1593aa647b212565ef1a76e1891a77

Observation 02e915b1-c889-4e9b-9790-b3b84e4d5e7e · outbound

This paper cites GPT-4o System Card.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.864335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.864335Z digest=sha256:ed8506d4c33f74742a9a0405323b5d969a359351e988e8fc39211c5526873bea

Observation ffcab89f-e46e-4187-b1c6-ef46b7c0af0a · outbound

This paper cites OpenAI o1 System Card.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.867754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.867754Z digest=sha256:e098596f9d720e3e1226ab33d7badbb2265872dbbe12ed53fefa12faae0bc9e9

Observation d022775d-2220-4123-95db-46bdbddfefa3 · outbound

This paper cites Look again, think slowly: Enhancing visual reflection in vision-language models.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Look again, think slowly: Enhancing visual reflection in vision-language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.871940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.871940Z digest=sha256:179f0aadfc766c64c9b9b9c5562deef49e68bddec863c281dd932079809faaf1

Observation a315bf8b-9604-4635-b73d-915505ca5c78 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Training Language Models to Self-Correct via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.875111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.875111Z digest=sha256:b0248f4308a32f1913e494e63d83571f8aad80f3bb5bb566fbf220f0a2f10dec

Observation eac0620b-09e0-4fbf-aab3-ad8008beed43 · outbound

This paper cites H., Gonzalez, J.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation H., Gonzalez, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.878279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.878279Z digest=sha256:79a063a092e8e55244b247714f58f3d8b085913df614f34955516b2c9e43f24e

Observation a9fdc6b7-6594-45f7-a494-178c7ef5778a · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.881298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.881298Z digest=sha256:9ec7a053925eb0aaec1dc042e7a33268e24ecd0bed3d1ea04254801a091ac8f5

Observation b8f5a1e8-506d-4662-aba7-03be9d80e269 · outbound

This paper cites an unresolved cited work.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.884637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.884637Z digest=sha256:44248c83e21c80f07216f5cb05f0c964e25b4e4bb02ed9879a963ba55d34cdcf

Observation 42dead94-8c66-4fb5-b628-9b550f2e79cc · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.887624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.887624Z digest=sha256:224bfab0b1bb93811cdaede1b52b606c682128b4225152cd580b64b61472ef1d

Observation 059a77d4-f74b-4a8b-a71f-78d97344c433 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.891083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.891083Z digest=sha256:735d02a7d33ef17d893fbdd845a3f93893e6c43725ab8398aee38851d4b74d30

Observation b43e7b87-370d-4494-844d-54a9bc84120c · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Self-refine: Iterative refinement with self-feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.894328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.894328Z digest=sha256:c2a637c1cce074c763f30d5754f58f701a9a4dc8293efcd46a527d48ea27f6f9

Observation 42d150b9-d757-4e55-b8d5-847ca62dff35 · outbound

This paper cites L., Tan, J.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation L., Tan, J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.897277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.897277Z digest=sha256:1836cf80f08b04ec94953779a8620880477c6c8bce66b9ae2e3c49426bffde57

Observation c85aff74-3635-4b8a-bb67-dcca7a2768bb · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.900352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.900352Z digest=sha256:7c28a8aea5d57468c6d41cc0c6624d6e751441a5ea71020e62dc2bfd640fd88d

Observation 750a83f8-29ea-4732-b46c-a8bf25f968a5 · outbound

This paper cites an unresolved cited work.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.903784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.903784Z digest=sha256:45450ddb339938135e5f1d624e19424fc4ff6ea9fb537a887a3996ba65b39e94

Observation 0a0de85f-d0fb-4ac1-8123-c374e7e4026c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.906826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.906826Z digest=sha256:d7b8eb99624d47c319c068df7a96fea3bab5a8188177350034d3da768c730c68

Observation 3aff086c-97a1-4861-96b9-0f125780ac6a · outbound

This paper cites Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Srpo: Enhancing multimodal llm reasoning via reflection-aware reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.909940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.909940Z digest=sha256:ca739e10912f7421e5a9bf095e6035988f987b1db08f321232203c6917b6b05d

Observation def79938-6814-47ff-95bc-e777b88801a7 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.913033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.913033Z digest=sha256:e7e57200315b50c43935efb82467a19eb153532fed9c1ad3df46e1a8340d45aa

Observation 4fde56f7-fc30-47f7-a5b5-25c842b4b746 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.917664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.917664Z digest=sha256:827d6cd93a7801103d261717b4317d4806c80d5c172d42f6dbebd84bb56d6ab2

Observation 1478b99e-35d9-45ef-b481-a9237ec9ca09 · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.921539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.921539Z digest=sha256:b5dfdc6cd4f6ac7c853ece4432852472f8007052ca4edee17ed02eb93d4555cb

Observation 76033a86-4052-42d2-bb78-d5c307e40699 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.924847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.924847Z digest=sha256:9f1c5eeabc234517f7fe3b36250b1543fac0e76afb8f7de445bc08b579f59d09

Observation c065b604-b9e3-4e69-84f2-dd8ab5341b75 · outbound

This paper cites V., Zhou, D., et al.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation V., Zhou, D., et al

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.927695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.927695Z digest=sha256:0ab132da73d8bd579a87f68ddba57571eabb45c822aad312fec7a1050f3b33a3

Observation 9f537819-a78b-4e4e-8494-e6c407fc8e08 · outbound

This paper cites MiMo-VL Technical Report.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.930637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.930637Z digest=sha256:1ee720b8b8cc1900ba484e43863376da78dc7901c0e46a75f015999be30b48b5

Observation 120a71e1-6f58-4d1e-822a-b6a170ba68d8 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Llava-cot: Let vision language models reason step-by-step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.933805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.933805Z digest=sha256:f9a40395e881956368a671c49d4f8969ebc7bbd0bb31233e6191bc09dbed2db5

Observation f3896ec2-27d8-4c99-918e-7b4aba42deeb · outbound

This paper cites Qwen3 Technical Report.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Qwen3 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.936891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.936891Z digest=sha256:0819934e2f303b4274786bbb5a473acc7606c3efe3482dd238a8d4c5a0444aa5

Observation 8f6e4d58-357f-4d24-a45c-e2b88533f0d7 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.939966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.939966Z digest=sha256:8d271f166c7b6e9800eb327048922934af433c081ec9a6f6ec2426e52951791e

Observation aaa2445b-b441-4f51-9e5b-f79ef4f19982 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.943200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.943200Z digest=sha256:b48474e851e5ab70286808302c87743fbd69b0a216574d8979cdf93ddd2718bc

Observation 66bcf109-c5c6-4e75-b892-2d5608f460bb · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.946237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.946237Z digest=sha256:1824ab7522bbf5a7862ec8f2d1daedba7824ad562b421070995e86eb625a870a

Observation cda9b2f5-09e6-4d90-b771-8bf4a4e68d19 · outbound

This paper cites Evolving llms' self-refinement capability via iterative preference optimization.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Evolving llms' self-refinement capability via iterative preference optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.949689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.949689Z digest=sha256:658259fc0985da241ce6cea6c2299e865202f931ab04874456d445e26d509f0d

Observation 775c1e07-9e2e-440e-aed9-28be4a1a3986 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.952698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.952698Z digest=sha256:efff44380314bdc5c164324c44366b618ea5f8e51df2680071a8277d25008c86

Observation 8f22f2d1-250d-4936-8c9a-e3f03266fc8b · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pp.\ 169--186.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pp.\ 169--186

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.955811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.955811Z digest=sha256:e94b099a5f48396d0a9f6158ecbc710ce236f26f9b188c6351c20cc46a04013a

Observation ace3ceb4-62cd-442f-888e-6f48df7e1117 · outbound

This paper cites Small language models need strong verifiers to self-correct reasoning.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Small language models need strong verifiers to self-correct reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.959015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.959015Z digest=sha256:1612f5afacbbfd5a273a0e74e938fd9fb53fb6a81bd20f543e4c2eed934be943

Observation 5600fd11-4511-4c33-8d85-fc507c4a1722 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.961875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.961875Z digest=sha256:020440b5d385ebc471cf60de0e68d0a79cde1dc899f813cfbb3d8ce8767f9663

Observation fe9308a1-b17b-4d04-9f31-13d65d4fd0f6 · outbound

This paper cites Group Sequence Policy Optimization.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Group Sequence Policy Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.965141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.965141Z digest=sha256:2bb5ef8a4ceb759b3fffe506db65397bca7ce5b3c3e2e476688a42ea455cf6c4

Observation b05f92e3-598b-4804-9a38-39866bb2bc59 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.968926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.968926Z digest=sha256:d669cfa17a88dbb78176100c569295349cc0108b8ca80d9a989280ebfd9ab7bf

Observation c7f138cb-58e0-4824-a831-82b22c490c80 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

Learning Self-Correction in Vision-Language Models via Rollout Augmentation Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:13.972406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:21:13.972406Z digest=sha256:2c11dd3fafb5549ef451c29fe082194685848b2406fea54fdb1c37bf9c6c0f37

Pith citing papers

Observation 787af9da-b4b9-412f-96b4-ce2c6b31134d · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Learning Self-Correction in Vision-Language Models via Rollout Augmentation

Reference 108

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T13:56:19.151657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-09T13:51:49.149342Z digest=sha256:4a6a232740ed05a07e096ffa8c8c86f5a8cb769abd3762aa5c1e7880d2d44951

Observation 5090e2b1-e65a-45bb-94b1-ff5bb52dd9a7 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Learning Self-Correction in Vision-Language Models via Rollout Augmentation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-07-13T06:45:27.857034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:45:27.857034Z digest=sha256:8bc6fc00eeeb8b7b84e4dca92dc77426352b41a26ccd1531fbc0edcf45b61c5c