Pith. sign in

Paper Citation Record · LEDGER

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.09931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09931 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:44.291837Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ae13e2f6-3382-48db-9a68-58fc8f6be3a7 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:45.105840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.093580Z digest=sha256:03a64de258ef3efaf42c27f4fbfd864448933b52ab9a4f1706f1307bf3311fb1

Observation 80076d7a-3418-46d3-a740-9737c0d2d841 · outbound

This paper cites Qwen3-VL Technical Report.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.098478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.098478Z digest=sha256:e53c0280ef3a5f5a312c2960efb7468a59ddba27521096eb11044e9b4a9a2853

Observation a900038d-cf01-4383-ac3f-4711196d3526 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?,.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Are we on the right way for evaluating large vision-language models?,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:45.092767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.103068Z digest=sha256:366b5228ff73a4eef227c2e5c26a5ae155e73cf396b7870874a7b825515c433b

Observation 5e56d924-b1b1-47f8-ae15-b64861b2dc4a · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:45.079436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.112924Z digest=sha256:2714a7cae0cadda52f21bfa38e05c762baae4f3ea741b0720c7bb1d098fec047

Observation 5b5082b8-5387-4225-a25a-dbce52dbfc6c · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Reinforced Self-Training (ReST) for Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.118064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.118064Z digest=sha256:02e982376a37720c86ac312a7b081a5ce721e8cc9a18600463a0cdfe49135527

Observation 69a6b4cb-7eb2-4dd8-9b2a-095e2666057e · outbound

This paper cites Active-Zero: Self-evolving vision-language models through active environment exploration, 2026.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Active-Zero: Self-evolving vision-language models through active environment exploration, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.122539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.122539Z digest=sha256:62657beb22805d570179ee1674baccdddea02c641ce25f1b527a206e93732c98

Observation 01653e1a-8049-494c-a812-9c7320c6e501 · outbound

This paper cites VisPlay: Self-evolving vision-language models from images, 2025.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots VisPlay: Self-evolving vision-language models from images, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.127261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.127261Z digest=sha256:f8d912837b7f820192b701d6b1cd64b6774b6042671bd80b1ff061b8f2d6a8f4

Observation cc1008c9-cd83-4da0-8597-d3369fc394c9 · outbound

This paper cites CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T04:17:44.764322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.132452Z digest=sha256:aaa27a8fd4d0e86551892f69c1215e011b1725529409f12efd1315842c1360d4

Observation b41e6217-f389-4304-81a6-a5c94676d9bd · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Distilling the Knowledge in a Neural Network

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.137755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.137755Z digest=sha256:05f8926c99c6feba8aac667f7f5212cda91de58c45400ad3207af6c5abaebc79

Observation 3c49111a-f39d-412c-ade2-2addec5408ec · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.142454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.142454Z digest=sha256:bcfabe38e65467a8947c4c71d9355773e1a80c89ffd1cb58379804a7569a9d2a

Observation 9de25cb6-c398-4287-ad24-2068d9dcdbba · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Reinforcement Learning via Self-Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.147331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.147331Z digest=sha256:9f38c2802f225a56d05c39acf1b00b76e8d5d01e0abd654ccc2e85a03e8a244b

Observation 69b91893-6b02-4442-b4ab-e69624456fc7 · outbound

This paper cites A diagram is worth a dozen images.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots A diagram is worth a dozen images

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.152810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.152810Z digest=sha256:616ad4aaf0050e53f852934255d9d3f7165d76d1ac9ef3dbad9545d609c2e75b

Observation 6ea2943f-50cb-4a27-aa55-9a70356eaf12 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.157228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.157228Z digest=sha256:6da9e8c5c8784ada0ae3e3084d4b7cfe24b60bcf0a666d849b8075d3817992e2

Observation 4afe182f-1200-4888-a95f-4f39c5a8fb0d · outbound

This paper cites MM-Zero: Self-evolving multi-model vision language models from zero data, 2026.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots MM-Zero: Self-evolving multi-model vision language models from zero data, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.161858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.161858Z digest=sha256:475485d7c0e10c2aae0c674c8779e7dadb4725ed73a2b9e2c47383b05f3e4103

Observation b5b85fbf-b64e-4942-b5f0-c960b968d35e · outbound

This paper cites Visual-Advantage On-Policy Distillation for Vision-Language Models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Visual-Advantage On-Policy Distillation for Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.165871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.165871Z digest=sha256:1573d7fd6151487d7bc1ff2d60f17a87cb0a896d8e1ebeaa477b55fe32e09454

Observation 706078c6-568e-4c34-9217-0d157d314475 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:45.053778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.170330Z digest=sha256:99fd20b29ffddb6096606490c306937f6f76abe79b9c6ca3999c9a6e0352969d

Observation a344f4de-710b-48ed-8232-b2b26ba2238a · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12), December 2024.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12), December 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:45.040516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.174470Z digest=sha256:5768e6465b97b580f4ce65c2f95f6958390f3b895563b9c133aae4b34120b68f

Observation 7736c1dd-ce60-41e7-b50a-8a264bae6921 · outbound

This paper cites Decoupled weight decay regularization, 2019.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Decoupled weight decay regularization, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.179553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.179553Z digest=sha256:4904486d1d5276e14c6a72ff5bad104f4df2d9f377c886f7a38f669e15ad7a83

Observation ce529e7f-a892-4f5e-9b7b-3af3e8aa0250 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.183963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.183963Z digest=sha256:3e2f27c388f761b1e8f66f203b107ece5ab2ec7d7dc3a0cad9b09d0d56cfc101

Observation bef55480-7c01-4a8e-b997-1fa6b4a33fb8 · outbound

This paper cites an unresolved cited work.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:45.014002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.188690Z digest=sha256:0826e88183f331ca55fed5ab6a3cded47fe1c83a6634fe219245343dfca48f1a

Observation 07284d10-0b13-4f2c-9923-4fe1077c87b4 · outbound

This paper cites Gpt-4 technical report, 2024.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Gpt-4 technical report, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.192669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.192669Z digest=sha256:d5a4cb3e5023d3fdc369984346076976cfba2ca9f544b55fd852614f8eac0221

Observation b9233527-fcde-489f-9072-e20468c0ab45 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots A reduction of imitation learning and structured prediction to no-regret online learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:44.992401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.196313Z digest=sha256:46ca3123ec14603a079a1150c28d77a6d15202e5bc5700ab69541fc9b4452ad3

Observation f3415c4b-9aed-4d6b-a6c3-c4406f885d5a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.200564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.200564Z digest=sha256:1a85e631822beaff7d5074756a666243f5ecdedd51f06b5fe1f7b6ed11a84ffa

Observation fbdbf1c4-2014-4070-a796-b8ef90a1612e · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.204392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.204392Z digest=sha256:20c148fd56bf1bb76abc7827a93d2480e5c487e0f8c1d578c96248f8d05ff0fc

Observation 72feca84-aa2d-4986-bec3-5b004ca5d1ef · outbound

This paper cites iReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots iReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.208861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.208861Z digest=sha256:6178feffc9b37fc5db43401993ea1aceb9b629bc99497624c0eeab69f7d1abe6

Observation ef8e20ac-bef1-4ef3-92fc-d5d834df0948 · outbound

This paper cites EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.213416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.213416Z digest=sha256:f98ecac6a58fe5deddd63f00f104b7db4610d99cdde09522528fec932bbd4948

Observation 62df7afa-b1e8-4b9e-a724-3fa77d74b28f · outbound

This paper cites Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:17:44.577412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.218120Z digest=sha256:d152c8b84241f9d78c1c2fcdd92a333974f5948b25204501b62a7182988c7d79

Observation 27ce28bc-f464-4c95-ad26-d3b1fccca431 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.222018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.222018Z digest=sha256:48efdf65cc8464599f07339e3bf68d18a00110ef49fef3848f0df989f4a36c2b

Observation 22c0bd89-a4a7-4c3e-b24f-bb1932478b05 · outbound

This paper cites Paying more attention to visual tokens in self-evolving large multimodal models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Paying more attention to visual tokens in self-evolving large multimodal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:44.970556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.225755Z digest=sha256:98e65d31f725bfd1d83665d1ee84b898c21300fd424213a876749a533a159d1a

Observation 7258add1-d78c-4f86-9fca-25162a9e1008 · outbound

This paper cites Vision-Zero: Scalable VLM self-improvement via strategic gamified self-play, 2025.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Vision-Zero: Scalable VLM self-improvement via strategic gamified self-play, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.229522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.229522Z digest=sha256:38c3f1ca73597446b660233d61c58e35f33590dc63beb60d3c776918ba68cf93

Observation dc8dd3ec-c85e-424f-b562-1b6f99920aa0 · outbound

This paper cites When models judge themselves: Unsupervised self-evolution for multimodal reasoning, 2026.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots When models judge themselves: Unsupervised self-evolution for multimodal reasoning, 2026

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-11T04:17:44.495048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.233726Z digest=sha256:88bd0253068db3fadc7952348615c3ea746c9b936e6920eb7845c02329b97b0a

Observation 72dad015-8635-4e31-9257-0451d72e240a · outbound

This paper cites Realworldqa.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Realworldqa

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:44.956552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.238033Z digest=sha256:ef761e0e3bdf739bbfe90edcd31eb0b1ae58f3cfb597f755097cac5984441e0e

Observation 5c91bc7b-6a14-4e5d-a237-a84c4d75b3a0 · outbound

This paper cites RISE: Reliable Improvement in Self-Evolving Vision-Language Models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots RISE: Reliable Improvement in Self-Evolving Vision-Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:17:44.403579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.242315Z digest=sha256:0cb6e10060e4138e1fbf829e6dc8ef9e1fbe2d71d1c8742fac2f8297f7b32047

Observation 23b8ef9d-c767-4ac4-bfc7-c32e96dfcd55 · outbound

This paper cites OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.247249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.247249Z digest=sha256:d9a86806274a2dd082929a6d4269e0fcfd1e1db81d2d5f0c5b9d151ca616c3a3

Observation 639c3775-a925-4144-b8bf-4a3764fe13cd · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.251370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.251370Z digest=sha256:c0e13c6d2da1d493e485da1ffc3ecccc3ef12b048e12159ddacab88ea9716fba

Observation cd36e72e-bf56-4a99-933c-dce48774ef01 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.255750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.255750Z digest=sha256:1db7afcbced77579ab0f1e7716af8c4eaddd77ca5f9cec5012f1aef7ba3628bd

Observation 736a200a-4218-455a-a978-591c4b01ac2c · outbound

This paper cites Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.259723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.259723Z digest=sha256:a64caaebe5a8258f9b3c5e06b6fbde5b80b0341b46fdc4ae6c51212f74aaa861

Observation 5594c7a9-0988-458b-a9d3-06cbc35e7371 · outbound

This paper cites Self-rewarding language models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Self-rewarding language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.264323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.264323Z digest=sha256:86e54c503c0f9002f2d9b1a8d89aa147280b6dc2c0d7197c06d5a817ecba0b10

Observation 49f1436c-0382-4bef-bf56-5b4c584e63dd · outbound

This paper cites an unresolved cited work.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.268126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.268126Z digest=sha256:b4e09b6c56105b49d44e6b705c9fa9424c02dc88b379becc06ece6808e7b46ba

Observation c3343f26-59cf-465f-b162-c28ac1924ac2 · outbound

This paper cites R1-VL: Learning to reason with multimodal large language models via step-wise group relative policy optimization,.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots R1-VL: Learning to reason with multimodal large language models via step-wise group relative policy optimization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.273112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.273112Z digest=sha256:1d59804516bb04d87711ae4a90282705333558aa5c96c18f92baa28e82362aed

Observation ed80c698-ea8d-4b35-b1ff-30817ceab881 · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models, 2024.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Lmms-eval: Reality check on the evaluation of large multimodal models, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.282830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.282830Z digest=sha256:a53e8c17b3f10e2a9255b238f757f402998c79192a34d46d9b121f562d4e344d

Observation 0dc928d2-923c-4d82-a5f8-9f024a03d796 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.287143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.287143Z digest=sha256:c226a84a261b578a99a0243350d87f96185b955f70c6ee2d58a8762d123a9a18

Observation 8bb9a19c-d6a7-4937-976e-735adde98c7f · outbound

This paper cites ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.291837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.291837Z digest=sha256:882367aa9a5527d52e45516905478cd76d76c994419dcd69b822745c746c89fb

Observation 756a531f-d99c-430d-9d57-cad1e6e2540e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.107841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.107841Z digest=sha256:332b9a7f5a7ad64fd1700dac41525b4685b9cb6cd3541d49636b1a513bebdf88

Observation f9102c83-a3ba-4dd7-8c74-3b0eb7023be1 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:44.278544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:44.278544Z digest=sha256:b59e6f5ee1bf8e8b65779185710d6d7926deaf5263881fb39b7aafcdf27239f6

Pith citing papers

No inbound Pith citation observations are available.