Pith. sign in

Paper Citation Record · LEDGER

A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2504.11343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11343 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:59:21.875260Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T04:05:55.456269Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f3cfd3a9-54be-4a41-b964-576ef5d6a952 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:46:57.054014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:3e33ec34237f051a709914945c6bf67c40b3dd0c99e9d19c019d6d58f2e25734

Observation a861f957-35dc-4297-b7ef-48b40dc9a42e · inbound

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL cites this paper.

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:21.875260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:59:21.875260Z digest=sha256:c31917d1d31c2f39c477c513ccfd22c67253fa8f924e8698c462bfd6c4e0f799

Observation 1044172e-f93b-435e-aac4-4ae269e5bcd1 · inbound

Scalable Chain of Thoughts via Elastic Reasoning cites this paper.

Scalable Chain of Thoughts via Elastic Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:13:33.633613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:13:33.633613Z digest=sha256:a38fbe16790e06ea772155054c42496413cae1711f010134d93de16ba77fb2d6

Observation 7825dc23-4a1e-4b4f-85c6-2153f653c563 · inbound

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization cites this paper.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.040337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.040337Z digest=sha256:7ef897809938f29a9433bb2a86a5940cb61e163e6d78e5367ee58fde6e96cdb4

Observation ead49eaa-8eb3-46c2-b3bd-ebb996ea348c · inbound

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs cites this paper.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.530228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.530228Z digest=sha256:f89f454a32a0266682ab19c12d16cdc06aa326528d7dba6ae9a34f6fdf4d5dbc

Observation 007b1964-85c9-4886-b1c9-f9788c4ad213 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.621729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.621729Z digest=sha256:fb62cb7f156506fe298bdecbe8a8880529ce8f8c688fa82196aa815b3dcb0947

Observation 55773bad-37ca-41b8-a1eb-2acc1865b435 · inbound

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning cites this paper.

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.196041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T13:34:27.152447Z digest=sha256:d7f02d4d3b0daaaae51f3817aa2d03b3d27565b711b2f025543090d53d038024

Observation bc56ceb1-ca1e-45ab-a9e7-534e7fc3ea92 · inbound

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning cites this paper.

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:35.860086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:35.860086Z digest=sha256:eb09159bf9f6cce8a35bd268e520d423c65cc547a8f5f43f66dd2b53860ad64a

Observation b7f103ad-20e3-4ede-b2fb-bba1a62b64ea · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.370913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.370913Z digest=sha256:e6126dce99e6302fa76bff31ab7430b2b32ee42717a167f922de7c71726bd2f2

Observation 962e3e5d-9e77-4b1e-aa22-9c407beb8a47 · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.239657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.239657Z digest=sha256:70c553854040e44603f2a7a9c1eefb18aee7b6b88c370b1765bb3319e7acdd7a

Observation 93597800-d7a9-4053-93fb-d358917a779b · inbound

Customizing Speech Recognition Model with Large Language Model Feedback cites this paper.

Customizing Speech Recognition Model with Large Language Model Feedback A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:21.148418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:23:21.148418Z digest=sha256:417c44dfef58bd21be2b78800c79f78e859ff829f7e5018b071dcfa0ccab5993

Observation 130bbb4e-c597-480b-8b5d-2e281be011ba · inbound

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation cites this paper.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.118325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.118325Z digest=sha256:8e7f5044924cafe8efd7ed3860e890908a7c0c23ac96ca5af629e90696641452

Observation 02e451e0-3775-48b3-b6f6-f7ce19f80b58 · inbound

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents cites this paper.

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:16.165898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:16.165898Z digest=sha256:5e9a2a00397eec9a1a15ae94e50e5d03e9aaabb5b3bb81e5f1f165a3b53d1763

Observation dfe89364-6f1e-4824-b5d1-26b29b027228 · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:40.648243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:40.648243Z digest=sha256:19228a0aab415ad1eae50a4cad841be6b7491f35ef70c789a6962af1d0bb289a

Observation 11c90913-72bb-4836-a257-2e2879251b27 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.253465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.253465Z digest=sha256:233be5446a8b4d93b556a6697abf52f6c39ca5c0ca9a3362e6eafe59fe1e62ed

Observation 448730c8-885a-4819-9491-e48d9a58ada9 · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.525759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.525759Z digest=sha256:e9f0534559cdceee636f3349bc9b84666fd93604c4176dd0f837d9fedf7c6bf8

Observation 8784bc97-94c8-4a9b-8e9f-ca944336d9f3 · inbound

AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning cites this paper.

AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:32:59.005705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:32:59.005705Z digest=sha256:0bd3bfc1716fc292fc6989cdb69b1dd4809c80f54ac1f00edaa5ae88b8867aaf

Observation 30cb60d8-9c69-45c1-8da3-8c4b6d8d5319 · inbound

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance cites this paper.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.684636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.684636Z digest=sha256:0756907382601c0d7aede487fb442c275169b8e279b30ed28fe9f787fc96cc9e

Observation bb2d36c3-c7c9-4a0d-9b02-8641cb2cc399 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.456853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:25ce8899088aac3af4346d7aee3fc64404a0805b73abab64a36e81a254c25bb5

Observation 0c7116e4-8cdc-42c2-9271-48c094d9c972 · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.244574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:f359aa82786bbef07a71eecda86c4efcb30ff23c1c04d198f2f1308a7ffbc39a

Observation 3a903c09-9f27-48b5-996a-8758fa183abe · inbound

DiffusionNFT: Online Diffusion Reinforcement with Forward Process cites this paper.

DiffusionNFT: Online Diffusion Reinforcement with Forward Process A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:54:30.986262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T16:54:30.953199Z digest=sha256:efab0888a9e294e8c58421da652cf27c77eb0c1b0260f11999207f2d5d7d92b3

Observation 9bb53df9-d8d6-4975-bdc3-ab161ec4c708 · inbound

Simple Policy Gradients for Reasoning with Diffusion Language Models cites this paper.

Simple Policy Gradients for Reasoning with Diffusion Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:53.871009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:53.871009Z digest=sha256:aa1a8a359800b3cc2515d4fcab5dc93f09cf19ee91d8af3943b29e2314daddf3

Observation 3d40d08e-cbcd-4cba-a805-8b25870ed109 · inbound

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting cites this paper.

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:49:34.587109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:49:34.587109Z digest=sha256:808b4a640e18d42d908ce22f108f0126cd622cc56766478bc052c7dcd008195d

Observation e2e87cb3-0435-455e-8590-bda005b55f76 · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:50.580515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:50.580515Z digest=sha256:bca93c70e01ff83d58c4e66b0194c68357ff13621637c1b59d687014aee79d74

Observation acc3680d-d292-4990-a7da-29ed62b2cacd · inbound

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control cites this paper.

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:51:25.636355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T13:34:06.461850Z digest=sha256:cc16d942be5822886c8a8b8912ef525397e397a4638d4718164dfe9e99dd60a1

Observation 73b77fc9-59fc-469e-91af-c9034311ff2f · inbound

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control cites this paper.

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:32.380758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:56:51.940356Z digest=sha256:1a39f1a20d2fd1749027c7d48c3cf02f02ca08fe2ab00fb7f8ba423d35261884

Observation 1755dd5f-1fb3-48ea-8ae5-caa2491a1599 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.859103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:28eb0b05643e8675a9b73726dbc1c9d37394af07c9bca07e2a81079d4d37bdba

Observation 630449dd-b80d-4aa3-a288-bfdc3eeaca58 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:49:14.948536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:dd0570ff79fcb9868212cf52e9d5b06e34b9da5fd05005a160b85d1559b6cac7

Observation 4e3f625f-f911-47dc-913c-9b7283007d79 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:09.087260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:a02d9e96bccddf29523d4445f64d31ef3fe2a8d9b937b9b82abcd705a7c7839c

Observation 27f16019-f5ab-482c-a3c3-229879c8f308 · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.409003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:c923dea14e9ea90196018769fb9de3a2896e826373b6487bab62d919b4633172

Observation 8904c37e-bae8-43f4-b1f4-e6b1b9576a75 · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.827682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:0bf10c3ab86232dd3b95268fb90bfa12ca3a3faca6a419fda0449c51a42d3da2

Observation ac7f4f7a-08a9-4ae7-9c10-ac45239fb51b · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.092691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:ffa031b1b72500a722b0c77131e014b73576900acb106913046bb7d867d23e47

Observation cffefd4d-5467-42c3-b666-e5c677216762 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:05.243802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:75f492ad22995e3568f17f6ad42eca36b7be66d223fbbaaf5f4ce1b0a9a7adcf

Observation 04ef7fec-55a2-469b-b11e-13f85b769347 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:59.403633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:5a04f80b4601a30a08dab97a4ff02bc0b33804b2aa64bbb45ba0ba6c11a0568f

Observation 6d45c1f7-88b7-4b72-9a5e-75ef342eb4fc · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:28.062118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:bb74c351d0b3a5c0433281056b728c9f8b16cf83efc9c9dea94cb51dfa48528e

Observation f3cd6cc3-8198-42f8-b0c1-5d701390201b · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.199683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:a4167a207bbbf21c5609852c152e6d216486648b578d625fb332970dceed45c4

Observation fd58e815-33b8-4c2b-a799-7d7328ae7e9f · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:41.063061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:84a5b6ff7f4776de05454a054c4fdacb730877a1e9131eb4da8ab9aea0a7e7aa

Observation 0aee1525-3fe3-446e-a08f-0a4820de4161 · inbound

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate cites this paper.

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.804526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T22:54:53.415067Z digest=sha256:e8e6bfe1a83d02310e18d8f8ebed03ec7fe24b0f004db0fe98d6a7ada0c53f2f

Observation c20fbf8a-6285-40a2-b777-36c0e8059a36 · inbound

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection cites this paper.

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:22.844157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:20:06.381334Z digest=sha256:91427371f1aa212c9cce74185ae6a70343432795dedc7401ba24f269e7606d2c

Observation 247f0921-085c-40d3-bb1a-5832dc6274cc · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.052519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:2d869a3d7e642d4a0853a1ee97a09df09a83d99b9676ef411efcd07a8b603956

Observation 92d6a179-4695-4e99-92a4-c596e43b8601 · inbound

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents cites this paper.

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.521293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:14f1b6b237f9a6a1a12d60bc2908bb223286793af7a746d12cb0460381c6f670

Observation 842d94d0-100c-48b9-8371-829985f2993b · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.067368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:77a067c37ca188c5c8341236deb24a2b86725b847cc8717814579a34abb3d804

Observation 08ff4005-a93c-4e45-97dc-15cb1ad52b5a · inbound

DRIFT: Refining Instruction Data via On-Policy Data Attribution cites this paper.

DRIFT: Refining Instruction Data via On-Policy Data Attribution A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:18:54.722039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T01:57:05.784589Z digest=sha256:d50b9e8a3556fa7d84cb4e425d5df96ccccc4495a962be578a25b2bada2ff021

Observation 7d607ce7-cda3-4fb0-9d51-f107ba4b1af5 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.575477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:0a362eb4e0e6bf269d5453fb933b4999c3f2dedf6ce908d95e79a3ba1d3c075d

Observation 6299d2f4-a2e8-4d50-99c9-9b112124ab68 · inbound

RL Post-Training Builds Compositional Reasoning Strategies cites this paper.

RL Post-Training Builds Compositional Reasoning Strategies A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.457832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:42985a89728e2eaf6d9ed77ee9f5c88240b5e0d71e2071e842aa2043927c6519

Observation 9d906ae0-a2f9-433e-821b-42cf4472e9ba · inbound

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity cites this paper.

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T04:32:37.522325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:32:37.522325Z digest=sha256:fa3b3be118b2ca3f5505c3bcb5a66cf56ae0c0f145c3ef53d8a2ec35a2b39e83

Observation 8b0b7d76-777d-494d-a718-80aaea0fcccf · inbound

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches cites this paper.

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T14:45:23.980380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:45:23.980380Z digest=sha256:28dc067b4cac91cefa70a772dc7025516d49bd41518350287bca8b2ba7773520

Observation f9475cb6-7154-4997-af99-64cca06655f9 · inbound

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning cites this paper.

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T13:54:23.273919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:54:23.273919Z digest=sha256:203e8f629a7c8548d0fba34dc5603ef59d7f9fc4752955f0ebf55a84430a99af