Pith. sign in

Paper Citation Record · LEDGER

Parameter Exploration for RLVR via Variational Learning

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2608.09805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09805 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:32:13.146209Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact10
  • verified fuzzy28
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 463ddaed-3d44-4da6-9ce5-80f49d67d6bd · outbound

This paper cites A Survey of Exploration Methods in Reinforcement Learning.

Parameter Exploration for RLVR via Variational Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.730067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.730067Z digest=sha256:250128d231f1dc9e15c4e3c941b412ef7283970b4a7499ff0e414cc4596ea23d

Observation 37c920ea-313f-4c46-8e5e-e92705d6dc25 · outbound

This paper cites Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025.

Parameter Exploration for RLVR via Variational Learning Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.735104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.735104Z digest=sha256:532f185412b1b1fba1853872d3e5ef68cdb63e6ecd0b25650fc69f434c4032d3

Observation 9f15f11d-376d-433e-a2d4-c7a09b942c1f · outbound

This paper cites Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards.CoRR, abs/2602.02555, 2026.

Parameter Exploration for RLVR via Variational Learning Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards.CoRR, abs/2602.02555, 2026

Reference 3

Resolution
verified exact
doi, observed 2026-08-11T10:32:14.059743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.739164Z digest=sha256:dd8454aa4da68141f31fe5b2a1517b226dd244bcfee0b5c446212557ce928b0e

Observation ba81f022-c73f-4977-a899-1156cd2224c0 · outbound

This paper cites Llama-nemotron: Efficient reasoning models, 2025.

Parameter Exploration for RLVR via Variational Learning Llama-nemotron: Efficient reasoning models, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.743020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.743020Z digest=sha256:462e5c304643aa3a38c12105170a9e7a6c545faa0371863198022892f710d7c3

Observation ae42e433-66d1-4662-a69d-e26076b8156a · outbound

This paper cites Weight un- certainty in neural network.

Parameter Exploration for RLVR via Variational Learning Weight un- certainty in neural network

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.747358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.747358Z digest=sha256:741f200cded2e4160e71820012b71e43def93fbbcf647aa107223328e7a382d7

Observation 646564d1-5bb4-46a7-ac84-198d06b87cc4 · outbound

This paper cites FullStack Bench: Evaluating LLMs as Full Stack Coders.

Parameter Exploration for RLVR via Variational Learning FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.751253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.751253Z digest=sha256:0f7ee004ca025d66f61e2b13a83e2e9b9eac38542d8f0b6cd8238d3d49ca77cf

Observation 06241385-ab5d-4963-b8db-254df635a039 · outbound

This paper cites Skyrl-v0: Train real-world long-horizon agents via reinforcement learning, 2025.

Parameter Exploration for RLVR via Variational Learning Skyrl-v0: Train real-world long-horizon agents via reinforcement learning, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.114998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.755926Z digest=sha256:cbe68455c652848257e8bb24293bbe2c0df8c73426c7be2ff7847ec054e2f8d2

Observation 21a74956-aa23-430f-82c6-254cdabc43f2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Parameter Exploration for RLVR via Variational Learning Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.759973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.759973Z digest=sha256:68af2be505bedf1591f7292f13969dab82c5f80b0c272bad576f67c06d6cf486

Observation 51591005-14da-446a-84f1-1dfa858b03f0 · outbound

This paper cites Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward.

Parameter Exploration for RLVR via Variational Learning Exploration vs exploitation: Rethinking RLVR through clipping, entropy, and spurious reward

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.102504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.763941Z digest=sha256:05f956f31a73951b8ff568c0877e30ef96ac99148ebc1c871d775de81c9c0c91

Observation 07287b1d-f253-4169-b518-6e5282df3d20 · outbound

This paper cites Improving LoRA with Variational Learning.

Parameter Exploration for RLVR via Variational Learning Improving LoRA with Variational Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.768570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.768570Z digest=sha256:a3c242a4b71797878cdd4c7a52baced3b1aa19e259cdddffebec532d4d4948cd

Observation 107d603d-cb86-4f0c-8ce4-f20e8e33bb98 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Parameter Exploration for RLVR via Variational Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.774265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.774265Z digest=sha256:d6c1b6cec618455b400163e1e54fb40f1233f644d1ddb9dc21dd71141fed4104

Observation b13e8654-d971-43e9-9564-14865009e741 · outbound

This paper cites Uncertainty-aware decoding with minimum bayes risk.

Parameter Exploration for RLVR via Variational Learning Uncertainty-aware decoding with minimum bayes risk

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.090956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.778746Z digest=sha256:3de90932dab4ccc4b920cecad5463d69373fbb4e390bd1a85868b599955224b8

Observation 8c5f775f-51d5-463a-9099-8ed57d78607f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Parameter Exploration for RLVR via Variational Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.783229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.783229Z digest=sha256:cf4f64b216c72df5e7f7ee5fa6b8b6e48be2651aaad629996a62dc8cd75f9da9

Observation 123cc001-309e-4151-aecf-ccfc5a20a76b · outbound

This paper cites A survey on policy search for robotics.Found.

Parameter Exploration for RLVR via Variational Learning A survey on policy search for robotics.Found

Reference 14

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.954304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.786642Z digest=sha256:8c9a65622b9beee5114d6282e150710088ff8b8cb9e69b37c30047461ffb0883

Observation 66d0116e-5f2b-4b80-975a-7f0c8792c0b5 · outbound

This paper cites Sharpness-aware min- imization for efficiently improving generalization.

Parameter Exploration for RLVR via Variational Learning Sharpness-aware min- imization for efficiently improving generalization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.076200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.791126Z digest=sha256:df43f9fb4319ff1a75f15c43d87d07884104a36ad4813d5cbb30b5894f13e6e5

Observation 0d56f047-1a95-4fb1-a99c-57e9f2decc1e · outbound

This paper cites Noisy networks for exploration.

Parameter Exploration for RLVR via Variational Learning Noisy networks for exploration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.063344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.795514Z digest=sha256:d6562b59d29e32c4d9dca797cfb3ca76c0d1d41604314ad960cfee9592bdd76a

Observation e18a1ea0-85d6-48b1-b172-5c81df055be6 · outbound

This paper cites Neural thickets: Diverse task experts are dense around pretrained weights.CoRR, abs/2603.12228, 2026.

Parameter Exploration for RLVR via Variational Learning Neural thickets: Diverse task experts are dense around pretrained weights.CoRR, abs/2603.12228, 2026

Reference 17

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.938318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.805368Z digest=sha256:42b7f25965d8fea3b88418c22541298a170f5e8bbdd1989d5698bac253491abc

Observation cc62d422-b240-4614-a421-cd1809e4a508 · outbound

This paper cites Practical variational inference for neural networks.

Parameter Exploration for RLVR via Variational Learning Practical variational inference for neural networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:15.039427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.810372Z digest=sha256:2485e0daadd2042c0144ae957e07ae0776e05809cb68da03e222649f14221348

Observation be16919a-7051-4fa5-b122-e5686d58ba26 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Parameter Exploration for RLVR via Variational Learning Skywork Open Reasoner 1 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.818963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.818963Z digest=sha256:8d28a84e29ddc9cc909bef748a14d84847c6160a2bbdec6aaf02cb9a99b8d9bf

Observation 39dcb276-4953-4707-9f8b-bd6e0c39851a · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Parameter Exploration for RLVR via Variational Learning Measuring mathematical problem solving with the MATH dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.823287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.823287Z digest=sha256:62c5c33d1ccd89be0922b6bcf07ca762251f774255de040dd2eab9c43a21a91b

Observation bbcfc14a-08ab-4b2f-8bbf-1ae24b45ff93 · outbound

This paper cites The curious case of neural text degeneration.

Parameter Exploration for RLVR via Variational Learning The curious case of neural text degeneration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.827658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.827658Z digest=sha256:a5b2f6550f9004f2b893fb971191548cddb21dc77c5962c18f7b80bcdadf1844

Observation aa45cfff-3eec-4ef9-b7ea-ce9b847dbacf · outbound

This paper cites Brorl: Scaling reinforcement learning via broadened exploration.CoRR, abs/2510.01180, 2025.

Parameter Exploration for RLVR via Variational Learning Brorl: Scaling reinforcement learning via broadened exploration.CoRR, abs/2510.01180, 2025

Reference 22

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.827218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.832290Z digest=sha256:ba22bda5da9fcd36408498df1f9fb787429fc0883f0720648c7b5c05a541fc48

Observation 45a2747d-aeee-4978-8f5f-d1e3705c635c · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Parameter Exploration for RLVR via Variational Learning Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.991288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.837036Z digest=sha256:28cc3ef53c0c2dce35cb17e41feb45e878726951a86078abeabfc5fe83163b99

Observation b29344e8-1bd0-4d1f-81af-47409f0aa669 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.J.

Parameter Exploration for RLVR via Variational Learning Near-optimal regret bounds for reinforcement learning.J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.842280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.842280Z digest=sha256:82814f98bf1248d14a5548b19d5c8f5b7ba8a3f7e6a71a4245635b562bc01438

Observation b589645f-91ae-44e4-a097-1f32185282a3 · outbound

This paper cites Rethinking entropy regularization in large reasoning models.CoRR, abs/2509.25133, 2025.

Parameter Exploration for RLVR via Variational Learning Rethinking entropy regularization in large reasoning models.CoRR, abs/2509.25133, 2025

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.757133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.846240Z digest=sha256:edb7d10d686f5aee2b85a57753c7071b5884b74bf83e0903cf2799abd53e496e

Observation 4fa8d917-90a4-44d0-8db7-8b6913d2551e · outbound

This paper cites The bayesian learning rule.Journal of Machine Learning Research, 24(281):1–46, 2023.

Parameter Exploration for RLVR via Variational Learning The bayesian learning rule.Journal of Machine Learning Research, 24(281):1–46, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.979814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.850986Z digest=sha256:a22525de75eccc0e3ec85a8a80c078374abd2f94bbcddcf5d67e170c8b845288

Observation f3337831-493b-48d4-8920-4c952f4b3dd9 · outbound

This paper cites Fast and scalable bayesian deep learning by weight-perturbation in adam.

Parameter Exploration for RLVR via Variational Learning Fast and scalable bayesian deep learning by weight-perturbation in adam

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.967885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.854665Z digest=sha256:fa80262f9e10fe3f38e0482edfba77c54711f70617cfd15769683d57e3e4cace

Observation 7af71b16-0a03-4670-a9b4-38097d635421 · outbound

This paper cites Generalized variational inference: Three arguments for deriving new posteriors.

Parameter Exploration for RLVR via Variational Learning Generalized variational inference: Three arguments for deriving new posteriors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.953142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.858567Z digest=sha256:ee12aa4c88e72d4b51cf3acb71dd1cb5efb480718fb41e3088b23cc4f0b0cc76

Observation e15d98a9-9f10-44d7-9cb6-9be57a09ff38 · outbound

This paper cites The Case for Learning Application Behavior to Improve Hardware Energy Efficiency.

Parameter Exploration for RLVR via Variational Learning The Case for Learning Application Behavior to Improve Hardware Energy Efficiency

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T10:32:14.357333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.863076Z digest=sha256:1e4b5cf5a544b785b8867c3e107379fc4a2ced2e4d13ca12f649bf97e94c8b00

Observation e0fefb07-c04c-48d0-aa5b-12d3b47a11b5 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Parameter Exploration for RLVR via Variational Learning Efficient memory management for large language model serving with pagedattention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.868294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.868294Z digest=sha256:a285ea06350eda6b61193c15797a766aaa9bff9fe671e6646a0dba5e42a6a90b

Observation 4f434f82-1798-4f91-af72-e455398bb5bd · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Parameter Exploration for RLVR via Variational Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.885592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.885592Z digest=sha256:4ed7217366900b4a554ea37ba7c1320f98a1d2136733ab2133f0437bd837730e

Observation 32432725-071a-4e0f-afec-715aa718c91a · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo et al.

Parameter Exploration for RLVR via Variational Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.941961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.890555Z digest=sha256:9ed968980535545b1fa4a3ac52c419e9225ac9ad7200c4942c77e2b9d37467d2

Observation 0563aac8-04f5-4dc0-80f6-170c4957c00d · outbound

This paper cites Verified taco problems.

Parameter Exploration for RLVR via Variational Learning Verified taco problems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.929410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.896238Z digest=sha256:ad8ee78eaaccc3a7c53922205a249430f8beff763358a63c84242c586a35fa8c

Observation 09e3f26f-11c0-46cb-8543-b0b7b43bf160 · outbound

This paper cites Handling the positive-definite con- straint in the bayesian learning rule.

Parameter Exploration for RLVR via Variational Learning Handling the positive-definite con- straint in the bayesian learning rule

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.916744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.901478Z digest=sha256:e52434e55b077268d5b2c97fd74aa04112a118e48cf42300449c4c3a5b3ee4c4

Observation bdc9df48-1918-4067-baef-34908d197424 · outbound

This paper cites When speed kills stability: Demystifying RL collapse from the training-inference mismatch.

Parameter Exploration for RLVR via Variational Learning When speed kills stability: Demystifying RL collapse from the training-inference mismatch

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.903579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.907014Z digest=sha256:4578c40332997420e6d7a9e88b3a7fc865331cc949d679edb725566215378fbe

Observation 48cddb81-4e20-40a9-b123-3e65c7504c61 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

Parameter Exploration for RLVR via Variational Learning Code-r1: Reproducing r1 for code with reliable rewards

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.911912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.911912Z digest=sha256:e35b0e1f20e650e0ea5eaa008c75adf22a5a6f263954e92b51ef3a74dc736b10

Observation 469e5597-9969-4364-bdd8-aefef327d7dd · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Parameter Exploration for RLVR via Variational Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.916702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.916702Z digest=sha256:005d2ed6535083881e34666152b16f73372872cc873151e3b699d6840e286340

Observation c46c7011-3bd6-4da4-b5ac-95641628bd36 · outbound

This paper cites Regularization matters in policy optimization - an empirical study on continuous control.

Parameter Exploration for RLVR via Variational Learning Regularization matters in policy optimization - an empirical study on continuous control

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.881665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.920919Z digest=sha256:4e9f3fa451398f3e2f5c18f3b7918e91682bc5fb1d8ab498254c0fbe5407ee47

Observation 7d98b97b-749e-4c50-9625-8c85bb0645d3 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

Parameter Exploration for RLVR via Variational Learning Understanding r1-zero-like training: A critical perspective

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.929977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.929977Z digest=sha256:d0de94bac3bcb67692c8386587406cfa5d1e61b93e9032bdc000db45be7a132d

Observation d1efa133-d7be-4794-bd75-3f0a4a52fb6a · outbound

This paper cites Decoupled weight decay regularization.

Parameter Exploration for RLVR via Variational Learning Decoupled weight decay regularization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.935853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.935853Z digest=sha256:99004fadbc4c897e06a44f090b96a00da498038b22f07df34769d06af94fc780

Observation 7e5db798-f81c-4539-9cb9-43dfa506115c · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa et al.

Parameter Exploration for RLVR via Variational Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa et al

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.834331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.940625Z digest=sha256:192d1f482813d7220bd26a7e2f51d7939b236b781cb2ff088b9d824b1b5eeeb4

Observation 8eba8178-36aa-4337-9dff-9572eaa26618 · outbound

This paper cites American Invitational Mathematics Examination, 2026.

Parameter Exploration for RLVR via Variational Learning American Invitational Mathematics Examination, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.945458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.945458Z digest=sha256:7d580bcd29075d551be2337c7e05c8a648bf5463170a5077ef81362b9bb2846c

Observation fb3d2388-c503-41d7-8992-9b9b8033965a · outbound

This paper cites SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks.

Parameter Exploration for RLVR via Variational Learning SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:32:14.270721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.952545Z digest=sha256:a58148dd74acf34c634acd6b3c1908f23d31640d0783e0386b5884576b182d2c

Observation fdfe7acf-c09e-43e7-9b63-70e8b6eabbc9 · outbound

This paper cites SAM as an optimal relaxation of bayes.

Parameter Exploration for RLVR via Variational Learning SAM as an optimal relaxation of bayes

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.812635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.957266Z digest=sha256:01508c633de77cde3ddec77730092ff611754a0e05e8f4549c34a6744b5bd5cb

Observation e36fba69-5a95-4a91-ac65-ca1b408713c6 · outbound

This paper cites Faster, more efficient RLHF through off-policy asynchronous learning.

Parameter Exploration for RLVR via Variational Learning Faster, more efficient RLHF through off-policy asynchronous learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.800601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.961285Z digest=sha256:19b4be0ea9232b3bfcd9df65f99f7e07f322bcd86dfbcec3871e8338f67ad4ba

Observation 31f11779-fdc0-4d42-8776-bacf1706d5bc · outbound

This paper cites Olmo 3.

Parameter Exploration for RLVR via Variational Learning Olmo 3

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.966627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.966627Z digest=sha256:7afb576643d71a06f0702f22dbb9bf315dbb0315663e2acaae4f6f2f6b03e86f

Observation 73ad86e3-942c-475a-800f-ef103c20bef4 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray et al.

Parameter Exploration for RLVR via Variational Learning Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.787028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.973226Z digest=sha256:6cc77cfd53cdfb8bfa80c0d3e3acae12ffa982f025b75c6b870d1e55d9b9345a

Observation cef92906-a67a-47bc-82b4-1a8377b882dc · outbound

This paper cites Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz.

Parameter Exploration for RLVR via Variational Learning Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.776599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.978337Z digest=sha256:56a13668c6924aca74494cebde610e8c404380bcb69b7a3af26975886166bde2

Observation f75a85d1-56bd-4d9c-9930-7d709691dc09 · outbound

This paper cites Exploring parameter space in reinforcement learning.Paladyn J.

Parameter Exploration for RLVR via Variational Learning Exploring parameter space in reinforcement learning.Paladyn J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.763936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.983370Z digest=sha256:9d6f2b4a2bb4d63286ef44d2afd07ca7dca9e3d820349b884c0e239135a97601

Observation cc300bdb-8913-4a54-b274-81d3787e8a95 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Parameter Exploration for RLVR via Variational Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.001257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.001257Z digest=sha256:f6baef24bf3deb97f8152d3b736f1f49c57bb2ec211cbded57a7f71c89dad0f8

Observation 2c1d7bd0-2199-4cbf-b47a-887012c3dd81 · outbound

This paper cites Parameter-exploring policy gradients.Neural Networks, 23(4):551–559, 2010.

Parameter Exploration for RLVR via Variational Learning Parameter-exploring policy gradients.Neural Networks, 23(4):551–559, 2010

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.006112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.006112Z digest=sha256:692902bbd24bf4b4122f1a561a6ef13bd08793e8343bc90adac9dc0e187629e7

Observation a38eab6a-6954-4594-88ba-7018c0acc17e · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Parameter Exploration for RLVR via Variational Learning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.013301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.013301Z digest=sha256:02051b823d1eab412bd8ddf06fc60eb28623eb1ee99466f149d759af6e5086bc

Observation 3c22fee3-d2a8-4aba-a5a5-480d4c63475d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Parameter Exploration for RLVR via Variational Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.019444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.019444Z digest=sha256:183ecdb27f94665dcab1916b7c4e38ff29679608a59f5fea22e2e17efd36950b

Observation cdc4d653-dc12-4ea8-b959-48296873be97 · outbound

This paper cites On entropy control in LLM-RL algorithms.CoRR, abs/2509.03493, 2025.

Parameter Exploration for RLVR via Variational Learning On entropy control in LLM-RL algorithms.CoRR, abs/2509.03493, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.024389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.024389Z digest=sha256:4b01dfca1ba022d3955195de497347b8540e15673c416b553ede614c7de2f887

Observation fc207249-b380-469b-8164-1bfd76a837c6 · outbound

This paper cites Variational learning is effective for large deep networks.

Parameter Exploration for RLVR via Variational Learning Variational learning is effective for large deep networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.752260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.030530Z digest=sha256:f080939c2cc5b01da4d2eef8ffe0ffcac9bc41501bc114055b09330bf7a7a96a

Observation 5c5e5b2e-d02f-4f6e-a25a-efb5754ff767 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Parameter Exploration for RLVR via Variational Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.039985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.039985Z digest=sha256:6a8b0af11c6471f90dcb50c1b63ffd25fe9c68fb02b3c17c52197d0ba7f9ce9d

Observation 01429360-ac01-4efa-876a-3e51d1e403cb · outbound

This paper cites Strehl and Michael L.

Parameter Exploration for RLVR via Variational Learning Strehl and Michael L

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.044273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.044273Z digest=sha256:3e123492eec2d9729d7430a95f189cd7384969d26cb6f570f3ec6e9c99d369bf

Observation 0ed17521-795b-4680-9582-a16643fe2b55 · outbound

This paper cites Path integral policy improvement with covariance matrix adaptation.

Parameter Exploration for RLVR via Variational Learning Path integral policy improvement with covariance matrix adaptation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.727527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.048926Z digest=sha256:4a9e788f0decb5723a04fcca5efb184bbf36d7f18bc8145af7cf2b6b5e031ed9

Observation 2603b6f4-9c35-4596-81a3-4981b964bb16 · outbound

This paper cites RL grokking recipe: How does RL unlock and transfer new algorithms in LLMs? InThe Fourteenth International Conference on Learning Representations, 2026.

Parameter Exploration for RLVR via Variational Learning RL grokking recipe: How does RL unlock and transfer new algorithms in LLMs? InThe Fourteenth International Conference on Learning Representations, 2026

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.714895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.053355Z digest=sha256:67db7064276ee892571cd0f0b855247302dc33b3e398980611fd284b7937bbd5

Observation 2f6d0311-0935-4286-a635-9965396094f6 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:14.703322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.059829Z digest=sha256:a0d0f8ef2481d4427327d145e19a16cf03f58c41e42caf28a372f399966a0c9e

Observation ae7c3df1-a809-46b8-bf67-5dce78e32c75 · outbound

This paper cites Sutton and Andrew G.

Parameter Exploration for RLVR via Variational Learning Sutton and Andrew G

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.691421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.066364Z digest=sha256:8cc72951c3ed1a0fe1433ea634df7d501523212536895efcb18f7f55f781f333

Observation 907023fd-33b9-4a78-abcb-d1b231d7cdbc · outbound

This paper cites Theodorou, Jonas Buchli, and Stefan Schaal.

Parameter Exploration for RLVR via Variational Learning Theodorou, Jonas Buchli, and Stefan Schaal

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.071427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.071427Z digest=sha256:67b038e3ecf499156ff9f7c5ce6f778015756236eb9fd759909cdaea159b86c1

Observation 90fa11a7-9efa-4a55-b408-9ac2e528f2aa · outbound

This paper cites Generalized exploration in policy search.

Parameter Exploration for RLVR via Variational Learning Generalized exploration in policy search

Reference 63

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.571135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.076147Z digest=sha256:acd91b9ea7baf018070efceea1c28019473302e4a5b338279cd67f2b4cd34f7d

Observation b3f7816e-f0b6-4c34-bb47-255a26c11571 · outbound

This paper cites Aletheia: What Makes RLVR For Code Verifiers Tick?.

Parameter Exploration for RLVR via Variational Learning Aletheia: What Makes RLVR For Code Verifiers Tick?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.081086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.081086Z digest=sha256:5ab6c08d72f77b2586e5d7c2c3a295c32092edcb7131c0808d521905b9a5f184

Observation 6a678b06-4ac0-45a2-bfd0-706db61a7da8 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Parameter Exploration for RLVR via Variational Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.087281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.087281Z digest=sha256:3efa2e5a318cbce181348587f38cf02137b7bb0e94ac4fee8d69e498813391ed

Observation b1068912-5ebc-47e3-80c3-77ecc4ab946b · outbound

This paper cites Learning from delayed rewards.

Parameter Exploration for RLVR via Variational Learning Learning from delayed rewards

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.092467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.092467Z digest=sha256:907ed5b7f4b16b3e55d2c0a196e653f87cd770b341a99ee5408f10dad76d2d2b

Observation 87d8fd76-b25b-4726-88e9-bfa2efb01d1c · outbound

This paper cites Adversarial weight perturbation helps robust generalization.

Parameter Exploration for RLVR via Variational Learning Adversarial weight perturbation helps robust generalization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.672871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.097808Z digest=sha256:4f89b2b3106796eb84260843c40b1290ed87751c9a14f0b65c64910c87f7b6da

Observation 15766fb5-f584-47d9-9500-f27ae32303ae · outbound

This paper cites The invisible leash: Why RLVR may not escape its origin.CoRR, abs/2507.14843, 2025.

Parameter Exploration for RLVR via Variational Learning The invisible leash: Why RLVR may not escape its origin.CoRR, abs/2507.14843, 2025

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.102018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.102018Z digest=sha256:06c9739acb4f0409374fb0c08eccea5c7950d069721463ac9f25c0d8dd531a96

Observation a2c6d4d7-5060-4d36-af4c-9dcc5e70c0d1 · outbound

This paper cites Reasoning or memorization? unreliable results of reinforcement learning due to data contamination.

Parameter Exploration for RLVR via Variational Learning Reasoning or memorization? unreliable results of reinforcement learning due to data contamination

Reference 69

Resolution
verified exact
doi, observed 2026-08-11T10:32:14.660615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.108060Z digest=sha256:a46b7090c86db3894e1e8da00fc76232061060c2e8b6ed961d7a2783b011e012

Observation 51e4aa03-2498-412f-89c4-f1a4775f3258 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Parameter Exploration for RLVR via Variational Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.111958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.111958Z digest=sha256:82cb75139870ec00ace8878733229c61a0aa69a63713002e3aa064664865708a

Observation c678ef5a-4c70-40ef-8213-ed6ce0076ec3 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, August 2025.

Parameter Exploration for RLVR via Variational Learning Your efficient rl framework secretly brings you off-policy rl training, August 2025

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.116229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.116229Z digest=sha256:9a327806e15f0a23dfa122a1bf321ab6d221d50a7c4bc828041e3e4fbd528c14

Observation 19cd3682-c4e9-46fa-8e09-ac3dbf307f84 · outbound

This paper cites The debate on RLVR reasoning capability boundary: Shrinkage, expansion, or both? A two-stage dynamic view.CoRR, abs/2510.04028, 2025.

Parameter Exploration for RLVR via Variational Learning The debate on RLVR reasoning capability boundary: Shrinkage, expansion, or both? A two-stage dynamic view.CoRR, abs/2510.04028, 2025

Reference 72

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.447823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.120490Z digest=sha256:1a1d2ae7fb327eeab8ab50c28b04985cfd487c95dd242dd4504ae6ca542364fd

Observation 7fc9da07-b433-44a7-943c-b608181a1e43 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Parameter Exploration for RLVR via Variational Learning DAPO: An open-source LLM reinforcement learning system at scale

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.638208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.124865Z digest=sha256:2501d0fa8dd2bd709a32d80f0689f9863875599e2cfd97809d032ee3846c9bce

Observation 91df57ff-200d-4d8b-b1c3-ad18867ce443 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Parameter Exploration for RLVR via Variational Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.129972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.129972Z digest=sha256:c62ee4ac2a44676765d60d029dc72342df71993423184335702c563d946a0887

Observation 93db9b12-6f5e-4340-865a-1dbfa34e2bfa · outbound

This paper cites Group Sequence Policy Optimization.

Parameter Exploration for RLVR via Variational Learning Group Sequence Policy Optimization

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.137555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.137555Z digest=sha256:020f390115866af1e929991ff94f34262746c3978c3ef0b074aa8d655211ccd1

Observation 7d653e8f-31f8-4780-9f55-d034ef0ebb78 · outbound

This paper cites The surprising effectiveness of negative reinforcement in LLM reasoning.

Parameter Exploration for RLVR via Variational Learning The surprising effectiveness of negative reinforcement in LLM reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:32:14.622844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.142111Z digest=sha256:6c76f0272ce94429927bd1b8777d74c581d639183647932ea8e663704440c406

Observation 053302a1-ebd4-4250-ac2e-12f4ab659af2 · outbound

This paper cites Exploring multi-temperature strategies for token- and rollout-level control in RLVR.

Parameter Exploration for RLVR via Variational Learning Exploring multi-temperature strategies for token- and rollout-level control in RLVR

Reference 77

Resolution
verified exact
doi, observed 2026-08-11T10:32:13.355286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.146209Z digest=sha256:12f95a56ada03c95c9f92506d6743f154d7956943c85dbbf1226eacff19e6174

Observation 0cbc3438-9d41-410a-8c07-9614d43308a1 · outbound

This paper cites URL https://doi.org/10.2478/s13230-010- 0002-4.

Parameter Exploration for RLVR via Variational Learning URL https://doi.org/10.2478/s13230-010- 0002-4

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.988600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.988600Z digest=sha256:8632712eadfeb76303e0759381aec68f8817fa58f618d8a831c18381b9ea33e9

Observation f8a980b0-da8d-4898-8282-325438106671 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2011

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:15.027341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.814815Z digest=sha256:248d5a13885c4979411965347c731b0c44d928ff4c7ce0a1d330345cb4f69383

Observation f282f0ce-7eb4-47dd-8d52-efd6e25bef35 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:15.052051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.801070Z digest=sha256:f1315e1ba7014398383254436f16cd261c245746d79a7de702d825da16fcba5f

Observation 937248b9-6481-4813-a1fc-2729e8ad8412 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:14.868197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:12.924363Z digest=sha256:62a54be3de9fadd5e4473faf5f61cff363b21c143e2f66f1d70c3dc41a817485

Observation a210d05b-a4f2-4727-9d7d-3edfdf420c50 · outbound

This paper cites an unresolved cited work.

Parameter Exploration for RLVR via Variational Learning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:32:14.739485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T10:32:13.035731Z digest=sha256:90ba5895b882443a8fb677682a242a778c5078c95e988ce72200b5546d32d0cf

Pith citing papers

No inbound Pith citation observations are available.