Pith. sign in

Paper Citation Record · LEDGER

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation

As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.09271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09271 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:29:03.988127Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48b53655-8338-455c-9072-cc43d51e5d70 · outbound

This paper cites Maximum a posteriori policy optimisation.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Maximum a posteriori policy optimisation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.154815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.752867Z digest=sha256:4ad8bee92621510029b80bca91fc385f7c0a4120797ec6d6a1d9c262fd6a5857

Observation b55a5c17-6a6d-44f5-821f-3c6cfdbe5ed4 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation On-policy distillation of language models: Learning from self-generated mistakes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.758603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.758603Z digest=sha256:a3e97bd6448b6cdb8fb339b12ac60b666bd8be6ed0683618ef568e9332ffaff4

Observation 525ca485-55f0-49a7-b7e6-8fad8d4166b5 · outbound

This paper cites Escaping the Verifier: Learning to Reason via Demonstrations.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Escaping the Verifier: Learning to Reason via Demonstrations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.762979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.762979Z digest=sha256:66a1f48e232fa16992385fe2fad0b17b8a27f91e878455ea541f4be5c0201a82

Observation 9c081b2c-b1cc-43ae-826a-154f1b865aa2 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.767752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.767752Z digest=sha256:7ecefdfae5883b664239f15ad3ed0b5d2b144931e2652e905b0afd5a6e1100d7

Observation e21b2de8-8e72-4244-a467-cbf86dfbdcbb · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.773626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.773626Z digest=sha256:33e01a5d06d55704c192735a1c36889bc1725031f8342ed81c8d1bcb71149e7a

Observation e92ab6a6-f12e-4c5c-bf29-15f1bc327298 · outbound

This paper cites What is the objective of reasoning with reinforcement learning?arXiv preprint arXiv:2510.13651, 2025.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation What is the objective of reasoning with reinforcement learning?arXiv preprint arXiv:2510.13651, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.779243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.779243Z digest=sha256:acbb8a2c3caf9f7b023d2c1262997cf2ae339289d444fac90adaea57ae1775f6

Observation f13f1dd0-43c6-4c6d-9695-30ca959afb0a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Imagenet: A large-scale hierarchical image database

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.783941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.783941Z digest=sha256:a9b23d2e5fc2fe801d66fa39f607a7cc0711cb69c57911985b4ef166b226b9c8

Observation 84f6e10f-4685-4e7f-8cc3-fb2d32f8e0f2 · outbound

This paper cites Openvlthinker: Complex vision-language reasoning via iterative sft-rl cycles.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Openvlthinker: Complex vision-language reasoning via iterative sft-rl cycles

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.111003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.790434Z digest=sha256:2dc11eaf5cb280504f163dbeef24b30922d56c42011c033ca69bc3b8a24d695d

Observation 20687639-1a63-447a-a10a-f63fbced52b2 · outbound

This paper cites Cold-startreinforcementlearningwithsoftmaxpolicygradient.AdvancesinNeuralInformation Processing Systems, 30, 2017.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Cold-startreinforcementlearningwithsoftmaxpolicygradient.AdvancesinNeuralInformation Processing Systems, 30, 2017

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.094454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.794665Z digest=sha256:60c6ae30e83b402bd990fa5c41cfef51ab59c85636c646de9c94e03ac82b84ce

Observation 26e738b6-6dd2-424a-8a0f-45140ae1f74a · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.802470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.802470Z digest=sha256:8dfac744e4ffb915da8a096af69e2ac6e188bc612296c8766d5b2b3b20322e11

Observation 7beda436-9439-4ca8-a003-b6981708edcb · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation OpenThoughts: Data Recipes for Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.807561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.807561Z digest=sha256:e4bd268c53bb4cf0398ade0f5afe944c4f9f8e6bfa78a691d9fb250c4ac1b9b7

Observation 4a17ffd8-7427-4f11-9cf7-79adc4121cf3 · outbound

This paper cites Learning to reason for long-form story generation.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Learning to reason for long-form story generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.075028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.813234Z digest=sha256:a872d636542322ad1843d7ae3e9f0272dcc633e3c597ebbc38c6cf5c83b798c6

Observation 48f9ec6e-9262-4274-afaa-6fb7501c6178 · outbound

This paper cites RLP: Reinforcement as a pretraining objective.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation RLP: Reinforcement as a pretraining objective

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.054235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.817892Z digest=sha256:fa2798dcf77340690206110f2a25f9c53ee2e7a155b8b331927c02781bde2f0a

Observation 05494d81-0622-43c3-b6f8-6e8f374f2ce4 · outbound

This paper cites Deep residual learning for image recognition.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Deep residual learning for image recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.823062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.823062Z digest=sha256:61a98f0db3c2f8887267c8fccc18154003db17e52890d5c175a42e2b62dea12e

Observation 6cfee7c8-6f4c-48d2-9d30-2b7b50364774 · outbound

This paper cites Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, and Vicente Ordonez.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, and Vicente Ordonez

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:29:04.545909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.827802Z digest=sha256:3558867cd2bc00852fb068e338a9f3cf5a4e73cc1ed11ae5ff79116ff576de60

Observation 35b152db-d9c5-4e67-be3b-5f387a809020 · outbound

This paper cites Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.022024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.832205Z digest=sha256:ea916284b4b254ec1beb3832deda9283f9df49e0538c2868fede79f7dbf84451

Observation 9b185181-cef0-4488-85ec-28312844286e · outbound

This paper cites Measuring massive multitask language understanding.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Measuring massive multitask language understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:05.003190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.837328Z digest=sha256:55f69ab358a20a220aaa34dec6116d17b4ec7b9be3ac5650e346b5be8c74902f

Observation 15b122f3-3271-47d1-a85e-1056e4c2a181 · outbound

This paper cites MeetingBank: A benchmark dataset for meeting summarization.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation MeetingBank: A benchmark dataset for meeting summarization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.841893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.841893Z digest=sha256:4cc01812c72e9f424b6f160d8a8027addc39bdb76afc8862d34b0f726ab24627

Observation 3c36242e-055c-4d68-baa6-1c03b07921da · outbound

This paper cites Answer-consistent chain-of-thought reinforcement learning for multi-modal large langauge models.arXiv preprint arXiv:2510.10104, 2025.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Answer-consistent chain-of-thought reinforcement learning for multi-modal large langauge models.arXiv preprint arXiv:2510.10104, 2025

Reference 19

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:29:04.467803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.846255Z digest=sha256:18817de44ec2912ac1fa470e33d7d3c5971c01aa0aa879cb7557daef46ea224b

Observation 795c7606-06d9-4595-9f60-48d13b5bff96 · outbound

This paper cites OpenAI o1 System Card.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.849865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.849865Z digest=sha256:ae14566a504f02eb2a05e6ed02981575e9a13ba0ca6c9e2f1521e116574d2153

Observation aa3ffc40-e62f-4a9e-ba83-fd47bd5756da · outbound

This paper cites Understanding r1- zero-like training: A critical perspective.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Understanding r1- zero-like training: A critical perspective

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.984900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.854986Z digest=sha256:bb8b20a1f3b7b65709a97097abf4f05661af3820020b4aa3ecaaffded07318ed

Observation bbf87a8c-ddbe-4498-8be6-dde95ef68651 · outbound

This paper cites Decoupled Weight Decay Regularization.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Decoupled Weight Decay Regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.858538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.858538Z digest=sha256:a009d78c4be7023265dfdd0ec452cd24e77c543b14dc9367c4003cbb874a6f80

Observation e4fd2436-2fc2-4385-b832-dbd69df5c953 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connectionism, 2025.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation On-policy distillation.Thinking Machines Lab: Connectionism, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.861966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.861966Z digest=sha256:5176b7625223e13efdac4b216be07f1b31c596b114842e8d9cf319cd9591667f

Observation 52b88f80-ef63-4ff9-8573-7d2bbb375abc · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.866650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.866650Z digest=sha256:024a7c94e94785432b8d9c97c150cf389abb30aa55375751543f740bf2f23ea0

Observation f8ec9b47-7531-485b-acc9-1ae53493c044 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Reward augmented maximum likelihood for neural structured prediction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.967759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.872377Z digest=sha256:712314910a03655a39eb39ce004e875a4a1220293ca2ce38d359e25a1ecb9e65

Observation a1ef6423-ea6a-4ad1-a018-93043bdfcc5d · outbound

This paper cites Iterative reasoning preference optimization.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Iterative reasoning preference optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.951654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.876731Z digest=sha256:4aa958b8e6dbeef1bae14d631ec65dbf0cedb4913084554cc1d33508872faa42

Observation 51f6a1b5-835f-47b5-8420-966b7dbfdea3 · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Rethinking the Trust Region in LLM Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.881817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.881817Z digest=sha256:bd68a794466d35968ee54ab8778b9c18e60394ccbb4cb85d1f9a8e59f7c0f9bb

Observation 55fe61b9-3264-4cfd-b253-47e751c87af8 · outbound

This paper cites an unresolved cited work.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.889681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.889681Z digest=sha256:a187e9b5853ad25a430417d5f62eabf5b34e835f96cd3555405cb37cfe36ab4b

Observation 179e6939-895e-4b35-80e7-2111e926cea2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.911990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.900830Z digest=sha256:9353c0323c934912c4e406e269d169a9c5d92c7e72eec306b2f9168f125c4318

Observation 3fe1534d-def2-40eb-a0e7-b2d07c77d067 · outbound

This paper cites an unresolved cited work.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:29:04.890952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.905739Z digest=sha256:373c0835c21f5aed4d466faee0edd3a69140b67ce0fe63f6742e52bc7df930f7

Observation 453e334e-fed7-4919-985c-4d93a7cf781a · outbound

This paper cites Optimal completion distillation for sequence learning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Optimal completion distillation for sequence learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.856233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.916418Z digest=sha256:effc819ac28666c03f1141ab15c9e039edda780c9b6b0a813a5c64f303dd9b3f

Observation c449c312-c2ba-4e89-8204-79e2a2a4606b · outbound

This paper cites Proximal Policy Optimization Algorithms.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.921502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.921502Z digest=sha256:4b0d7573e4dd509fc00b8f105dbd1e13b78b1ad6af5452fbe756e4c0155e6bb1

Observation 8f039705-4656-4b94-8dfe-0fa5e4a2df2b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.926196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.926196Z digest=sha256:244d8c8a17398158d72e1c32a1a1003f018917cd1d27cde2e50ba45a5dbbf1d3

Observation a0f34d5e-7215-45a8-a781-3358ef13b455 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.931344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.931344Z digest=sha256:d95dbd6ecf9a665f4b6cca9634c7ed2e8402482690354a1205319350764ee49e

Observation c91d0b00-6ba4-407d-9136-bf043f5db81e · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.935373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.935373Z digest=sha256:200155c24723bc83844b8c9ca369d212394595f3cc4112d7b9cf831e4a5d5c1c

Observation d1e7ee22-279d-4e3f-b327-ce76786d6063 · outbound

This paper cites Maximum likelihood reinforcement learning.arXiv preprint arXiv:2602.02710, 2026.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Maximum likelihood reinforcement learning.arXiv preprint arXiv:2602.02710, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.940136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.940136Z digest=sha256:086535d09c9a6d83a11e94a5572c1ece6a10829456ffb36591782003e2ea43eb

Observation c3f0d3d6-0693-4bf6-9d57-009fef5ea153 · outbound

This paper cites Vl-rethinker: Incentivizing self- reflection of vision-language models with reinforcement learning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Vl-rethinker: Incentivizing self- reflection of vision-language models with reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.844077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.945359Z digest=sha256:9e086ce17c7e248b8a433bf12e99a112516f724750b56d44323331acd51878cf

Observation 5b43ec2a-b9e5-4ed5-a44e-8c85cc8456b3 · outbound

This paper cites Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.830899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.949476Z digest=sha256:bb69e1a8e942f78176724e2191b2e1b2bfef913391de071205bb1df626ecdbcb

Observation 002ece37-9b4a-47f5-9574-3e38074cefb9 · outbound

This paper cites Thoughts are all over the place: On the underthinking of long reasoning models.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Thoughts are all over the place: On the underthinking of long reasoning models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.812235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.953507Z digest=sha256:782fc912f720ba5a82282cb0c5bae0b5c545ae480df3348ad06ad62597562f8f

Observation 5eaaa617-1b79-453c-91ba-0f1d303e262f · outbound

This paper cites Sportr: A benchmark for multimodal large language model reasoning in sports, 2026.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Sportr: A benchmark for multimodal large language model reasoning in sports, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.957406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.957406Z digest=sha256:302bfe03226f6e3435a7db4836b8044ab40a89c654c9062fea59fe0e094bd893

Observation d40d4dbe-5ab3-4b13-98ad-cca8b8785645 · outbound

This paper cites Proxythinker: Test-time guidance through small visual reasoners.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Proxythinker: Test-time guidance through small visual reasoners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.797570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.961440Z digest=sha256:e84701d9012f6fa78c700f7aba545abadf5af170c048d482baee384b6ebf7ece

Observation 2e12425e-c792-4fb2-82e7-bcae2d64769a · outbound

This paper cites Qwen3 Technical Report.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Qwen3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.965787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.965787Z digest=sha256:72ad7c94b9b3a2d4e40ceb3a7f009dedfb41cf3b4e2526608e78f7321c60509f

Observation b3a193c2-0a9a-41db-833d-8db2ba8ccc8a · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation DAPO: An open-source LLM reinforcement learning system at scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.970117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.970117Z digest=sha256:49f2e3b21268cc5e56bc06c9ee94facd5c567ec73d3fcb4af1d95b2bb5c5d0f9

Observation 0d210592-509d-4f11-8a92-9c657cc31ef6 · outbound

This paper cites STar: Bootstrapping reasoning with reasoning.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation STar: Bootstrapping reasoning with reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.974838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.974838Z digest=sha256:6f5a7973ae8f67012a2a820ffedd974adb9c15be58b6c0e51936020a50c45739

Observation b00b7bd6-e5af-4951-b3f8-389e1fe0f17d · outbound

This paper cites KAT-V1: Kwai-AutoThink Technical Report.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation KAT-V1: Kwai-AutoThink Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.979079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.979079Z digest=sha256:0a865ad406cfc48453005f3298729ababf8dc2049d85bd7acb30404fc40c36a3

Observation 7b41180f-ec8e-476d-ad7e-6b1f00cae5c5 · outbound

This paper cites Reinforcing general reasoning without verifiers.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Reinforcing general reasoning without verifiers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.757871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.983724Z digest=sha256:1e69947bae1abe451e207ff3456692c5d8924fba610c2bec9c3869be299be99f

Observation bb7d2340-b2b4-46a5-866f-c891faa92bc3 · outbound

This paper cites an unresolved cited work.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:29:04.871224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.911135Z digest=sha256:b4c57871e7e08a3b31eee38c3a63df1e60a46111bc32a37da4cb3367ebf919c4

Observation 0ab855d8-2657-4615-997b-dada58adeac7 · outbound

This paper cites Please reason step by step, and put your final answer within \boxed{}.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Please reason step by step, and put your final answer within \boxed{}

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:29:04.742339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:29:03.988127Z digest=sha256:985aa5837c53314c568a8f6dea313902cc7762dcfa1066c5e4eeff42338a06a9

Pith citing papers

No inbound Pith citation observations are available.