Pith. sign in

Paper Citation Record · LEDGER

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

As of 20 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 31 inbound Pith citation observations for arXiv:2506.03106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03106 v7

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:15:35.267794Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:11:38.110587Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:29:15.119817Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved11
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a579f800-b35c-4935-90c5-4a76dec37ed3 · outbound

This paper cites needle in a haystack.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback needle in a haystack

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.779353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.192749Z digest=sha256:a88c4ebb3b6a9009ea0dd354684ac2687706e2dda117cc7c84b2a6140204b172

Observation 8f534efd-4ea8-42f6-8d0b-09cda7837070 · outbound

This paper cites While the worst-case complexity remains dimT E(H, ℓ, ϵ)≈O(|S|L), the critique acts as a pruning signal.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback While the worst-case complexity remains dimT E(H, ℓ, ϵ)≈O(|S|L), the critique acts as a pruning signal

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.769972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.195999Z digest=sha256:0f643d755e3e40f69319c164a9deb2724fb2786f0c7de02bdacc381c105a8231

Observation b7ad4a6d-2136-4acf-a79a-1b4869b9c1c7 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.690194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.222317Z digest=sha256:b295922867d6cb58630226553f35810326f2f863d425c191544f4aa2c64fb40d

Observation 90b9b898-9e89-4f9c-ac0c-dbff6c9078c9 · outbound

This paper cites Wang, Y ., Yue, X., and Chen, W.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wang, Y ., Yue, X., and Chen, W

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.170943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.170943Z digest=sha256:324cd4d03b61250484654f7cc6af5d258da6602963b5fe510413827b9595f96c

Observation ad29d770-e3e6-4bc6-877a-0c21c185e548 · outbound

This paper cites The correct maximum value, as derived from a proper analysis, should be 10 3.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The correct maximum value, as derived from a proper analysis, should be 10 3

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.556971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.267794Z digest=sha256:f1c25e84d7c6614122a6e16e91ea2717a00103aca2bb53e35ff3b01350964d61

Observation ef39f899-e63f-45a1-b692-4364d4fae2ea · outbound

This paper cites Qwen3 Technical Report.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Qwen3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.178558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.178558Z digest=sha256:c14767a8e2341c1b8d20035ad86180cc927e511834a3c473a93a223105141ff7

Observation 39b23d80-f7f2-4675-ad68-235586a58802 · outbound

This paper cites Zhang, X., Peng, B., Li, K., Zhou, J., and Meng, H.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Zhang, X., Peng, B., Li, K., Zhou, J., and Meng, H

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.181921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.181921Z digest=sha256:61edab152506bc3769b064b612dc21b5d6b83150b36411ee188dbc64e6d1f820

Observation 0d54d17d-9fa2-43c1-bf9a-a3de9d99d77a · outbound

This paper cites correct” and “incorrect.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback correct” and “incorrect

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:15:35.185214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.185214Z digest=sha256:518845894992606736c4fd2546290e07142bce2366e212ee2b09c08f513b8b8c

Observation 5f77a2b8-0532-498f-aa9c-3d95e421f3f1 · outbound

This paper cites As illustrated in Figure 8, this function is bounded between (0,1) , where x represents the token probability of the policy.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback As illustrated in Figure 8, this function is bounded between (0,1) , where x represents the token probability of the policy

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.789361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.189419Z digest=sha256:1c54925ae9fbc50f4782e06d29c7da448381314c50d9c7ee9d37240c87193e00

Observation 7ffef02b-96c2-48ce-a18e-b35fb89a5553 · outbound

This paper cites In this setting, the feedback function is simply the reward itself, fη(a) =r(a).

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback In this setting, the feedback function is simply the reward itself, fη(a) =r(a)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.760582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.199170Z digest=sha256:1d78407c1ee0a9e32e897bb067cd637a8b9073dd22dd0b05a1b60598078590df

Observation f008ba4f-474c-4083-8873-e2fee8bd2e6d · outbound

This paper cites The probability of finding the unique optimal solution a∗ is equivalent to sampling the correct element from a set of sizedwithout replacement.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The probability of finding the unique optimal solution a∗ is equivalent to sampling the correct element from a set of sizedwithout replacement

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.750597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.202422Z digest=sha256:37d2b5c82edac7465892f4430b6dfdd474c6ed5bc28f9edd794f6fa82ad174a4

Observation 1b7aa275-33cb-445b-9173-d95966e02ccb · outbound

This paper cites Since γ≪1 , the gradient magnitude is significantly dampened.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Since γ≪1 , the gradient magnitude is significantly dampened

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.739780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.205872Z digest=sha256:8400e2f056d86f7ace39166b525f239139516b7bd44e2153f6717da69e1dfb23

Observation 9f3ca082-9d4f-44c2-8694-c307e8481ab4 · outbound

This paper cites weaker refinement,.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback weaker refinement,

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:15:35.729061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.209183Z digest=sha256:cbee78749cf672d21311f969e525a495c06acd7a77bf623499dd6cefdff574c6

Observation 119a2374-8a19-4e5b-94a6-3c2f71405470 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.708975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.215979Z digest=sha256:df60ec402a6b3b61917b870dc5ae5e6fba7e8e53b7f20c6079213432a4eb62cf

Observation c9c10e5b-01ee-4e8c-9b10-4407e713a172 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.699270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.219150Z digest=sha256:f23a5f412b439c6f9020269ecfa4955c24d7190766d2b4fa755a6dd6f0e579e8

Observation 7b28c74e-0063-4c28-baa5-b1582b90f90b · outbound

This paper cites Conclusion:.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Conclusion:

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.680966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.225726Z digest=sha256:9750cbf22fb8e60497dd8ebd8043fb5b9ebe48887254ec318442416e3bac750e

Observation 459c337b-f614-44a2-b336-4b4a4742ad76 · outbound

This paper cites Calculate the cosine of the angle of the axial section of the cone at the vertex which is also the apex of the cone.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Calculate the cosine of the angle of the axial section of the cone at the vertex which is also the apex of the cone

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.671451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.229585Z digest=sha256:d3545eb8cf51e1b7e970d7eb5384261232f8cf740aceb86af9d35c6bde002d27

Observation b3b5fbd5-10bb-4732-bb84-ca5c18c21d16 · outbound

This paper cites Wait, let me check that again.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wait, let me check that again

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.661156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.233555Z digest=sha256:63fa073569f633fa1217de7e1a519f2779ecd3cfdff78bbb0070dc0208941a2c

Observation 28757a41-8f8a-4084-a7b3-6f5e700fdf57 · outbound

This paper cites an unresolved cited work.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:15:35.651769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.237179Z digest=sha256:93d95f6d9fd9a77101461ef41b5c4a8858d8d499ea53f0a56cc1244ff19680fe

Observation 13b69af0-af44-488c-b3ad-25abf3a3b693 · outbound

This paper cites Therefore, the exact value is − 9 100, and the approximate decimal is −0.09.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Therefore, the exact value is − 9 100, and the approximate decimal is −0.09

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.632214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.244109Z digest=sha256:4c5d70f277f4ff3c8ba852128aefd6d77b89e5b921b36064c3046fc59d381097

Observation f59290c6-e1c1-4789-ab3b-44eb96d8ac3c · outbound

This paper cites Alternatively, if I think about angles: A is arcsin(0.4), which is in the first quadrant, B is arcsin(0.5) which is π/6, also first quadrant.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Alternatively, if I think about angles: A is arcsin(0.4), which is in the first quadrant, B is arcsin(0.5) which is π/6, also first quadrant

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.621418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.247160Z digest=sha256:936189a8bdfb5bb9b5e931adbdfcf1525af700d9fc1158f09f5bc5befa796ec7

Observation 93a02610-6565-4c74-b972-b2944de36dba · outbound

This paper cites Alternatively, using complex numbers or other methods? Maybe not necessary.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Alternatively, using complex numbers or other methods? Maybe not necessary

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.610767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.250430Z digest=sha256:b9cd5152ba14f44897a928bc907844d18dae3b3ce02b76792e8197de7aa73466

Observation f43a9e87-3503-4e60-8d7d-88999f07724e · outbound

This paper cites The trigonometric substitution should be used more carefully, ensuring that the constraint is satisfied throughout.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The trigonometric substitution should be used more carefully, ensuring that the constraint is satisfied throughout

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.600273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.254352Z digest=sha256:d1f9aeff96d4f4685adfd9e86594004b71cf99054512f63982277b0c1b68e43c

Observation 38bb0bc3-ed9d-43a2-a587-2e083088dddc · outbound

This paper cites The identities used do not lead to a valid simplification of the expression.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The identities used do not lead to a valid simplification of the expression

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.589374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.257744Z digest=sha256:2ceb6dbc90b201b6bcbe29d5e0164b4ff9f15c70ecbf6a8e7a78abb0cb034c4d

Observation 1857b7b4-4f3e-4372-939d-e47bc3240fc5 · outbound

This paper cites The derivative should be taken with respect to the correct variables, and the critical points should be found accurately.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The derivative should be taken with respect to the correct variables, and the critical points should be found accurately

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.579589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.261000Z digest=sha256:55b7a20714af29a8196349194f60d0694e713e425d8624c6a95e1c6d98b16ded

Observation 47df313b-3674-4c95-a2a8-6db297b6c724 · outbound

This paper cites The values chosen fora,b, andcdo not satisfy the constraintabc+a+c=b.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback The values chosen fora,b, andcdo not satisfy the constraintabc+a+c=b

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.568539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.264265Z digest=sha256:f2e099810e818234ed80bbf0ced27cf8eaf2980789697b9e69ab0047c059a7e2

Observation fe9f3e00-79b7-48ae-8afb-028fe77f3064 · outbound

This paper cites Therefore, the value of the original expression is −9 100.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Therefore, the value of the original expression is −9 100

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.641688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.240692Z digest=sha256:3185c3d1a0d5626e377eab4a254f6718437fd78e83064c47efb2e91620513293

Observation 9a8caf7d-b4f7-4f1e-96d6-2abf3d3b9665 · outbound

This paper cites Wait,.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Wait,

Reference 150

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:15:35.718737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:15:35.212533Z digest=sha256:946cc35652d04c94da64937fb947bca575b164561dd8b5d62dd704d94b198ca4

Observation 25931b4a-6bc6-4e11-96a3-e127664f2d9a · outbound

This paper cites Training language models to follow instructions with human feedback.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Training language models to follow instructions with human feedback

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.163327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.163327Z digest=sha256:5fb7289a389a0a727d7005aec9600200a846eaa612d861a673de3ac0d3cf5237

Observation 8ffcab6a-b67c-4d0d-9558-9ec2d5b83103 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.167088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.167088Z digest=sha256:e0eb706384e8f89ab7b879b69766637a36cb46cffae8bc6b6b743838653e5d88

Observation d350a58a-8bf7-4c42-8595-7e3dbfa237f4 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.158495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.158495Z digest=sha256:6aea5c7e53199717a52def3755b89b1325bea83d3dadc48029dcbf0cb219adce

Observation 45364875-fb18-435c-8de7-197aa7aacc8a · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Learning to Reason under Off-Policy Guidance

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:35.175237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:35.175237Z digest=sha256:9c0021268bf7f9b3e219f84446e16865218eef619d8d972ae8108790d0bfd6d6

Pith citing papers

Observation ef08bf37-6769-454f-b0b7-bcb326090e57 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:a6823113094a91b2184157e2fd2a5e2208a05a65c74c14159954bfe159e209f6

Observation 890f6533-5d3c-4937-8e75-fd63a407581d · inbound

MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning cites this paper.

MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:19.009023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:47:19.009023Z digest=sha256:70fbbe5d5c53cc93d1d90b3f28caaf7b9f805dd460703246126da2574ad1feaf

Observation d5df3f36-6a6e-4d02-a269-d7f95777e969 · inbound

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning cites this paper.

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:53:36.062617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:53:36.062617Z digest=sha256:d1dcb8ce55db61a3c04907e9234fc365156ddc5ba5bd0beda2d06b5dfc25e76d

Observation a04af6ba-5ecc-4578-9d72-008d996753aa · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:12.915233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:12.915233Z digest=sha256:fde973388b70dc5865cc8ccc9e0a3f431f0f86247730de5753d56323a524dd71

Observation 27f1b512-88a9-4cbb-929a-98e9068fa0ce · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.877236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.877236Z digest=sha256:21d29eae1192f77305406fcb2d4ec631527aa83faad84b2667f8c9de071f9e49

Observation 83044b5c-6cf1-4702-a64e-c12687fbd697 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:a2a5613b3de61ef65b177e53cc58de8087cad1ece1e6d55c87192085df84916a

Observation 9af02442-d983-4e7b-b991-45aa307c7da0 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:3a9cf5a9ed7a7f6f985bc82a59a6de98943a79f7fe8da46f0a90a6f1ea1928b0

Observation 3baccb9b-70e2-4b70-be63-a0c3d0af0ffd · inbound

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens cites this paper.

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T15:21:55.984824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:21:55.984824Z digest=sha256:b834ce2e709b30deadc7fe5af12b669eb057e7db0b444d8fa210b991cdd680c4

Observation 9fcaacfb-fa23-4a17-aee4-351c5a09750b · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:22.180784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:22.180784Z digest=sha256:e4026570102e10f171d6497d8b916ea3e62ae73b956b4272bd14f935ab018272

Observation 00596522-fd42-4123-951a-941771bc4d7c · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:18:10.087258Z digest=sha256:8eb4f68425bd7d4c57148adbee1b764d92e8638c424f4af31a9f82fdf9791a71

Observation 61bce089-f9ff-404c-b03d-b31eef3ffdeb · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T06:31:13.232880Z digest=sha256:1dac76f6baaeee76ce99d3fe9091b77990a1e4d48815f575916f087b908257ea

Observation 266e7451-b303-4c0b-a695-4a57d85bcf95 · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:5faeeb9b5b1033de79e8488cc337eae4f3472a662f59802cbafbf96f9d35f78a

Observation 11c880c2-cb3d-4c40-8b8a-947b0597696c · inbound

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance cites this paper.

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T08:05:36.206946Z digest=sha256:ea2f3896a7332820b4d5a435f8b2e432f83243edd2f310c07ac309bb93158b6d

Observation 0be93c8b-7545-4a9b-938e-05314ec92922 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:8a62f280cd684ebc67fcaecee66bad3247af4ee52c4036fdd328dc0d4b64b2eb

Observation 48110f5b-b2ff-4350-abed-0302f3f83bdf · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:ce49eb4ea640abb4a2a676bacd58ebff75d1046e88b6021de0adad11b070b0b1

Observation 49fb552d-fa19-4b23-9a97-b8ac7eeb583e · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:e0eb13b98f338e296e6321d678255720c7098c17fc53a474eae83264616fac36

Observation c92ccc7a-cd02-4fe1-bf07-cb4b694efa8a · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:a6799f45bbbab5fd1e1f608f64bf2e4bc70c04a43691d6360d315e9c4dc6ca75

Observation 88a1d0f1-93d7-4601-875a-9bc7a2fa9e35 · inbound

Multi-Rollout On-Policy Distillation via Peer Successes and Failures cites this paper.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T21:29:36.803832Z digest=sha256:9413abb1f7971e83f98b8733419363365b3e8fd39e6aa460651ca5e404b6b78f

Observation 1079fdf3-9892-4d3b-833e-04649fba1d4d · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:0dc89f0176a971dda96873eb6e6d3c889135c780be6a2a3cd0104344e960eb56

Observation 655f3118-46d8-4a0a-bbeb-1c7b7de284f2 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:41:25.205368Z digest=sha256:8028612a4beb32b0904c1f4ff3128f59c54e19a2af6a0b3a7a2d1fabb5b3c66d

Observation 257ced7d-efc5-48ce-92a7-13ba8301d3d2 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:32.244687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:32.244687Z digest=sha256:52ff5fd5cc55c0d8d67611df79bb80401ad3d2dfdf4f21f48474279a614903c2

Observation 903fc1e0-9e56-49c6-8d49-e5453ff4554e · inbound

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning cites this paper.

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T22:51:56.666980Z digest=sha256:58c8f3d3de7bac2ce9fa5f5402881ca7c98dda889f43a211328fa023dbf8a632

Observation 0b8c0dfe-6fa1-4a06-81dd-689eda059d4e · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:f16c7f41dd7933603c28b53777c90873c1dcf6eb94f95a9767e792097f6462f1

Observation 83749c53-f9af-400e-8a4e-5f7b7a0b53fd · inbound

Reinforcing Human Behavior Simulation via Verbal Feedback cites this paper.

Reinforcing Human Behavior Simulation via Verbal Feedback Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:03.715193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T07:21:48.649289Z digest=sha256:cae512a554430acb1bf8db317ad746402d508ea6e3b8fb0941a4be84acb91a63

Observation 2c37c286-2716-4642-a16e-14e9f9a86d98 · inbound

RL with Learnable Textual Feedback: A Bilevel Approach cites this paper.

RL with Learnable Textual Feedback: A Bilevel Approach Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.361285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:36:31.047856Z digest=sha256:202b386796a499a1c32bd23993cfcf1540aca4250642206c76a3e2ae76038381

Observation 48d2b711-066b-41b2-9cc9-1acf83b95f91 · inbound

Credit Assignment with Resets in Language Model Reasoning cites this paper.

Credit Assignment with Resets in Language Model Reasoning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:53:59.255224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T21:50:18.827822Z digest=sha256:dc41524364c48d6555d4e4e3b00f9950d3ec70d2f58e152924ae3b1f27d2c81b

Observation 660eefdb-b031-4778-b04d-c84b18bac9e8 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 189

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.705760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:eae235223a43210f52ac083bce4460f46007e9cbce0eb2c261d9ac8564bf2eb8

Observation 5fe53e6d-5bfa-441c-8a6d-0c93514215e6 · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.122126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:29abd787873400de7fecc7dce5b31306b30e7e32c90979ee92406f54228d93a8

Observation c34e7e5f-a8a7-40f4-b444-3fab060e95ae · inbound

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry cites this paper.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:f9bf91ee13354893117208b9e4774e2cd839a8fb085568b6430733a0e30dd0ab

Observation 483e6977-b08e-4d13-b184-6573aeae10e3 · inbound

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks cites this paper.

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:59.623378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:59.623378Z digest=sha256:b22b0edfd93ad9716330af98612a9ef7b12bd5894ef8b0232f5276eeb7f60e53

Observation fe2f8e82-8e4d-4241-9d0c-09b7c8812357 · inbound

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning cites this paper.

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:38.110587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:11:38.110587Z digest=sha256:a5b38592b319e9717e5a8a797c7134732ca136f8534b9747d446771bbe65df06