Pith. sign in

Paper Citation Record · LEDGER

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2506.04821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04821 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:36:35.797792Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T05:46:26.938277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 60bb10ba-e173-40fe-a364-2bbf2a34bc19 · outbound

This paper cites H.; Ol s \'a k, M.; Yang, X.; Nguyen, H.; Menegali, M.; Jung, J.; Verma, V.; Le, Q.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning H.; Ol s \'a k, M.; Yang, X.; Nguyen, H.; Menegali, M.; Jung, J.; Verma, V.; Le, Q

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.553539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.553539Z digest=sha256:60fd669f705eb3ab99f3e8d199b69911641caf562d58ae66744dfbbc3dd72051

Observation 3a740ac3-fda4-4759-89b2-cbc2857280b9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.587424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.587424Z digest=sha256:8fa78649a06b54d4eb08bcf5af7983e98f933bb5ae5b6dab5dc9e2818e655a98

Observation 823dc748-acc7-43ca-9ddc-209620e0dde3 · outbound

This paper cites an unresolved cited work.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:36:36.222441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:36:35.669262Z digest=sha256:272c9e44e84463ab00f20100310b3250bfdca733a1bf7f25181297379fbe1f2c

Observation 1682acc7-86f5-49fb-b6bd-999fe06814c6 · outbound

This paper cites Puzzle Solving using Reasoning of Large Language Models: A Survey.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Puzzle Solving using Reasoning of Large Language Models: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.720092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.720092Z digest=sha256:eec840fa39fecc93dbc5e43024a8f33c6cb21c4ce6f652b39e054147148c79cb

Observation db483bbe-601a-4225-a1e6-25cf69b12a15 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.729013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.729013Z digest=sha256:c8ddf43abce74a8908e94c4b8f78d5afeea515b66793310f12b99f6be3129e3d

Observation eda3e062-da76-44d8-b78c-2af5ad068bfd · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.733352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.733352Z digest=sha256:e282c63f902374b86376584edbbfb4d81f8e7dc922262d89060a805e9ae0c1fe

Observation d6dee929-919a-4567-b871-7e1eaa86b420 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.738540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.738540Z digest=sha256:5b7bddb6c03f9262f50a0a1ca1b3e561a56f50bcb283b034be8820c6cff770f3

Observation ddb1a080-64b2-4276-aa7c-5464b4f3a081 · outbound

This paper cites an unresolved cited work.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.742751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.742751Z digest=sha256:0566f6ff9b77024cc3e5e88afb88a8465fa04c1c95482fefac1f06f463cb7072

Observation a0974290-344d-4535-8144-9be15832edb5 · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.746562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.746562Z digest=sha256:31736b6b361d07bd8137b3f0659666461d878918e4a49dafc8eeefd9e6113998

Observation 436dfa2e-86e0-4c0d-ad85-be963254fe7c · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.751083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.751083Z digest=sha256:3865ffebe7a7156973be577afbeb41448b2d427b52620fa288faeb50e8998375

Observation 013e1b47-08f1-413b-91cb-09753a8c9a6d · outbound

This paper cites Instruction Tuning with GPT-4.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Instruction Tuning with GPT-4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.755270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.755270Z digest=sha256:7cd655d10ddb7c10c906866f12c3943b4fc03f0c71c8da4abb356c19ba6baaaa

Observation 80a8c158-7aef-491f-89e5-9a5cb7030b09 · outbound

This paper cites an unresolved cited work.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:36:36.199387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:36:35.759659Z digest=sha256:01308e7322862e1f041278da03489fb3db404dc8a1e31697d1cdcc39c8ca5c61

Observation 975b78e7-bd62-4b63-be5b-a04e75c8c1de · outbound

This paper cites DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.764026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.764026Z digest=sha256:a9f2972f0fe26ec4ba8a025101a16319abc75384b518408d7c60fada87739ac5

Observation 6f1a336d-4493-44be-abff-52b47b38eadc · outbound

This paper cites an unresolved cited work.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.768129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.768129Z digest=sha256:353f3baff9c6c60931a9dd5d484b6f91a55df1e96eea4dca8093cc7096bba450

Observation eb642055-069f-49a9-a58c-65798539348c · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.772152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.772152Z digest=sha256:2ba1556aa53be134e5b7e13a7624474428fee61b04d237416f5307656c4589ed

Observation bf4644b6-9ade-4cd6-be97-82153afcb6d6 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.776289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.776289Z digest=sha256:6cb4cf570e0eb7fad4b76360b373f6eedc38adeb8a8bf3f39940aaf05b846ab8

Observation 80cd305b-df40-4fd1-a4b1-858eaebb26a4 · outbound

This paper cites Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.780423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.780423Z digest=sha256:2b0213fc958e9efc7de201bfd19a4d2a3ad4185a126471926f7e963775fb54be

Observation 887d9805-7ec3-461b-a203-4cb49fa4379a · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.784341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.784341Z digest=sha256:105a2f3083f7bcab257276f3470e0395434a87c2f956017faf9eac0fe5db40ea

Observation af310fad-6d59-4dd3-bc84-cfa9d1623182 · outbound

This paper cites Leanabell-Prover: Posttraining Scaling in Formal Reasoning.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning Leanabell-Prover: Posttraining Scaling in Formal Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.788897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.788897Z digest=sha256:808239fe4ad68a81c62d15fc4f8c505648b5951ea463dac34bdce9e2c370cc8e

Observation e6724016-5825-423b-99c1-3eb68387ef20 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning , " * write output.state after.block = add.period write newline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.793331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.793331Z digest=sha256:4b29633a7295d301ee3355e1e80c3518e1f49ec992bd1ad3e664fd5a7fb5cb8d

Observation 0b5b9d75-a6e3-4e33-a920-9d2850aa4bf7 · outbound

This paper cites write newline.

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning write newline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:35.797792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:36:35.797792Z digest=sha256:07d3a8fa719bffbdce9f7fa5926bfb6117a46a6b0b953f5dea49f2e0e46fe094

Pith citing papers

Observation 623f87b7-4031-403d-98a2-f37cb1f0fdf9 · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.323771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:b34fafa266759f75955c197f0ee9f8b367dcb5ce4ce3e2a82a902b061476346d

Observation 999d9f8e-21e8-4887-9e79-2dd05598ca63 · inbound

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models cites this paper.

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:18:15.745450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T23:17:29.025222Z digest=sha256:afa427d618f1df58d0db8a6a7aa5747c16cc5b42172e6072122b4c103da665af

Observation 07ecb561-819d-4d51-8b35-05de69aa6164 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.925001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:0c2042f8e15eebf266319c3680bb42f6c2a8bf519e5ee1992ce9fadedfec36ec

Observation fda495d6-d9a8-4a96-a2ac-0282e47ae368 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:11:28.884924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:cf1a081e869fd9b087f7b91f1a42c0efd42e46f592d6b18d0b4733022dfef1c4