Pith. sign in

Paper Citation Record · LEDGER

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

As of 13 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 13 inbound Pith citation observations for arXiv:2506.08989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08989 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:24.353008Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:17:38.690339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T18:17:33.758346Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved86
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e21e5946-6e2c-4958-aa70-05157d3d6c78 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.089154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.089154Z digest=sha256:0fb57a60596d94307461ec18ab25a735291d700ec14a241e6521950851fac7ae

Observation 21aae386-2ef6-4aa4-8d78-b48c9020da11 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.092784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.092784Z digest=sha256:72746b2ea9d69bf0199eafde961f16249ad40bd98ed94c747be32235a7847aac

Observation f7f6efc6-d15b-4409-b085-79fcfb820213 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.096216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.096216Z digest=sha256:8be3aab01839a385b45af3cfd7076ea0e96d0bc44a7c9401c17066cf18c39a0a

Observation 795cfb9a-3440-4cc7-b00a-dfeb9e47aef8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.099163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.099163Z digest=sha256:88a4bcbd5ecffec57083b22494183ee5bc992e59c2bca676c41c48b0b3b3fdc8

Observation 444bad16-8172-439e-af77-37911851bead · outbound

This paper cites Nearest neighbor pattern classification.IEEE transactions on information theory, 13(1):21–27, 1967.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Nearest neighbor pattern classification.IEEE transactions on information theory, 13(1):21–27, 1967

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.101834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.101834Z digest=sha256:efcc58581eb270ff0d0482e29760f7e5f695c280c2d520313ed9d498dfa59d8f

Observation 8494ff6c-dfad-47cd-b313-b7eec9c5ea07 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.104449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.104449Z digest=sha256:a0e8138de7e5f74c26008617ba4c0361b98404c9e055a1b412756deb2c2f25c4

Observation d618a45c-bbd8-44c1-b0e7-8ceaa4afb208 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.107393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.107393Z digest=sha256:f0b42a3d491134df11cc97434023f754c4159bba56661f73b5da4b315499d70a

Observation 65d781ef-83f7-4f1c-8f7c-0d316c9fd3c2 · outbound

This paper cites The Llama 3 Herd of Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.110140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.110140Z digest=sha256:093b043309700106287897bc0cc2d665b21bd615e40052ae8373e53e44b67987

Observation 06e02db3-49d4-4303-b2bb-2dae77510a5a · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.112886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.112886Z digest=sha256:ef1241b8786bc25c83e73fec02bb428b8c673f802226b7e1e7d3eb0cbf0aabd2

Observation b6a6c548-6f5f-46b1-9fa2-e82483501392 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.115411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.115411Z digest=sha256:8199cffba7f05d1f024bf939b88670a604334b2bb8c2862ba1857ab5fcdf5a51

Observation d58dc3c4-d2cf-4252-a6cd-fa0e7567d73e · outbound

This paper cites Olympiad- bench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Olympiad- bench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.117978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.117978Z digest=sha256:421b0be88afdb2d4e8f590118dea4a3d2b842bd9487baaa41d5d722c070cf0cd

Observation 9d429c5a-2a21-4c17-9e9d-333aca6d0183 · outbound

This paper cites Skyworkopenreasonerseries.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Skyworkopenreasonerseries

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.120427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.120427Z digest=sha256:2c87e492512dc4869ce96518a6b698c1b4d70f29c25916b8cde9f8bcbacb2928

Observation acdea8a8-df57-4221-90d7-adb65cde3e04 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.Sort, 2(4): 0–6, 2021.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Measuring mathematical problem solving with the math dataset.Sort, 2(4): 0–6, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.122804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.122804Z digest=sha256:47cfc9f289770b143ecd85fc28dce69ff309c6a3d5c7c3e67244b1514603d6f8

Observation 5fdc6bfa-d305-422f-9113-6ae52e74720b · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.125078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.125078Z digest=sha256:5ad04036cc5a046710f012651e1383aa0504a650945426cc3d224dd1220cb65e

Observation 408a24fd-cfcc-4564-b8ff-ff9c928c58d3 · outbound

This paper cites Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.127813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.127813Z digest=sha256:ef63e26851399a7b1261bd0f74b0c035dc13ccd118a1841d402a67f6f07cbcd3

Observation 81d44e76-1f7a-41b7-a26f-723b456eaa4f · outbound

This paper cites OpenAI o1 System Card.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.130388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.130388Z digest=sha256:611fd1f847a5a8467b6d75a1c4135ce1be1cc7b31b418a2ce6ae9c343c300dd2

Observation 9aa4b2e8-32b4-4c68-bc0e-e7d79fd00ed8 · outbound

This paper cites Knowledge- augmented reasoning distillation for small language models in knowledge-intensive tasks.Advances in Neural Information Processing Systems, 36:48573–48602, 2023.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Knowledge- augmented reasoning distillation for small language models in knowledge-intensive tasks.Advances in Neural Information Processing Systems, 36:48573–48602, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.132956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.132956Z digest=sha256:5a19163dbc127ef84a831e50cc9192dde6fe5686a486d5c6b36d9dc47a174cf7

Observation dffebfb5-adc4-4029-9d9b-9559ba82485f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.135196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.135196Z digest=sha256:0cc3470385175a66c35eaded724b99602790b7d60794a2882fbeb4b39e90e361

Observation f6e8813c-4150-4f42-b6e7-46e2507e6ec1 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35: 3843–3857, 2022.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35: 3843–3857, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.137465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.137465Z digest=sha256:70ef603a0deb0111fe4ccb8fa7875dac3db2edda1d45604cf4c5e79b1a8228f5

Observation f8f00e26-8ce6-46d5-8511-c212abc361c4 · outbound

This paper cites Common 7B Language Models Already Possess Strong Math Capabilities.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Common 7B Language Models Already Possess Strong Math Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.139853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.139853Z digest=sha256:d660c27fee25101a4760db72f6d7780e6d8e40e78950dbe0b54e771b190ecc40

Observation f00847c0-fec3-4aee-bc29-da9ab492b095 · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge.arXiv preprint arXiv:2411.16594, 2024.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning From generation to judgment: Opportunities and challenges of llm-as-a-judge.arXiv preprint arXiv:2411.16594, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.142570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.142570Z digest=sha256:502d5d3707cb775740fadf901e3197765d6cf11c3a06007c2313d945a36ee99d

Observation 67b2ffa4-f2eb-4445-bf66-67a400742f9d · outbound

This paper cites LIMR: Less is More for RL Scaling.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMR: Less is More for RL Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.145061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.145061Z digest=sha256:642d089f29d08b9d67ba5327915841117d60d9e002488b1f254b524b23fbd66c

Observation 2c93907a-1052-42cc-a055-453eef601138 · outbound

This paper cites TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.147834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.147834Z digest=sha256:233d148782e13f8bd29aeae5ec115eb91a7bd5375e824673e996e831377078a5

Observation 0527b2e1-f295-4287-a70d-9bbb64f064f3 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.150655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.150655Z digest=sha256:cd2924193a70c8d31172134052f5163c054a0b9897dd50782b6431edc59a91a4

Observation 0e47cd1a-4d78-45bc-ada1-90e2f1782f55 · outbound

This paper cites Task Oriented In-Domain Data Augmentation.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Task Oriented In-Domain Data Augmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.153512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.153512Z digest=sha256:bda8fbeb0b24caba34a159ee9143a9d58b81e617b7c5e896f2e507a3182d9ddd

Observation c0173a58-a462-435d-ab26-f430656a5616 · outbound

This paper cites Let’s verify step by step.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Let’s verify step by step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.156285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.156285Z digest=sha256:4cd19bd16fa8f91056985dd8d5759d230a0440612e63c06cd882c22b5546beae

Observation 0c1205cd-def1-4f25-bf6d-668cc55db813 · outbound

This paper cites Augmenting math word problems via iterative question composing.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Augmenting math word problems via iterative question composing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.159002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.159002Z digest=sha256:ee8834d603523aaeb1c682c1f2304a4ef165ce9ed4cca131b4acfb29d6758f8f

Observation 5b1eb7ed-159f-4a59-980c-567cb2b38c6b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.161345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.161345Z digest=sha256:4bcd2bbd71ad62f998cc1972066c5a992f2d286ea7f2f6d94d1491ce70bf0cfe

Observation 36f27c6a-6278-4eaf-b813-a876df3e527b · outbound

This paper cites SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.164213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.164213Z digest=sha256:bf8420f42b00d3b39e1e72edb671e20392e33cd8e1a7e5acefad81e05f0c9e4e

Observation 426ff685-7541-4cb4-a2f6-2a7abf44b8c2 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.166865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.166865Z digest=sha256:57a137ee00abb62cd2bbcc761fd62071a9e8809eba1aea9cf3421d1fde31f53e

Observation 0e9c5958-01a1-44c0-8ace-e662bbc76d89 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.169433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.169433Z digest=sha256:476fd4a078e64e0ca0ad0fa7cd10c806f74b9b1ed1d8fe129d523b36b77b9711

Observation b1ee18de-6c4a-4f6f-8ec9-ad19fbc7fa63 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.171772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.171772Z digest=sha256:2552248a47335cbd694b5e3026445d38065c5a369ac19c6a4654de15e745cd20

Observation 20c33ad2-872d-43ac-8203-ccf0d3296030 · outbound

This paper cites American mathematics competitions (AMC 10/12).

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning American mathematics competitions (AMC 10/12)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.174644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.174644Z digest=sha256:b197a77a9ae07c64130a30264f65d47c2f2930a1b718332cab94d8ba4e964a1c

Observation f3ea30b0-2db3-41fb-b297-c75c47c18ae7 · outbound

This paper cites American invitational mathematics examination (AIME).

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning American invitational mathematics examination (AIME)

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.176910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.176910Z digest=sha256:fb36db4e38fa1f9f4391cd61a680f370813d4c0ff07ee775a354566b321660e8

Observation 9e2c2ddc-9a31-4f0a-bdc1-83ad9eb5649e · outbound

This paper cites s1: Simple test-time scaling.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning s1: Simple test-time scaling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.181910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.181910Z digest=sha256:790b797dbb207e37fa5dc6bd3b175df25f0434608f2c9e923a507f7e3ff2f9b4

Observation eb1e8ecc-1bb2-448b-9ec6-5a8c2ea6156b · outbound

This paper cites Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.184746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.184746Z digest=sha256:0350b7ba5942fefcfc0b993c2ab4c63f0874e127493299785c8fc518d147881e

Observation 3e828667-bcf3-48a7-ad23-a9bd85c9dcfd · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730– 27744, 2022.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730– 27744, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.187499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.187499Z digest=sha256:6e0a77a6d17ea24a31aef7b45f00c9400ff416b2aab76510b2145566fafcee17

Observation 54129201-6111-4211-851a-87ceeab991c2 · outbound

This paper cites MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.190807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.190807Z digest=sha256:ba7683ec216f5996050b6a9b0888778165df3b1d60f477aff6e89cbd6aff11fe

Observation 544df257-d2cf-496f-a09e-d4f549208a86 · outbound

This paper cites Proximal Policy Optimization Algorithms.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.193880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.193880Z digest=sha256:518b97627120ed9ffd28ee9dfdb5fd650acb3d914ac40a0163033328830cb497

Observation 8736f22e-dae6-4c9b-8faf-a1c93d6de001 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.196848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.196848Z digest=sha256:bd1d0e35df215dc1905808cf43b718e7e5462d0c2176a81c8ea050c0421bcc07

Observation 24662ffb-c264-4631-af2f-7f3314206f60 · outbound

This paper cites Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.199659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.199659Z digest=sha256:132c2f072bf1e854c52d42b99eeea1d58f551ed195cdd1446ccab928a98fc22d

Observation f93e7aa4-8f1a-4dd0-8290-698a6f6bfce2 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.202855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.202855Z digest=sha256:3e2f936e699d8cb99342fd13f836838dc1203dcfdd02bc49ba4ff0075f89ae59

Observation f461f536-92b8-46b9-a689-3f5ae74fa4e8 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.205768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.205768Z digest=sha256:8e75647011affa82120e0b2e5dff82802027d18b61f5f906119be9c764f180a2

Observation af7c036a-e277-46cf-bd1e-95ab1ad8d0ef · outbound

This paper cites Large language models for data annotation and synthesis: A survey.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Large language models for data annotation and synthesis: A survey

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:25.156802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.208637Z digest=sha256:f2b6804ff40a361ab2873f4a52289f5acb7a2eb86454b4b1e8ee807dcbb11af7

Observation 3680da20-6462-4d17-83d2-5ab0e4357c1a · outbound

This paper cites Mathscale: Scaling instruction tuning for mathematical reasoning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Mathscale: Scaling instruction tuning for mathematical reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:25.148509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.211171Z digest=sha256:973c4e41443ed108efe2516ff917a34c6dc903ce4b105799d20928fb0e21d779

Observation fbe9324c-2fb3-4fd4-b76d-3acff595b7cf · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.213883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.213883Z digest=sha256:710dca60a7af820c858a34cd2bc67f87530e1caaed49b49524c1419814b98b0f

Observation 7e925c12-d9b7-4362-8e23-4fa43bbb41d0 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.216939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.216939Z digest=sha256:7abaa5d074738600833a31501b7e125f831fb41694c9348859884bea98afa1fd

Observation 2e2b9996-07ec-4440-b989-c866a5647eec · outbound

This paper cites Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving.Advances in Neural Information Processing Systems, 37:7821–7846, 2024.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving.Advances in Neural Information Processing Systems, 37:7821–7846, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:25.134887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.219750Z digest=sha256:c9334daffa3fef7ad4ff335b026102c7ab69996bd9338a20985dc2cac396c517

Observation 84ceaeff-0d18-45aa-8f79-707c113c317c · outbound

This paper cites Openmathinstruct-1: A 1.8 million math instruction tuning dataset.Advances in Neural Information Processing Systems, 37:34737–34774, 2024.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Openmathinstruct-1: A 1.8 million math instruction tuning dataset.Advances in Neural Information Processing Systems, 37:34737–34774, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:25.126951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.222297Z digest=sha256:7fba0428dca1565ee169570d146a021dd877005852f4c9023c276c6cf33774c2

Observation ead05707-8f16-40f7-9349-cb3cef9a9d93 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.224905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.224905Z digest=sha256:5a798e391dabfcc6c38ccb21e8da3403f128ba5ead56a3b494f465200b903812

Observation b30c748d-c604-49f7-9c3c-aa054484011c · outbound

This paper cites Explore the Reasoning Capability of LLMs in the Chess Testbed.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Explore the Reasoning Capability of LLMs in the Chess Testbed

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.227678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.227678Z digest=sha256:70d4dbea951ea9c3abea31aec1e291f4a813e646e44e55da83ecec015e67e54b

Observation 2afe12b7-5507-4f09-9017-3eca11d63e40 · outbound

This paper cites Examining false positives under inference scaling for mathematical reasoning.arXiv preprint arXiv:2502.06217, 2025.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Examining false positives under inference scaling for mathematical reasoning.arXiv preprint arXiv:2502.06217, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.230830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.230830Z digest=sha256:b652398d484e426ad8ff2c9b520628f5cf6b03ceece61e8a5d6863ff777a3224

Observation bf457c15-fa3a-4d33-a7ad-bcdc482e83be · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.233618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.233618Z digest=sha256:15f053d9aa7096f6e07675ccb143eeeb41ce9ac37689d3a269a947c902e440a1

Observation 3792c869-4ccb-4a8d-b2fb-37e61dfa4142 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.236534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.236534Z digest=sha256:cc6a516d36678fd6c1650952f47461ad8ef1f78ef88b412e3b1afa85feef28ab

Observation 962e3e5d-9e77-4b1e-aa22-9c407beb8a47 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.239657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.239657Z digest=sha256:81f3d61ab3f9151d09079a1780288210f9839e783dd1d011b15c1e144b86a734

Observation be76fbae-93bd-43b9-9954-a0c0b08d0d4f · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.242658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.242658Z digest=sha256:18be1c8ede07cd78bdf97b02a53a32c729576c3a082369937c5db11755732e85

Observation 76e5be5d-b6b1-4df6-9000-713d118aac47 · outbound

This paper cites Qwen2.5 Technical Report.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Qwen2.5 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.245535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.245535Z digest=sha256:2ab9d1c56dac86a13166067569c2567e3ddd48d97d915e2f96374d74cf5510cd

Observation 7cdadef9-18d5-456d-ab77-de1091d4c497 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.248156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.248156Z digest=sha256:8caac1f1c8cc5049322261adc1609f5ba161df043ab8d3090f0e3852f28b255c

Observation 49f89318-4f59-4b71-b1cb-89104f99129c · outbound

This paper cites LIMO: Less is More for Reasoning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMO: Less is More for Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.250842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.250842Z digest=sha256:014c98d55f15316a84f49e8e8813e3124a8dfd5bf4e920e02322beff3df51eb7

Observation a9f9b1f8-3f34-4f5c-b985-d526825b24e3 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.253700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.253700Z digest=sha256:a7fd6ec76c1b093468fd4a0c5f3b49761c30db9fc1c50495a7ca2e5eccb8d96c

Observation d1649c7b-2163-4771-b4ba-dd88ccf752bb · outbound

This paper cites Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.256779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.256779Z digest=sha256:9f9424e81f38301c9c3ccab68e3197c1016a6c1ef6250a707e5ed7f77f5538f5

Observation decbd585-0da8-4383-9b40-8653b184b4ee · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.260331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.260331Z digest=sha256:0b669ef41cef55ab309db8d7c2f5d4f4b5b77b2c8982edcb9f61e7bc7e75a87b

Observation 42b9d5fd-007a-4d0f-8532-9d229f7347a2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.263254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.263254Z digest=sha256:d0acc5cf790c2f936d614f02ebc2ffa3fe0af00f788c92b746d8cfcc72c43100

Observation 357d701a-568f-4f66-9ec6-53b779eeae76 · outbound

This paper cites Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.266890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.266890Z digest=sha256:0625d2a49d2f76b1c3bc8902f46d801328477de5ba43de604d65cb4d7464c98b

Observation f9d31dd3-77a0-4581-918f-ffdecd649ba1 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.270146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.270146Z digest=sha256:f2d3376c8b9ea9dc6ddedf3afb19c6cb4a5183425ee86c173b9e133cace22375

Observation 89a8ed69-2f13-41d2-b7c5-05d6d8c1b80b · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.272905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.272905Z digest=sha256:541151f9c96236a48da9dc959a8f0511d64e80cf4e50b63e85ffa8a422ac6d83

Observation f3a71132-7d0a-4ed9-8b16-0e1a34f67b49 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.275765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.275765Z digest=sha256:575cbd87b6716f7bdc4560b5be0c34c80fbae9325166f3af30b851343c5552ae

Observation 0c633bf1-ae61-4c18-a6d4-4047f6ea3b29 · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.Advances in Neural Information Processing Systems, 37:64735–64772, 2024.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Rest-mcts*: Llm self-training via process reward guided tree search.Advances in Neural Information Processing Systems, 37:64735–64772, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.278944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.278944Z digest=sha256:a194745242f7b7c79b36c1951d90283ae37e1b580bfadba6f5a97040b4bddd63

Observation 228a7f3b-d51f-4453-892b-52afc8a2d0bb · outbound

This paper cites Balancing speciality and versatility: a coarse to fine framework for supervised fine-tuning large language model.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Balancing speciality and versatility: a coarse to fine framework for supervised fine-tuning large language model

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:25.108731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.281621Z digest=sha256:2e43f47d8fbd1002f8931b03cd2bd934b9dd36996b05ded0c3fc8dfc604eefc2

Observation b8e0183b-dd8c-426c-b887-fb5a31b60961 · outbound

This paper cites Process-based Self-Rewarding Language Models.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Process-based Self-Rewarding Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.284330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.284330Z digest=sha256:26201f368a981a0353d62d5481f460c36ce6e945ba3ebdccc03e7217673a60a5

Observation 9aad07c4-6676-4390-ac88-39428790589a · outbound

This paper cites Evaluating the Performance of Large Language Models on GAOKAO Benchmark.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.287403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.287403Z digest=sha256:e7235e0b6783704709998b4c402306c26fe9292736cf30885912658b28b4eb71

Observation 3777fc28-c445-4af4-a4f2-55d684f645d5 · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.290580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.290580Z digest=sha256:49bea6f8feb0ec070bb1bb35c73f65a8e79260a8846d45e4fe625314cd904fa0

Observation 593556dd-f49d-456d-9bcf-f6e075cae8d2 · outbound

This paper cites Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models.arXiv preprint arXiv:2503.02324, 2025.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models.arXiv preprint arXiv:2503.02324, 2025

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.293756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.293756Z digest=sha256:c5affaa047c97fe85c534d393d861cbac0524eaa13df5176e2f0222965940f84

Observation 2b7f6c82-2274-4773-bfb0-c0f83c69af82 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Fine-Tuning Language Models from Human Preferences

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.296727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.296727Z digest=sha256:a9f814114298a8b6eb764e926e5327ce07d02ae7ca1e71ad64544cedc6bdff28

Observation 608a7665-1ac5-4b48-9ee8-a877948827b9 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning TTRL: Test-Time Reinforcement Learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.299735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.299735Z digest=sha256:d8c5c624359f0ba98342123be40d87043758f84ec4288fe794f15884b4f0e41e

Observation 85ac589d-1f86-4088-b7ac-c1ef40899a16 · outbound

This paper cites Let’s think step by step and output the final answer within “∖boxed{}.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Let’s think step by step and output the final answer within “∖boxed{}

Reference 77

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:01:25.100791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.303152Z digest=sha256:24bcbd8f5bb2597b6247437afac5e69ca2ebc92aba950733bb4df2902eeac19a

Observation 4acc41aa-0735-45ab-9090-ced5baf37279 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.092909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.307054Z digest=sha256:f2cb60c4d995c247259918f56fab2ec91f29d97a7a41cdadc8c306a98a39cc50

Observation 3893a523-8b05-4b9e-b204-8ab6420d5ab1 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.085821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.310240Z digest=sha256:3ed87c55c1494a17c3cc01c23534e445b234616d92a7a99cbeb32e5590117775

Observation 906fa85b-88d0-4375-baeb-cf59e85ddbc5 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.078266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.313777Z digest=sha256:2c5efa0a79d87a4f87add54577be2c92a089fbc8fef57a8f134eb6894f7c7f1f

Observation 765d8801-9e27-4e21-a3eb-90cc2a0225de · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.071266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.317134Z digest=sha256:af34c3ea797af6ed13312dd5222fcfaed14eab2b35d5625931ae18dc4eb78036

Observation ac286313-68a0-422e-8629-f9813c2d82fc · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.064054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.319986Z digest=sha256:8b47076c5727bbaa4c38ab85765fc2003d6a03da6b9aee8c68dfe048b53e1cab

Observation 488ecaf7-548b-4632-aa20-a3822232e1b7 · outbound

This paper cites ### Output Format : - First , provide your brief outline and planning for the question design.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning ### Output Format : - First , provide your brief outline and planning for the question design

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:25.056425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.322640Z digest=sha256:28cf3ffef88ab0f1ad568abc557fbc839f2da0c9ef9f3c710da8c29591104117

Observation 7d10985f-22f1-41e3-89f4-74447f4600e0 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.048156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.325562Z digest=sha256:f7f537aea3d8d6723346e1a26b279a696fc95d31477d2d07b900c5aa9db72f5e

Observation cc6fcc66-609d-4c2d-b844-5f0cc2574456 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.039882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.328306Z digest=sha256:d433227a183c27d83d9a768c828d92d6e7c410295fb4c880d3d0c4016afdda56

Observation 9f239392-f56b-4ff8-8874-8024e6abbd73 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.032590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.331186Z digest=sha256:bc6079eb39c6e4fdc5ef5bec02bbab89d906244c75d866345bbc59e2262068c3

Observation ab9efd30-3f79-4d89-b83c-a3af5ccf71fd · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.025142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.334078Z digest=sha256:1d6569fdd0852565450a8ae13d5fc47c62df60d12d87f49907673a9d80f8e01d

Observation 9b151e98-c755-4475-86d1-5df25b4d329a · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.016741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.336829Z digest=sha256:aea3e5f282705f9f381debbd6c542b590845924aff20a3feb24dd4a1e9bf3d93

Observation 1211b350-b2b1-4cc7-9083-ef9bcb61a25e · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.008134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.339956Z digest=sha256:725fb83a2667a2ef53dccec842cf9a1554474d6507a05f617d6f07e5d59a6448

Observation a7b43e2e-2ffb-437d-9b50-cef7c5b7c573 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:24.999815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.344131Z digest=sha256:9b461b60cc0ef05975334d6c01fc5628b6d3c831f4cae22871cea9a6aa1b473e

Observation 72a69c35-1e4e-4365-87e9-1ef6359e27c4 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:24.991456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.346982Z digest=sha256:9c05a8e1ca28de0826601374221d281dc9d3cd3ff6b629cc61f646b067280079

Observation db2a18fe-e500-49f5-a2c1-614161821f13 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:24.983773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.349775Z digest=sha256:7181acf708af6e40f701fc904cda225940397ef531142b3823ef5e5ccc65f78f

Observation 864def1b-ffd6-49e7-ab28-d8d31f5552df · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:24.974829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.353008Z digest=sha256:985d2078fd089cdccd53387adf3ebda5d574b4a2ff72388cc9aa2eeca6611f48

Observation 944059f0-1116-4255-9878-3e7afa29d930 · outbound

This paper cites an unresolved cited work.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:25.170896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T05:01:24.179317Z digest=sha256:60e4d1de8fb3aacdb4e744945a5992389ce17c7ce21eda18dc090867bed74272

Pith citing papers

Observation 6ce4099e-64e1-4d5f-bb86-20d23d32d5b1 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.690339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.690339Z digest=sha256:6c3f68178112b8a271fc64ce27cfd61d935d5192b9a92a2f53060b4cc711d5a4

Observation aada9b5e-40ca-4420-be05-92b822b7db85 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.776855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:b1c331ca23e18128291e33342c46bc4e4220909a7ce7d8238aea31a8b0f66b95

Observation 33f605ae-eee2-4ce0-8cbf-94d14178471f · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 299

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.823728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:5baec3969cbf7925873ed5cf6dfa62cee37a05b3cdf8a2cb3fb1277208908134

Observation 6776643c-97a0-4199-9eb2-0fd15a3547b6 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.895807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.895807Z digest=sha256:f86778fb3b37244aadf8f8037d75a3b70fbd9658af542ee0e08ebbccf7cd671d

Observation 8fbca871-fffc-476e-ac85-e9ceb31743d2 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.689821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:693c620c75e74d3f64183e0f85e5fe5f60d4548892362b54e98a85e3aeb92a8d

Observation 70ec431f-26c0-44d1-a8bf-90c3a40bb5b1 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 182

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.785339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:ea63a29fb9c15167d1838aa6df1d2df834210d743dde5364c80fd881dc21ac76

Observation 015fd05d-9846-488c-9baa-c6ea3f247377 · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.233413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:a7595bb39f885a4c1cc005eec8a76f23cb9a14c561d277f2b34b596f41f20eba

Observation baf5f158-3f8c-4031-9a2e-2ea31bf95a55 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.987663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:bee999e9b0b567c7f37d89db59b64af014cad4a9704ac409653ed7d3a8a26e2b

Observation 16caeec2-af6f-44dd-a223-a6cb3ee161ec · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.431684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:1cc0eeec1a22189cc33f691f65884dde07480f22e6f93869b310606bb60d737e

Observation 1240ba79-4253-49fc-98b1-5c45025dfec3 · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.126478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:ad5e4d6dd8895a739be43afa540af60f7be74e26e9debd6fe7f824dafb8306d2

Observation 9c0575ca-1c65-4e26-a911-8601810c1a2a · inbound

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier cites this paper.

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:17:33.759692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T18:16:31.176239Z digest=sha256:f295064a5553a73845c33acef7d703251e34b7b290c75d6531e7ac32959149da

Observation 9211432e-5b7f-4b22-82cc-94924bc16437 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:53f4fcdadda78c53ce5a41dd678f623c1c1fef19428d9d072e321aa2876b91e5

Observation 320d73ca-b445-4c57-8b2c-a322c9e2aee3 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:24.739824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:24.739824Z digest=sha256:7d6a6505813d8090f984c470fbd7c7625541eea1f9a58ddb7057115663c8845b