Pith. sign in

Paper Citation Record · LEDGER

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 10 inbound Pith citation observations for arXiv:2507.17512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17512 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:04.687922Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:30:37.345453Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:30:01.593067Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8540c431-317c-4b6c-a656-283e69872171 · outbound

This paper cites Program synthesis with large language models, 2021.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Program synthesis with large language models, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.403637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.403637Z digest=sha256:40a345260cc635d8139342318b3192331f74a1485252dfadb28ee60c50db7715

Observation 464c56b1-320b-4356-b72a-904b3294ae01 · outbound

This paper cites The logic puzzle baron dataset.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning The logic puzzle baron dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:06.233755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.410008Z digest=sha256:ff9d31a14112c622404af2faf4f1cd89a31c609885ff69f87a9791906f5139c3

Observation 36c4b4ae-a02e-4d87-8706-ad3c159a4e66 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.416178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.416178Z digest=sha256:f40a668a39293b85a7ee7e3225cedb9b3b9df81c442123d380bbe97ce3b03150

Observation aec1e64f-9f69-469d-95c6-bdf1145e592d · outbound

This paper cites an unresolved cited work.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.422667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.422667Z digest=sha256:1600adb9da4727d034cfeff607bca740edb817e22ef0d76dd9bca28d4316111f

Observation 50bab17c-3824-428f-bc18-762d56417cda · outbound

This paper cites Self-evolving curriculum for llm reasoning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Self-evolving curriculum for llm reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.428206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.428206Z digest=sha256:523a7f486911be94337abc807cea8b3ad54e17a4ea4469577302c3bb85aa5535

Observation b9bdd950-b7f5-4e21-a924-d2cb42fa0ff3 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.433304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.433304Z digest=sha256:baf0d94f0940ef47670eab755be9c653d60890bb9ec5505140fde3937e72a451

Observation db7dce01-0a46-45a0-bb37-0a54dfe30f21 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Opencompass: A universal evaluation platform for foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.440151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.440151Z digest=sha256:e5c6c8020d65fc5707ab56fc5706273330d8ce660c71e1a567ead6db9bbfc5db

Observation 24e56876-24b6-4871-9f3c-a7b0858efbec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.446447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.446447Z digest=sha256:e530981a4fc93a22c95a9002febf95d125c403169adb6c9cc863e9054f8ab223

Observation 5e7aed1c-a2a7-43bd-b6f3-d7e3c7508252 · outbound

This paper cites Does Prompt Formatting Have Any Impact on LLM Performance?.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Does Prompt Formatting Have Any Impact on LLM Performance?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.452317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.452317Z digest=sha256:e8f0457e2350c728401af74217c01e1bdf7aa83dbf48fc9cea55e2238f67d295

Observation b9856231-50db-48a3-b3a6-f705f7c8002f · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Measuring mathematical problem solving with the math dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.459338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.459338Z digest=sha256:ea7d844739e6d2fd229d13f9b9ee12f9244e2fe7f8a14d610392691e4cf4d09d

Observation feed8f47-1a9a-4384-a9c7-ed405b8148db · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.466420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.466420Z digest=sha256:d5d1da2c33b0440b574ab8fe362e5553c92d3fb0c75a7f2432c9093f7c94796c

Observation 51949e84-a341-4780-b585-e1b79be1963d · outbound

This paper cites Curricularface: adaptive curriculum learning loss for deep face recognition.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Curricularface: adaptive curriculum learning loss for deep face recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:06.164700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.472280Z digest=sha256:153876ed84924a74a828eeaca6f853e7f0f2eeb4dde2fe1a535ddcfecf5a49eb

Observation 3b5b79ad-7d6a-4e51-b6ee-a8edd4a7b281 · outbound

This paper cites ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.480460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.480460Z digest=sha256:5269674cf7bdebb80c5088f0fe7b75d2ee6b0046558ee65346914e32770ebb62

Observation b249235c-fc04-4432-860f-86db0789011f · outbound

This paper cites Adaptive curriculum learning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Adaptive curriculum learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:06.142850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.485559Z digest=sha256:37bedd245b3fe02bf6563d657840103147dd97105bbec44f9c6766111c3b2343

Observation 261eacb5-8c18-4f70-be85-b4b4db4152cc · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning TACO: Topics in Algorithmic COde generation dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.490769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.490769Z digest=sha256:6bfa72dd310b5ca1ed64461cbbc2493fad908c9b287556f9634bc6cc080ef418

Observation 5794391c-270b-4b53-994c-3bcafb3f61d8 · outbound

This paper cites CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:53:05.622671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.495978Z digest=sha256:d2ce907076deacc631a0d68c6cc9f297a939f3849c6dc287725b47ab27d8c575

Observation 0c02e523-1b76-405a-a274-961ce3ec2b8b · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.502280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.502280Z digest=sha256:0dacf45dd9e686b207a1663f2040b8a7b7791bf4c2a4b759317aea44993400aa

Observation 0c83de3e-6965-4c84-83dc-a110bbc069b0 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Code-r1: Reproducing r1 for code with reliable rewards

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.507621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.507621Z digest=sha256:6d60880cecb1293d03033557c047198cd18ed14d5235628eded55f8b1d7d0c7f

Observation 3dd56042-bd00-4b39-9500-092af60b8646 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.512694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.512694Z digest=sha256:bbbefbdb94fabc6f8e647f15ea968f82f50cf43d3ece3ea4faff0857b3f575f5

Observation f4a4d1ef-71e3-4ae2-8d71-3a57ec56a047 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.518303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.518303Z digest=sha256:7e04995cc609f64918dc4c887b42fb5efe6149848e1cea23320cfab97f7daf2a

Observation 27c200c6-d8f8-460b-9b1f-604d7cd6228f · outbound

This paper cites Deepscaler: Sur- passing o1-preview with a 1.5b model by scaling rl.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Deepscaler: Sur- passing o1-preview with a 1.5b model by scaling rl

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:05.956790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.524453Z digest=sha256:38ffe37d7824aa4f222843210f583ab02be18fa1b8c25f40010da53206baa2dc

Observation c6d95a42-802e-42a9-9bd9-fdee215ffeaf · outbound

This paper cites Tinyzero.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Tinyzero

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.529918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.529918Z digest=sha256:d13c1c23c566f84bf4ef278ce987261e8316212d49c4e8dff81a0337c6b96fe8

Observation 7a590a8f-5024-4596-8016-5eadc3dde457 · outbound

This paper cites LEMMA: Learning from Errors for MatheMatical Advancement in LLMs.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning LEMMA: Learning from Errors for MatheMatical Advancement in LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:53:05.531253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.535669Z digest=sha256:83c40bd8af27a7ccfad0d609051bf518cdd8128c029e33b8525295f781981db9

Observation b4698c59-ecbb-4409-b14c-54f31f80b53c · outbound

This paper cites REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.542163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.542163Z digest=sha256:2d8dd5f8f4592eb6966c732dfb301be6bc73c553385b447a08c3f6c96c0e8c42

Observation 4600ec2e-47c3-431b-93fc-a06241d32518 · outbound

This paper cites Curriculum reinforcement learning from easy to hard tasks improves llm reasoning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Curriculum reinforcement learning from easy to hard tasks improves llm reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.548107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.548107Z digest=sha256:1da90c5b4c5bb3030d40d9c092039763db51a7eda61c8829ab502b1c228ea085

Observation 53c1ef05-4ca3-4a33-b9c0-7412ce87bb1c · outbound

This paper cites MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.553391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.553391Z digest=sha256:7c75bf56478ff945de90f1cec947c4b9daac62abdea4b0d5282cb4a597cdb511

Observation 077eda26-b9da-431c-8fa9-41f7b564eaea · outbound

This paper cites Qwen2.5 Technical Report.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.559294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.559294Z digest=sha256:35340263de343586b0b3d2658c1451765e1744c885c188cf9b7a8cb499e2e740

Observation 47fa10e6-114b-4b08-b746-7b13c998cd2c · outbound

This paper cites Magistral.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Magistral

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.564942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.564942Z digest=sha256:8dcca1066882600c2ff204fc2766fbf4996989a042e926b21ccb62c827d70e4e

Observation 7957b7bc-3abb-4142-aca0-64e57cae02a1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.570905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.570905Z digest=sha256:4f4388e0312f7a6b8a867038e4a655becd39448df0b95b9f662b49fa0646e916

Observation e594da3a-4ceb-47cd-ba87-bc49fec21a34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.577219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.577219Z digest=sha256:a81ccd7505bd05306f3ed28074a54f07b7ee6b28a915ecadbde0e257addede2c

Observation d22d86bc-084c-4173-b982-0f7168f298a1 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.582727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.582727Z digest=sha256:ebffddb637d2f69cbcc932e399eb7ed736e08d12d92f6e2d72dab5104b80a341

Observation 4997eec7-ee4f-4ac5-9d6e-a3b5f46a16d6 · outbound

This paper cites Template matters: Understanding the role of instruction templates in multimodal language model evaluation and training.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Template matters: Understanding the role of instruction templates in multimodal language model evaluation and training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:05.922540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.589065Z digest=sha256:ed5021665150bca4dfb5d494035e67228565d792ec1f8ed89dc39612aa2a9ead

Observation 94f85e59-22da-4291-a53f-247c92d11336 · outbound

This paper cites Dump: Automated distribution-level curriculum learning for rl-based llm post-training.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Dump: Automated distribution-level curriculum learning for rl-based llm post-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.593968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.593968Z digest=sha256:14def068ed2f932cfb533a7173197ad86b3bc4d22cead53fb27ba3b0dc5f2a95

Observation f940d917-eda9-44d6-8168-e3c54efb1326 · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.599526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.599526Z digest=sha256:0d1781ec329f243e16bacf4070683ac48d3b67ee82fe70da7c544c9bfc8967bb

Observation 69731711-117f-4cc1-9a86-23b321ce36c5 · outbound

This paper cites Rlvr-world: Training world models with reinforcement learning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Rlvr-world: Training world models with reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.605833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.605833Z digest=sha256:31f850db1fd0cbf55610fbf898f6660fd0c5e7399ca13ed16adde3356630ac73

Observation a669a212-d3d2-46c4-8b50-39da3cbb3553 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.611175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.611175Z digest=sha256:77e567d3819a1f92f7b6d7a50991c0f09b11ce4c01efa79a8426bf7e39be8fb6

Observation 6d631812-2c58-4900-962d-551e04168dad · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning On Memorization of Large Language Models in Logical Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.617685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.617685Z digest=sha256:c43b53ca7128e73ba8cc6a2d09a18b33a4e2b1bf2e1b75b73bf3ed6e0775cb83

Observation 4b79c033-5867-4008-8c43-bec4f22f60a7 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.623265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.623265Z digest=sha256:2a96f41a2e6470b152cd7ab0ebfd1b6b32f74f4d5a63f074069d340d7b666bad

Observation b28d8b18-0fdc-4bd4-b0b3-287c237b2b98 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.629639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.629639Z digest=sha256:c0fef32290d576c1ff18a4fed4a7449b225db9730eecac2fcdfeb6af5dd25ec1

Observation 63e26268-c21f-4897-a037-d1ba1f9f364a · outbound

This paper cites Qwen2 Technical Report.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.635611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.635611Z digest=sha256:68adb4994864b12fa2e7f9cf633e656300e4619fa833e7813f7bf4354cca0242

Observation c6b7f1a6-fa47-440d-b4b6-a80fb28074b2 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.641309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.641309Z digest=sha256:13bdcf63b26dd366f88114b52afbb53db0b1d9e804864a077915ccf53d9883d0

Observation 79755e74-e6e5-458f-b5dd-2e89aacd48ca · outbound

This paper cites RLPR: Extrapolating RLVR to General Domains without Verifiers.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.646543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.646543Z digest=sha256:6ed53cefe94b5b30ee213103c71f8a433e7abd658393cddc1aba94bdd7001837

Observation 59e5017b-f635-496d-8b23-47ff0b60d0d2 · outbound

This paper cites Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.652198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.652198Z digest=sha256:9f0328e9dcbd629b52cec0d7e46379fd75c9d35ed870eb03ffe7faa2defd35ca

Observation fe1abd00-ca2c-45df-9a78-b2c81c3541b0 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.657958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.657958Z digest=sha256:45c1a2c2bcc62cc599082b19bffb106b2e36eb26d7fe42a3b987c37baf1186aa

Observation 1d29a49b-588a-4212-814d-3619556a7145 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.664420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.664420Z digest=sha256:0dddf2b7e7f01d54ae4fa4b1150572e09665d9ff6fce13ef2163ff2aad5c8429

Observation 470444a5-3dd0-47fa-9d4b-5ef8075be893 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-06T14:53:04.670074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.670074Z digest=sha256:2078b93fcc8c5b1f6cee2796deab28068068961a5de831c0b5b20941fccca74f

Observation 512ecb7e-3785-48d7-9f99-53c430f3bf9a · outbound

This paper cites an unresolved cited work.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:53:05.904785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.676524Z digest=sha256:dd8b505c89175cdd3940f4becb404667479496204be7403700d2495bd3a5051d

Observation 14022813-be9f-4096-be30-4c9619cce5f5 · outbound

This paper cites an unresolved cited work.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:53:05.887391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.682076Z digest=sha256:2ea855339beb948e4a56a177e8ac31728f63846f50a88a54f584d242320aea71

Observation 95e095bf-ebff-4d3b-96b0-f247b1a9de44 · outbound

This paper cites ## Answer to the Example Puzzle { ”reasoning”: ”Given Clue 1, we know Peter is in House 2.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning ## Answer to the Example Puzzle { ”reasoning”: ”Given Clue 1, we know Peter is in House 2

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:53:05.870440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:53:04.687922Z digest=sha256:f5054476b69c23e1e0d8e508e7630af08b983f358ba558a42df0c92f78f320a2

Pith citing papers

Observation fd3b1ad9-d1ca-4d57-b2a2-a2d6d004ee21 · inbound

G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning cites this paper.

G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:03.334852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:03.334852Z digest=sha256:52ea10aebed71a02767c6799fe44810f24cf9b11c674d80f884f3f4ae325f9cc

Observation 9b320f60-a5de-4d56-91f8-c19802459a26 · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.214745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:7887e478006f54c13d5eb346357bc4012b19c23a4dd5db33fd5f7ae4bd4f1c71

Observation 868a42ff-7788-4660-89ec-71efa5ffe279 · inbound

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs cites this paper.

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:06:00.589898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:14:38.371700Z digest=sha256:98a9b2e05d32796b179ee72c007e04dc45be6786297cda565b46c61d918c176c

Observation a0e06f9f-0451-42f5-839a-a005cbc642e2 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.383704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:2823314ef7f027909f2271a3101f2953a0bf71a0f8faf3878cd361ccc5d1c07f

Observation 580515ac-9341-46d5-9a85-50c4309ae4a5 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:58.333309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:9ca4f51d7ed2ac175af002954936d15047f45c0c100e389264d9a3ea65eddc8b

Observation 5b202722-a2a8-4412-82fa-adf9ed5d7520 · inbound

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR cites this paper.

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:30:01.594743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T22:47:09.330723Z digest=sha256:68d17b3f00447e3d3eac958e462b51d2ff63dd95c518b64682db51df03582113

Observation e3968d58-d029-4d7f-a656-7917cb6cc049 · inbound

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR cites this paper.

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.490774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T09:55:19.804589Z digest=sha256:49517c425bb47fb961f889fcca867634631cdedc32d260932fa28ba9153f2ecc

Observation b0b85819-9a32-450d-990c-f7710a6cf889 · inbound

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning cites this paper.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.286875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.286875Z digest=sha256:cab19c71e9928bffccd4375ee4954193eb96e488f4d5b75b56b1ad45bf61ce07

Observation c2ae80be-b48c-4a30-8693-dbf3e39bd428 · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:25.044392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:25.044392Z digest=sha256:aaa291be4f9648c860dce50eaf84e805d9c1791b141526d2d2bcc579196c2c74

Observation 8025b3f4-d303-4815-81a3-d6d8fd450357 · inbound

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training cites this paper.

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:30:37.345453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:30:37.345453Z digest=sha256:4f2dd640103c5d1b3650b8a52ec6bb57293d9dd2ffd91246aede7cf512c27d20