Pith. sign in

Paper Citation Record · LEDGER

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

As of 12 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 21 inbound Pith citation observations for arXiv:2511.07317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.07317 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:08:10.236337Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:31:42.448654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T18:40:03.168966Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 711a9378-d43d-44ae-85ac-9eaee17f53f4 · outbound

This paper cites write newline.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.359917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.359917Z digest=sha256:b4810b29b3e5eea209f5ce47a2dd8033b6f567fd35682e54efbfc05fd67de649

Observation 41507d03-bcbb-42eb-bab7-f069ab16b6ad · outbound

This paper cites Aime problems and solutions.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Aime problems and solutions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.445615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.445615Z digest=sha256:d825ad6b2a66c0f573f5da67f947724f45940b4dc120aa3a95e3064a7ad8786c

Observation 7c049207-1df0-446a-aa0e-d0a5684a3135 · outbound

This paper cites M., Wu, Y., Powell, G., McGrew, B., and Mordatch, I.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments M., Wu, Y., Powell, G., McGrew, B., and Mordatch, I

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.518494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.518494Z digest=sha256:bb8a513bd429a11fc0218197ec71555fe52c462491cb0a30773ff7199be0b1fa

Observation d0f403d2-1bb2-4db0-a725-8eeb63a00929 · outbound

This paper cites Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.572067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.572067Z digest=sha256:350de738b105f5822797c5e65f6bd73a3499d0f3cb4dc2ce77f74a137a2350be

Observation b04e6ac3-fb27-4529-a655-615ad5de2d78 · outbound

This paper cites Self-evolving curriculum for llm reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Self-evolving curriculum for llm reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.630272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.630272Z digest=sha256:4fbccd7a26d205023ecbeb4c07934162a0600d0acdd66cc872be9fb174dca8dd

Observation c29c4463-2279-4cd2-901a-de145fdfcccd · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Leveraging procedural generation to benchmark reinforcement learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.748584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.748584Z digest=sha256:81d6931f521ee367cc26e205979d93125f8e72ef73a4d06e99ff87fc800bc781

Observation 9535e49d-2159-455a-a9de-101bbf0ad069 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.802059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.802059Z digest=sha256:ad5ed4a00290fbd79363afd709502ba7b7766fa05d2cbe4a06bd107c01732a5e

Observation c0ebb816-2aed-4d66-8ac4-8c34f78ff225 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Process Reinforcement through Implicit Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.849138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.849138Z digest=sha256:1ac524f847a8ab5e1e4d49679d072f65c8b1a2630850d9870b26b260d6e7be1e

Observation b4f9eeb1-52fa-43cf-84bf-786875497883 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:04.933968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:04.933968Z digest=sha256:a9c0b0b92f5c8981896dac2eba8f21bef880312ec3ac61356f188e89f0d93c42

Observation 83c99156-8ba9-4c91-aa52-838e1cac10e4 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.021984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.021984Z digest=sha256:f0b893f0ce04a4dddac1355392bddc5b11ea5421114d27430f97c64020b7121f

Observation dd6500fe-364f-47cb-a18f-034c9a757fcd · outbound

This paper cites Y., and Tan, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Y., and Tan, L

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.083453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.083453Z digest=sha256:0d80728f193e8e9ef97c9a53f7c23a74096c172c9e2e3872ef724cc4510da853

Observation b88fc6d6-9a7c-4967-99de-c508fd59c761 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.189419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.189419Z digest=sha256:f30fc5e0ffa444a17f295c5f62db8e35d35d6e6820a49b73ad8c0cc0d090f2d1

Observation 6241cec8-4ab9-4249-ac86-d3ee2cc9e701 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.247460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.247460Z digest=sha256:2fa80f4a7c8e90850cd6ae9c572dabb5b52940fa24e5836d8372d67bcc38aeea

Observation bfe257a2-376d-4514-a611-19c0b2631ef3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.337250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.337250Z digest=sha256:6f89d58fd130dcf00605f5dc57b7ad52d29f06b503fd0b32c3e7c0c2fc1422a9

Observation 63845372-c90a-463d-83d9-a1ae4dedb9cb · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OpenThoughts: Data Recipes for Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.394862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.394862Z digest=sha256:7d8b42159994d5cdeaf9497f31e4462c85081cd8449762a28be1638a7e0ce7c5

Observation c0722e1e-dee5-4d8b-9c9c-7859f02264a0 · outbound

This paper cites L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.472364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.472364Z digest=sha256:36acad459daf49c6b2f576d5170d95cd804af75781e691c11fc7f039c5db0fbf

Observation 7eff8bab-4ce0-4f8b-b813-c430af938f21 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.560042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.560042Z digest=sha256:779f8e8600cb38bf0e50dc02fcf907d91d102d92d0d6eb530872e6a493a8d524

Observation 5ba791f7-8638-45e5-ba25-680772fc0b65 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.673978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.673978Z digest=sha256:f7d909d2b14d4a982b82a128ba461a1b8aad9763ae22eda71f62bf3a04d6ed05

Observation 42f9225f-c8f4-4bc3-837a-6100450466b2 · outbound

This paper cites Prorl v2: Prolonged training validates rl scaling laws.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Prorl v2: Prolonged training validates rl scaling laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.744103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.744103Z digest=sha256:daf3c2f7a0d8a713d3caea4b17aedc7cf160035524e77a430850b09dbb0f4bfa

Observation 54ca0bb3-93fc-4c14-aab0-4a2359dd51bb · outbound

This paper cites Brorl: Scaling reinforcement learning via broadened exploration.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Brorl: Scaling reinforcement learning via broadened exploration

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.823057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.823057Z digest=sha256:0413a585f7a0dc85c81bd08a7508d676c2d70115c8c55ff183c0f49b23e1072f

Observation 38fba38e-0463-4c0c-82b2-238b28506783 · outbound

This paper cites Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.895267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.895267Z digest=sha256:71aeb5ea6572084b483bca6176356b1c22a4b12436a6158ecae4010cad2d7cba

Observation 68b9e502-9711-4063-87ce-6832ab832e08 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:05.978760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:05.978760Z digest=sha256:01d77a552ff82c62803d210f95f75d55dc96809561c88cdb04377b25a9282e07

Observation b1b10ddd-5362-49f6-9954-0259c47ac4f9 · outbound

This paper cites Prioritized level replay.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Prioritized level replay

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.061589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.061589Z digest=sha256:8ff75e228d26883db5e3b4b8b4d39fe94369c30d12d26fee03407c7892ac050a

Observation 59a2cb9f-857c-4e43-8d08-a06b88f9b181 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.138272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.138272Z digest=sha256:35425a4639eb1ed5ff8fe05c30ce738764dd160bc4d4b620450a7863f250a3f3

Observation c9b3361c-8da1-4f76-9532-371b05ce2414 · outbound

This paper cites V., Jain, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments V., Jain, L

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.221024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.221024Z digest=sha256:adf820dec2bbfc52a7df99b64eecdb023c9846686b59bcf37c3199d94066bb81

Observation 7132b63c-3286-4b58-8493-a9a9e8218e2f · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.327269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.327269Z digest=sha256:361782ab7a623454be9c053e60e29b481b727905331131dbb4eaad6bc27a1f5b

Observation b1bbae63-0a3f-429b-b99a-05baa91c1571 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Kimi K2: Open Agentic Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.409104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.409104Z digest=sha256:af79e6fd75f9c47578a8f26469764b4c1a00010d2701d6ebbf3ad57da95fa5a7

Observation b0dc3d5f-5cf6-4b9b-ba7f-6fd593e0749e · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.494187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.494187Z digest=sha256:f241828fd3732b2507939658e6a23ba598203c0e08c3dd46f87bc17f70971274

Observation 84a955eb-c0bc-432a-89c3-a50398bfaba5 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.605289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.605289Z digest=sha256:b44d14d0f319cbeb8f653c2a442235be230d397a6a26ce80e5904fc07bec5b74

Observation 8b87b70c-90e3-459e-ad40-839f969c78e5 · outbound

This paper cites The Need for a Big World Simulator: A Scientific Challenge for Continual Learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments The Need for a Big World Simulator: A Scientific Challenge for Continual Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.689110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.689110Z digest=sha256:bb348734a31b81a0b92c9ee99c1c3fd28efda1e19c9f970066fbc3af5083bf5f

Observation 0c65452e-c59e-4c8d-a26f-34f68ec3cdec · outbound

This paper cites D., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments D., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.771218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.771218Z digest=sha256:55e2161edc2c49b22d8bd24b5dce837afdc876f793c3d94c2a67fa24e1b727c8

Observation 0c8e6dbc-9ba9-4cfc-b8db-a4b698fc7381 · outbound

This paper cites InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.836334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.836334Z digest=sha256:8f85f37cd669a640d071e4f923cd5897ce8466642512a10244346fc4f7006f8f

Observation ca4b7185-7809-4763-8279-75bd5f975fb0 · outbound

This paper cites S., and Jaques, N.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments S., and Jaques, N

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:06.902982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:06.902982Z digest=sha256:67ca7aebd70a73499d37853f09d87d77ab715d1379d515aefff870c4a697b987

Observation eb63c31e-e793-413c-b0f4-ef3d588965c7 · outbound

This paper cites an unresolved cited work.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.071246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.071246Z digest=sha256:5b79b5dc9d6acb37177234425185f1cf7a1473ad9491bff76cb8c9b15bf359a8

Observation 75ec6c54-8332-439a-acf0-24e5e12fed8c · outbound

This paper cites Saturn: Sat-based reinforcement learning to unleash language model reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Saturn: Sat-based reinforcement learning to unleash language model reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.185149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.185149Z digest=sha256:285740763634e05e4c883f28ec46ceff5397c34f1ca4967c1da90cffd2839b05

Observation e2683b59-d82d-4645-8929-4126cc34d343 · outbound

This paper cites Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.393807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.393807Z digest=sha256:3fe8bafdc6d7c27bb2775ff427f0be90e19222becb693d032d989cade0b949d5

Observation f8b28bf1-c0ac-4f36-805a-8d86fc864981 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.481047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.481047Z digest=sha256:d72c99878c672f3f770409201e1927ef6d211abf61c2d9fa36c78ef99b74f4a3

Observation bb1abdb8-287c-4a71-9ec2-a779e459fed4 · outbound

This paper cites Y., Roongta, M., Cai, C., Luo, J., Li, L.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Y., Roongta, M., Cai, C., Luo, J., Li, L

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.569271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.569271Z digest=sha256:37d81f3abb563d2fa18dd0e77186856ac24e3ca2bd891b7a690ec0f0660626de

Observation 9f0efe4a-2ac9-4b6e-b7b6-fc352bbd9358 · outbound

This paper cites OpenAI o1 System Card.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OpenAI o1 System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.683868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.683868Z digest=sha256:37245c770d877c7d4d292da0def8a4c8ae13e7b849d9f167c68f6c4bc486829d

Observation 500fd292-5a3c-466b-b6e0-e7e7133f8384 · outbound

This paper cites Deep research system card.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Deep research system card

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.806197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.806197Z digest=sha256:726aa17f8f3f9ab0d0aa80ac3150c97e0ebd61194a7282ea5370bbd83813138f

Observation a58544e5-1854-424d-9ed9-b49781722c25 · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:07.929035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:07.929035Z digest=sha256:555cbdaa4f08b875bf2d3a78db6e5cffee8b7148390c86de7bec0841c1c664a8

Observation 12d60c90-3bd8-4418-8b28-8628cd7c5e9d · outbound

This paper cites Tinyzero.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Tinyzero

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.053532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.053532Z digest=sha256:c9341d1cdf16ff924a4267b1543d5b8cd330b9aeda3d432f8650bb2abeaac218

Observation 76dc2670-35e2-4c5b-b56d-266d55562eb5 · outbound

This paper cites Automatic curriculum learning for deep rl: A short survey.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Automatic curriculum learning for deep rl: A short survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.135012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.135012Z digest=sha256:47307f4b6403f17d1c77cd4075badf2a778a1ac523b8daa3fb4b95bdc1d27de2

Observation 181a8df9-d81f-457a-96c0-fecd3eda50ff · outbound

This paper cites Qwen2.5 Technical Report.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.255858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.255858Z digest=sha256:088faf8cd84e692001ca124557045225082cef9e8afac6083ca4d9b9b45daf15

Observation dfdb21fd-fd72-49dc-99d0-2698d935ad28 · outbound

This paper cites Qwen3 Technical Report.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.418042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.418042Z digest=sha256:49c7d7cc5befc5023965978612ab9c17811103692f3100c95358644b5a848764

Observation 34c31c20-9cfa-44c2-b1d6-5420b9cae465 · outbound

This paper cites M., and Littwin, E.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments M., and Littwin, E

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.537091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.537091Z digest=sha256:4a4c14be31d2f92453cc2cbe4d0afb1c49a0b59a67a60a451926697a8536147b

Observation 5608e14a-e46e-4d74-8c30-9b628f50b09c · outbound

This paper cites D., and Arora, S.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments D., and Arora, S

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.623207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.623207Z digest=sha256:60fc920a596b86cd7601642d18ad1afdd60241d4da834dcb0695034dd9b6dee8

Observation 40eaaeff-fdc9-462d-ad16-c0d6842b87f3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.747991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.747991Z digest=sha256:18e8e80cf2b2b7e90553b50fecac31bf88624e1dd78cfa0a35f7210e1b82a104

Observation 315a5eb3-8e26-4444-9d11-3368db7147eb · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.872364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.872364Z digest=sha256:cacfaa01781853a963299f219d942441099d19fbffdb1eb36546aeb20a79834c

Observation ece12c58-8a02-44be-a5bd-3f495b3650a7 · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:08.957659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:08.957659Z digest=sha256:8736d0302a763dcfa581a0cf84e50feb303e6c328f5c0f55be3542d3aa92bef5

Observation cde4d9cd-cad7-49f7-9f5a-6b546170dfe9 · outbound

This paper cites A., Zettlemoyer, L., and Yu, T.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments A., Zettlemoyer, L., and Yu, T

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.054674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.054674Z digest=sha256:7ff08cd3ab60bfeb215f50124eee30a6671f793cea370709be0ddd3bdd89436e

Observation ff53744e-1a34-46c0-a3cf-94671faf3c08 · outbound

This paper cites OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.137256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.137256Z digest=sha256:3873d868d1887bcec9296645204edb2d1bdc2815f8aa09060d791c0009c8a5f0

Observation a106d9d0-8c57-4f76-89e9-cca574c655b8 · outbound

This paper cites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.253737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.253737Z digest=sha256:3571c47ff90fb08b2a2489588554f34d9071d975d420d813b0145bf68ba73ff0

Observation 1cd2ef87-67bb-48cf-9a7e-6e047fbd4d9b · outbound

This paper cites S., Arunkumar, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments S., Arunkumar, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.362784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.362784Z digest=sha256:e0751c2fb553d6bd0c798c30d0a535332036c313c3d1af9a30cf255dfe16bedd

Observation 0f352196-bd35-47c5-9741-09d84abd32fc · outbound

This paper cites A., Khashabi, D., and Hajishirzi, H.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments A., Khashabi, D., and Hajishirzi, H

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.480977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.480977Z digest=sha256:8ae830ae8290010b518e2be55886df2f9b4cea3139dc9ae0047902713484680f

Observation e2b3afbe-65f8-42f1-941a-0635d7f30ff1 · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments On Memorization of Large Language Models in Logical Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.608509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.608509Z digest=sha256:6028a814fbc084216b2e43e358814b0c0d6396b00a5f3be401cb9d2ac94fd45e

Observation fc43979e-56b5-47c1-857e-606e250eaff7 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Your efficient rl framework secretly brings you off-policy rl training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.724972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.724972Z digest=sha256:841e6f5dbb63af95e6d2829292f4ae6f84dc8822088c6cb850c213b805adede2

Observation 32b4327d-65f7-4ff6-a194-9c1365e4c8fc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.847499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.847499Z digest=sha256:2d514d579414354c88dfb9543c47551510f3ddf982296f90beb2e21b4caa022d

Observation c1794073-6282-49e4-814f-6ae36e646daf · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:09.963111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:09.963111Z digest=sha256:de2509ef3a3e3fa928f0bff9aae6aadc5d598a6db17fb4517ea2594902df9f7d

Observation 8cd0af95-2e55-49a1-8e49-7722f2dc7d11 · outbound

This paper cites Absolute zero: Reinforced self-play reasoning with zero data.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Absolute zero: Reinforced self-play reasoning with zero data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:10.022415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:10.022415Z digest=sha256:0c0751e1e9b9ea3cc6dca40090a56448b4a609db43b9b84076cd057bf10612f1

Observation ce03bf9c-b9ee-4b57-b1a2-cd0e63d90554 · outbound

This paper cites H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:10.134897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:10.134897Z digest=sha256:0da3faa60ea4b12aeb18eae512b7ee0fa3801549bfaade398475065ecfb16a43

Observation 316f22f8-a7ad-42db-9f99-f2ee44fe808d · outbound

This paper cites APRIL: active partial rollouts in reinforcement learning to tame long-tail generation.

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments APRIL: active partial rollouts in reinforcement learning to tame long-tail generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:10.236337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:08:10.236337Z digest=sha256:28d30a08e14c34ebcd704c1ba83de71a4a4fa999ef9f346498e855b436e7bd0d

Pith citing papers

Observation 73b8833c-a07a-46ce-bdce-b308209207d5 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:253513ea28f5378dccdcdab40efb213d822c1bb905a96abc7dac72d44e8c7078

Observation c38355a6-2f37-4df5-b1ea-47e7e3e27f18 · inbound

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration? cites this paper.

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration? RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T19:12:23.254065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:12:23.254065Z digest=sha256:cb3a045c7c064dbf5d2e1d94476776979802711c98aa0788bad7a7c81b9931e9

Observation 1a62abc8-ae43-41e0-a18b-09106001c2f1 · inbound

Gym-V: A Unified Vision Environment System for Agentic Vision Research cites this paper.

Gym-V: A Unified Vision Environment System for Agentic Vision Research RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T10:05:09.049846Z digest=sha256:3e3b667d90772afe5e7d4baebccdfe464ad8528657adf210a8637a1b9326ed43

Observation 54ad51cf-ae04-47f6-bc8b-ee32a8f4f506 · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T15:08:53.731480Z digest=sha256:416e1a2040ff4e8edf2ff8815f8eb431871f12ed453168a26548b26461c1495b

Observation d762091a-30e5-4ab7-b095-96ba5a8a2c6e · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:55:12.135512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T00:48:54.797750Z digest=sha256:c139e10a0d32840e5b992672018e788680d1bfaebbf3489e81dc063fa6df2d73

Observation 5141a38d-f1c6-4c9f-bf84-84ca7f3dad6a · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 222

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:7cf80a13a075cda5ec998f1522c329e5543bc962a77ac9d884108d00fe5e0a22

Observation 34638b7f-9db4-4b4f-a216-9f243aa5a349 · inbound

ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes cites this paper.

ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenes RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.704541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:19:04.708277Z digest=sha256:f9e14d8dd63f011088aafd5f814daf527e66060630882a03a8c17a2e56f012c4

Observation a31563ec-0660-463c-bd50-6a643bd37c69 · inbound

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis cites this paper.

Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.654889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T20:58:07.921910Z digest=sha256:d7ae271212e3900505b614b8c98aba6c32f9ab77de438ff60225c11f939b4988

Observation def3f0a8-eae3-476a-8705-e4276688909e · inbound

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL cites this paper.

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:19.981261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T14:55:07.045625Z digest=sha256:e9bc0a57b06209fd11dbf3d087f813de1da6089ba25da60914e9fa9816b0440d

Observation 1f8654e5-305c-46fe-9781-a682b495bb4e · inbound

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning cites this paper.

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:29.080522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T10:30:42.057301Z digest=sha256:1f7310beb6fb6971f106313e7f508b5481fa07e747f319ced78d0d236d533672

Observation c9047f27-7620-46b8-acb0-88b701d01ba7 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 263

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:40:03.170224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:8e97aefa87ca5aa6de84a32ed1f81e3c4f426a000dbab6527209b53f890cd8b6

Observation 9fafa30a-1157-4851-84e3-bfeb04c14e04 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 263

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:15:59.080655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:6e1982ec5a30a87743d9d5e39abca89388769d73fa0cd7ffcfd68ba4a343126d

Observation b9405cd7-b621-4d4b-aa7d-e2fa6c9684ed · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:53:26.544967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:ab80b8ba35463c565910e62f6aa55a488dcbdafefb4b8d21ab1eb8fc6e07125e

Observation 55d954bc-7085-45db-83f6-904ce94409e6 · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:59:51.886210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:2702f09851483d85c80dfb8273d60cc66afd7fc672faa370e02fc2cb39d6214a

Observation 88541c9e-7c6b-43db-ac95-2fa12382a28d · inbound

SETA: Scaling Environments for Terminal Agents cites this paper.

SETA: Scaling Environments for Terminal Agents RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:39.248762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:39.248762Z digest=sha256:fe0695d717dff0d362b7ec9db4cb1b9503412b28935ed1efa2daf25c023c5d72

Observation 7d008aa5-dedb-4519-bd8b-fe0a4c2c5431 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:10.993520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:10.993520Z digest=sha256:f70ea33029cb15804e6fa09a78d631f5ef3ca24855fb9d694eaa118e8ac890c0

Observation 03046ac1-6a00-4528-a8a9-70b08d86f949 · inbound

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning cites this paper.

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:49:38.226910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:49:38.226910Z digest=sha256:74082d216a6a67ba7d1be33edd2f18463ce2dfd1bcc5369b368333b0f95616bc

Observation 79dbd851-04f6-4497-971d-815a2b4a0b7b · inbound

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning cites this paper.

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:49:50.113959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:49:50.113959Z digest=sha256:9b42cae2b878c1b2422fdc53a8d5b874463fb6a538f94aecfe4d84b237965bb5

Observation 71867253-54ed-464c-97bb-90248c9e3e5a · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:24.989053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:24.989053Z digest=sha256:66df38b52b3dc6b88796aac0eb5eb974ec8a6f5f1498c830b4bc9fc4fcd09cac

Observation 086199cf-c30f-4789-adfc-39e11916fa3d · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:35:01.545456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:35:01.545456Z digest=sha256:ac0f4e3b73e5fc4fa0b3d576d038f812582b7da512223b393a7128a0e1ee79c3

Observation 2c808672-be58-4cfc-8581-9cf6198aab35 · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:42.448654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:42.448654Z digest=sha256:5a49e023361fbb53ddf96b87e04fecb6f0e07f530a0305b04ece91570e5cf1ec