Pith. sign in

Paper Citation Record · LEDGER

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.21653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21653 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T09:49:31.823521Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 169d5e94-0897-4aaf-ab07-4d181591a69e · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:24.955761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:24.955761Z digest=sha256:93168e298579a259055c112919e4a90e484eb1fc85a86a2a89dc4af00cc21bad

Observation 57e14c27-f090-49f1-956c-2c70226cee1b · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:25.105821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:25.105821Z digest=sha256:4b9fd9930039146142fd7aaa947ee5b7bea1ba7c76e52ff721773e621844c52b

Observation 2733aa50-7cd1-4214-9f4c-d5e308c9d401 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:25.301575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:25.301575Z digest=sha256:e625127cb77d5b27469ae07b440e4a9c0d969bcf29adbed2ed67baa9d0b0a086

Observation 524d3b2b-94cd-4465-b5ae-79b1b6ccea87 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:25.446644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:25.446644Z digest=sha256:a36d3dcebb6b5f2e1fae6cd76762fb3d723589846c778cc3442d9b2302c1eb09

Observation 43cc60cc-2c5e-4c77-b040-5df844775a04 · outbound

This paper cites Laminar: A Scalable Asynchronous.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Laminar: A Scalable Asynchronous

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:25.584160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:25.584160Z digest=sha256:f14a6d18c03a485d187e8480f1c5adc11b6edf555fb78c36b05242b3bf17ac2d

Observation 0ee07cb2-abc2-4351-83b6-0b530f24a4d7 · outbound

This paper cites Stabilizing.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Stabilizing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:25.735099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:25.735099Z digest=sha256:f1db87a7e811e76953bab264a13387cc47651419b03978ba2d10e8fc387cc44b

Observation 1f495a41-c1d5-4289-8078-4685b1cbb706 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:25.921083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:25.921083Z digest=sha256:3faeac77e8ba3fb73f1cc1d914a4bfaaa0ad369b0bf237304f9c43573d57896a

Observation 4e0e083e-fdb5-4295-b431-8844095c498a · outbound

This paper cites Polar: Agentic.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Polar: Agentic

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:26.109608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:26.109608Z digest=sha256:63198fb7f1d63ac6570613fc38efa5c6d3c8eddef1266bcb8c4fbab41cad4efa

Observation f8f51995-d4e6-40da-b849-dd27d25edb63 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:26.337132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:26.337132Z digest=sha256:2eed1c7a681c45bb1c8a860f336905e29492b70fa7f94abeffbcdb3a157a5c0b

Observation 39b9d0b4-3124-48dd-ac64-8e3fbf0632ce · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:26.481742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:26.481742Z digest=sha256:283e42d8be410d7216edb59cd493eba2416561f54fad824b44d3c568cb1a358e

Observation 3a614a4e-908e-4769-a152-dc7129e941e9 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:26.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:26.635910Z digest=sha256:14664e50c3561f75c632f9cbc131732a8fb8b9a4fe7679327441555abf348be4

Observation 8bf6ffa0-fc0e-4a64-97ac-2ced3c95de80 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:26.802812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:26.802812Z digest=sha256:76610e2f56ca4e6a0fd573d2b43d20d633ca8bfb265d77adeab7cc6e813aae63

Observation dd0b2436-ff5a-42a4-8c99-a25afa36678e · outbound

This paper cites and Barrett, Clark and Sheng, Ying , journal =.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning and Barrett, Clark and Sheng, Ying , journal =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:26.983344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:26.983344Z digest=sha256:b3e24fc905d7541f177c687b7709c68c4c84223a07063209594d93d5f032521a

Observation 85bd74cd-14a9-44a5-9adf-f519f40301d5 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.162372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.162372Z digest=sha256:6a90905eff2449a45f5c50db41515bf05cc4d0d2a0c2731b41b7d49603c1e49b

Observation f7dae37a-047d-41f5-a905-123bdaa42063 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.311015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.311015Z digest=sha256:148efe0581741960dfe5e9699150e6fd8e11e8195919630ce25fe7907b01658b

Observation 909a516f-9848-4f18-83a7-0b6d55bde502 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.367566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.367566Z digest=sha256:44a0e386dba3ed398e7465afc6df2b08b12fef16519c9806bc887872d4e2851c

Observation 20da812d-47eb-4afc-8971-3c887eefa3fc · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.517223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.517223Z digest=sha256:dafc78786ddf0914e3650ea529a521a48b27b80342b8a2ddfe848e69ec7a72f0

Observation 1465e5b7-aa28-4ebe-b50e-424149d8b2e2 · outbound

This paper cites and Zhang, Hao and Stoica, Ion , booktitle =.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning and Zhang, Hao and Stoica, Ion , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.648442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.648442Z digest=sha256:46fcdd40563bba9324edd931c5fb6591272b200fb3bcc51ea866a145a02e64f9

Observation ff47368a-5bc5-4128-8c07-6cc64261bfae · outbound

This paper cites and Stoica, Ion , booktitle =.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning and Stoica, Ion , booktitle =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.728162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.728162Z digest=sha256:ea3f734ab722c981d746102d1628a9e88b0611bf43a5527b053013ecc3a9fdea

Observation bd3548b2-1621-412c-8605-c5f3e1d8a8fe · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.796451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.796451Z digest=sha256:ecc543abe7af75981781ede25b1b64f963e12af9e33272beb90fa011e15ef4b1

Observation 7058c24d-2e1a-4947-aa6d-f1628b4c1888 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.862469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.862469Z digest=sha256:9b908fa317ca0110f1f2479c0b5f10b2878e7c1f5f87feff6ed627c5e1ab8563

Observation 3818632c-8277-4616-88f2-2f5aba1d0357 · outbound

This paper cites an unresolved cited work.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:27.943534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:27.943534Z digest=sha256:d7df1ccebc7ec47eb71d045b4504404dfb7c490955d14dac63ac7a30b5c7122f

Observation 95a1e4b1-914b-450a-a949-ef468cdb288e · outbound

This paper cites and Yang, Yuqing , journal =.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning and Yang, Yuqing , journal =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.044364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.044364Z digest=sha256:dc6f49312444d79c1d31e41e1dbf37f525fcfd4abbf15b9abcf1a8f83a4c4270

Observation c3f02915-804a-4d72-99fa-761805eefe7b · outbound

This paper cites When Speed Kills Stability: Demystifying.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning When Speed Kills Stability: Demystifying

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.128160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.128160Z digest=sha256:9f4270131fc8f0d23839286284c58bf1b5d24c105fb17b27d43f50baf2a0955b

Observation 2694c4b7-3202-4cc7-b893-f0fd11245bc9 · outbound

This paper cites Understanding.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.288163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.288163Z digest=sha256:a706ca9f29ab8a53b43dc6bf002a791824d097ec8c7a3406ca37e7e523eee72e

Observation ce0c8919-f800-4303-aa4a-0813e4f916a9 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.368498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.368498Z digest=sha256:343d7cfcaa21ee63348c0ada818b7190886216b73605627db846307a2a717823

Observation 4f74df15-cba2-4cb2-a4c1-20b56a18c131 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.534055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.534055Z digest=sha256:3a809ed12e5de50420886b51959c2cddef47ca6e6c7a79a7d41c94216877fdc7

Observation 8ac76f72-7b09-4b87-ad9b-7fab3beeab8e · outbound

This paper cites OpenAI Gym.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning OpenAI Gym

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.640558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.640558Z digest=sha256:e2dd0592e8995029af5476d0f5bacd03795be08a2ac30c3d931ace5c9e1e8707

Observation 7164737e-8b03-4d42-8053-986528121718 · outbound

This paper cites RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.700382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.700382Z digest=sha256:eac32403af26b28976e903ec80f27777c2da9f071e54f6df7e30272ea75403d4

Observation aa3c69e7-334c-4751-850c-eb974ac2e515 · outbound

This paper cites DeepSeek-V3 Technical Report.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning DeepSeek-V3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.768314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.768314Z digest=sha256:86757a0c1f410955572818f5870351b5da91f56a9e1e97bdc1ba3c16b865d37a

Observation 7f2db113-7f86-4a24-8845-7519a8acf1ce · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.853734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.853734Z digest=sha256:9d6611a3170e23363521d110918b2f0f6e7e885714bd5f94a005bd6c9b2c3f5f

Observation 65851560-8aec-47ff-a848-82a693f9d033 · outbound

This paper cites RollArt : Disaggregated multi-task agentic RL training at scale.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning RollArt : Disaggregated multi-task agentic RL training at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.934558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.934558Z digest=sha256:a62da99c15a42bf80eba9b77a4803ab65af16af7ceb5de869f31ee86ea8dab9d

Observation 4c8d6e8d-30b9-4e5e-916c-53c7ca0360fd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.942046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.942046Z digest=sha256:b845121bc790c3dd3289c52f36f85c740c0193c143103732d3a014d46d29f6d9

Observation 6f80a6e7-f801-412e-bda9-2a853337e2ba · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.947139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.947139Z digest=sha256:21c29b3681da9901b23bacec114555b1ffbf3e1be028c38358194312cad96058

Observation 3a54e877-88d7-4288-a924-63f9b6dfc584 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.999598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.999598Z digest=sha256:ee88dad3bad767c1c157bc84510d83496a323c86cea3c76446ce8c806dbbfb23

Observation 032c9881-1857-46f3-b18f-ef9548bc4cce · outbound

This paper cites DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.089164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.089164Z digest=sha256:bb9ecb895507d3e7d64df9d01c2646ef12815f47efe36cf219be3fcacf7e9aff

Observation 1c2096ac-b5bd-4510-9b9f-88eeaf63604e · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.176487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.176487Z digest=sha256:7a5a777ec663bcd7cd9b703167f7e92c1493b5c3bedfcb2a98bdcf728b15cbb9

Observation f977c3a0-f5bf-425b-ba57-c2fcc1bcfd04 · outbound

This paper cites When speed kills stability: Demystifying RL collapse from the training-inference mismatch.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning When speed kills stability: Demystifying RL collapse from the training-inference mismatch

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.306484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.306484Z digest=sha256:cf195d881fe47bd07dec1c94c94e146b062eb9f52801347d2e30ba8a775d6b2c

Observation 32a94d8c-bfa4-461b-aca9-a886a35ff35b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.455614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.455614Z digest=sha256:14077edecc47d54618b314a4dbf1aebfd867f8b87a565fcc1c050c9092fff8d2

Observation 10f4918a-182a-47d9-bff4-716a6e3ab1ac · outbound

This paper cites Agent Lightning: Train ANY AI Agents with Reinforcement Learning.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.589477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.589477Z digest=sha256:fe4aafa3ae1cca05b03303a512553e5984622b11dbe1ef74d5367eefd4da51c1

Observation 3766b57f-6236-4697-82a1-7c89dd2c0d35 · outbound

This paper cites Stabilizing MoE reinforcement learning by aligning training and inference routers.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Stabilizing MoE reinforcement learning by aligning training and inference routers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.743736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.743736Z digest=sha256:a0efab60312ec6441c105c6caa91c236e9bb01cd632fd0dd9aeb6aa77e7bec5b

Observation 1775cd88-1944-43b8-a6ea-9d0d19736d4c · outbound

This paper cites Jordan, and Ion Stoica.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Jordan, and Ion Stoica

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:29.883445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:29.883445Z digest=sha256:9c68e25bf76787e6e25303f95ac3a205f0d8379651bfa745c4b10157a504d6ff

Observation 53ee4424-189e-4cf7-9f4c-3eb8fe59f823 · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning High-dimensional continuous control using generalized advantage estimation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.016758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.016758Z digest=sha256:52e5d0e53d41aecd36dc24d3ec72350d3ae752af3f38ceb784c09a93a0077aeb

Observation 2f052a68-1122-4f63-bf7c-c808b0a9f3dd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.124552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.124552Z digest=sha256:18feaac63e9971c5b8729c40d6f44df1cac151c961dd7ca76926201205f46271

Observation 53615c33-5cd4-4638-b25f-e2d606bffee1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.274180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.274180Z digest=sha256:65301e5b2599b67bedcee5c646859ba4044052e0d936a47288ddf83e2203fee5

Observation a18c5adc-9c31-416c-bef8-523fd5b64570 · outbound

This paper cites Laminar: A scalable asynchronous RL post-training framework.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Laminar: A scalable asynchronous RL post-training framework

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.428726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.428726Z digest=sha256:7f9901fb34083a775d8e9507d1ce8422e4285ab9a9469eb3ca6116c1dbb9c660

Observation a85b83f0-2982-4164-8d08-03090b49d7d6 · outbound

This paper cites HybridFlow : A flexible and efficient RLHF framework.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning HybridFlow : A flexible and efficient RLHF framework

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.576484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.576484Z digest=sha256:e6c372a83860108b9823b3e0b134b6e393f3d8faafc69093e23605aaa956eee8

Observation 67da6d4a-df23-49f6-a9a6-395678902471 · outbound

This paper cites OSWorld : Benchmarking multimodal agents for open-ended tasks in real computer environments.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning OSWorld : Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.729399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.729399Z digest=sha256:896c6c0931e523c5043db593a1ea89ea108226273c728cff8b451c067ff9d622

Observation 5b21a729-fd86-4cf5-b580-1adf8cc3ca45 · outbound

This paper cites Polar: Agentic RL on Any Harness at Scale.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Polar: Agentic RL on Any Harness at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:30.880046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:30.880046Z digest=sha256:7d219e585621623cfa26f194c35d87d54bade9900e72b117e8782c21b8e2bc5e

Observation 6a869fb7-4c00-430a-973c-afe7325b5047 · outbound

This paper cites Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.023088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.023088Z digest=sha256:a906630120de0e170d9334441e6b120b185634565bd781a6fc62b6377bfe0b35

Observation ae069bff-5684-4e86-9919-8ab9ea6f4ab6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.174986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.174986Z digest=sha256:f7b73e1e5c09d523df1b2b045d114e24306073ebf974b751e96bc12f3b97a40c

Observation dd042aea-7619-4a94-ba34-9e89b795b488 · outbound

This paper cites ProRL agent: Rollout-as-a-service for RL training of multi-turn LLM agents.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning ProRL agent: Rollout-as-a-service for RL training of multi-turn LLM agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.294037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.294037Z digest=sha256:2f73aa36f3d772a6e221a0e5c66415bda69efff111a2441ff1679c782573a033

Observation 7b849eeb-c909-49a0-82db-ab62884f9c63 · outbound

This paper cites Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.413822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.413822Z digest=sha256:9b38f477b619e0245321964a6261125caa1be3cbcc4ec9f2150a81575a2c8046

Observation 265e3e5a-e547-4283-a722-a98add457ae0 · outbound

This paper cites PyTorch FSDP : Experiences on scaling fully sharded data parallel.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning PyTorch FSDP : Experiences on scaling fully sharded data parallel

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.502294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.502294Z digest=sha256:bbd8e7238b41a16cd53396d4b876b87b51e73d1ccaf8430a3897a4fda04b7060

Observation 37981a8a-1c61-49b0-ae8a-b08f35d2efda · outbound

This paper cites Group Sequence Policy Optimization.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Group Sequence Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.626939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.626939Z digest=sha256:b29eeb7d6ce4eafa41145a486d52ca3e12fe92c24f373f165a2d5d98a85ce836

Observation d4c6c9c6-09b0-4fdd-b732-68d274133d64 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning SGLang: Efficient Execution of Structured Language Model Programs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.726786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.726786Z digest=sha256:72be82d058f82a2a7606b64e95db84c85545f8f0b16bf7654cfcf253131523d3

Observation f5568b77-114e-43db-aeff-8c01158870ee · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.823521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.823521Z digest=sha256:5bd5dc2b615096dd187d77c96a0511bb742e18021b80b7550e0f09404477af23

Pith citing papers

No inbound Pith citation observations are available.