Pith. sign in

Paper Citation Record · LEDGER

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

As of 19 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 4 inbound Pith citation observations for arXiv:2507.19766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19766 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:11.351063Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:13:50.246705Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:46:26.428426Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c735a32-6238-426a-975b-5fabe3121819 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.543215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.543215Z digest=sha256:5ef2a6593cf1a1c6142d6ab3c0eab0a97891f6a6b9116f29e31520aef36c0686

Observation 2734aa74-a458-4f4d-93c9-0945d4426584 · outbound

This paper cites On-Policy RL with Optimal Reward Baseline.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities On-Policy RL with Optimal Reward Baseline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.581791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.581791Z digest=sha256:89567e0265f5381c87e22b0d68b0a2ff932efcb1a29b6b07e7b29680b5104e31

Observation 0a55ddd6-06f7-4b78-b3f1-ae50f8f80e8f · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Skywork Open Reasoner 1 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.663365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.663365Z digest=sha256:a1699171028f2500182cb7a3aa8bdc56ed32c992bfe64ca5c769b6c24c3ef0fc

Observation 629a8ce9-f93a-499d-aa29-f067afaa48d1 · outbound

This paper cites OpenAI o1 System Card.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities OpenAI o1 System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.716359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.716359Z digest=sha256:094c47f38702d2a95568166a676e4da1aeb28978398682d197591ab369649934

Observation 06fb2d2f-c07b-4e5a-ab10-7bd775961a99 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities YaRN: Efficient Context Window Extension of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.758510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.758510Z digest=sha256:1527c6751439c3da0f8edb10b047ba3796cf1669df41104561b3141a549fcad0

Observation 8f1f2719-f220-49e0-9a92-fc578c72971d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.948538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.948538Z digest=sha256:1829ec1f4b37f341bf640bfe64d6b924dc523cb876e54530ebfc09c9fe0b929d

Observation c388865f-7eac-4003-8b93-e5f5eeacb09b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.053622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.053622Z digest=sha256:f6ec3a090b43405167c32256ba16e5ecd63d44bc55cdfb9fc6eb0619ec767bde

Observation 659d4409-3b4c-430c-9528-b3798a32d09f · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.097353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.097353Z digest=sha256:4e72131ea2948fa256eacc42e384b0f09ca2dc0e7b09cbc6e9eca4bdc2421179

Observation abcb13e0-4ae3-4c20-9f7a-9f70e88cc368 · outbound

This paper cites Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.143364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.143364Z digest=sha256:625d9688c227b3fe958733350f5ed2cf76a416af4635d7fb6cb1cebe0ce2cc6f

Observation cf226d8a-bc79-4441-8796-7c44343c35a5 · outbound

This paper cites Qwen3 Technical Report.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.221814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.221814Z digest=sha256:4c2088e279fb29ab0547e46507102ab60dd83481179b31723ba83b31af9fe2ef

Observation df889235-2c8e-4aa3-b8cc-e72b7019f29d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.266713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.266713Z digest=sha256:c8744d52c4e622c743d8d479a6c87e9191936ec1fa5d7d661a1795d609e07a8e

Observation e386364b-93b7-4e54-9f92-cba37ce75c91 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities The surprising effectiveness of negative reinforcement in llm reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.351063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.351063Z digest=sha256:50d26080b1537bdb2e2c948dc1fc68062c521d7997c4d7c8b7ad84f85def3bab

Observation cab312e2-9caa-465c-be13-b432e97a1a24 · outbound

This paper cites Proximal Policy Optimization Algorithms.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.876183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.876183Z digest=sha256:3ab79f367add0f01d4be2fd2c526f6ee8041ca36ae56c26de93203437476f54a

Observation 748b368d-2259-4f61-9894-70d0e2434516 · outbound

This paper cites an unresolved cited work.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Unresolved cited work

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.904156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.904156Z digest=sha256:20b362db29a304e965bda4d2a242c92e368125a79bd90b4adf4382c5ba8626d8

Observation cdc45cd7-dcb4-45d8-b03f-c4e476396495 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.324859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.324859Z digest=sha256:29266754f00e67dbbca5f9672f690ae5e0c149336f3cb7ffbf03a3b1289b7df3

Observation b4dfdd6c-3709-41fc-b4ef-5d11d911b5c2 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.826129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.826129Z digest=sha256:eb5dcf17f824d590158f4d230f9bef44c681fb47b7dad357bc12dc9bf56b0672

Observation c28ec755-8b39-4d14-95f0-12869927dce3 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities HybridFlow: A Flexible and Efficient RLHF Framework

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:11.025563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:11.025563Z digest=sha256:7cba46807c28a8b34e39bd868f1ba1b601b7fc98f89d13ca26df4f82a7562951

Observation 6e135b1a-7ed4-4bf2-93b1-d2872429679d · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities Reasoning with Exploration: An Entropy Perspective

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.485653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.485653Z digest=sha256:b4b0021a2389c37ccb4a05342f186f33208b31ebc231c920219ced3e2f09e389

Pith citing papers

Observation dfd12a22-b3f9-4d70-9b11-dfb5301856d3 · inbound

Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis cites this paper.

Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:35:31.077918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:35:31.077918Z digest=sha256:4a1cce02bb2c0829e538c11707afa68c4ed1b1b3fffbf782e4504779afdc5092

Observation aa62aa5c-4434-43ff-90ea-d3e69ecc663c · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:46.677712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:46.677712Z digest=sha256:b3603ff4e54cb233b182069efc955558b8a5a578ccfb8d466ab6aa07cacaca32

Observation 3c70c37d-ed93-455f-96a2-7c731d5cbd66 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.431711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:5ede20c4dc7f1f76cc4578cbcf7ab5c83ceff72d45eca7352c025fd96629a206

Observation 785f8f85-1cd2-46a6-b982-d95e14da7afa · inbound

Scheduling Mixed RL Rollouts Beyond Prefix Locality cites this paper.

Scheduling Mixed RL Rollouts Beyond Prefix Locality UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:50.246705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:13:50.246705Z digest=sha256:a2fea1de289f3ce2e108afef58aece95698353f1383113fbcb3ca123f320383f