Pith. sign in

Paper Citation Record · LEDGER

Learning to Reason at the Frontier of Learnability

As of 11 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 2 inbound Pith citation observations for arXiv:2502.12272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12272 v6

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T02:41:21.571824Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:35:11.631063Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:35:19.254627Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact46
  • verified fuzzy39
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f7f7a4c-d5f1-4184-8a17-8f69d175f469 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learning to Reason at the Frontier of Learnability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.069190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:26a750b5374bb314f0a6d4a2c88ec54e951fc83cb254f1a8688195889d55d752

Observation c2cbddbf-206a-4b79-9e27-6848eb81e78e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Learning to Reason at the Frontier of Learnability Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.038058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:89edb224011ffdd1ee295b48b32aedc0758d2d9e601bcc6c2623065a240d1731

Observation fdaf63a6-feec-45ff-9f96-01aa0bad26f3 · outbound

This paper cites Learning to reason with llms.

Learning to Reason at the Frontier of Learnability Learning to reason with llms

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.556981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:5d78fb55f37997937a3920ce8ec1e323e1b060686c698f2ff662526d9ea985c9

Observation b290130e-c07f-4672-9a7d-35b03ff5ed0b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Learning to Reason at the Frontier of Learnability Understanding R1-Zero-Like Training: A Critical Perspective

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.030208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e049a4fa5901d9a97061f252ddaf70f96ca0a88886fc474172b17a461d4e7b93

Observation 7046d60f-b54a-4c5a-a42b-cca48dcecaad · outbound

This paper cites Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment.

Learning to Reason at the Frontier of Learnability Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.553779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a7af7e6aaa780464ee45d10a21a05ebae4681fe2a1c607cb7e80eed37c0a1602

Observation 66fa54c1-421c-4938-8538-0c19b137cf03 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Learning to Reason at the Frontier of Learnability VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.025966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8ff73c905b7e65ec80b5cf65d8293143de686f115f55e2a75e4c86225136bda8

Observation 79f22151-c393-4174-970b-f15eaa667fd7 · outbound

This paper cites Group Robust Preference Optimization in Reward-free RLHF.

Learning to Reason at the Frontier of Learnability Group Robust Preference Optimization in Reward-free RLHF

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.236094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7b18e1b233607e75ddd3633be0dd070ce53f5869dea7e37e744d8c4f12dd3b2a

Observation b2f396c2-77ad-4b05-8ab4-ceb0fcc0b99c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning to Reason at the Frontier of Learnability Proximal Policy Optimization Algorithms

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.231736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b313d755c07887f91da1ffe11771351060abf73f73d36fd666e0f5d7e939abb6

Observation 137dd64d-fe7f-46dd-bbdb-6ecd5ba5c84e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Learning to Reason at the Frontier of Learnability Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.227411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2da36bc4b92a0a96ddf8adf1e67e930f6892a6e4b5bab8d5f53e689850af0197

Observation 42f79281-0270-42b9-9207-899617751327 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Learning to Reason at the Frontier of Learnability Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.081240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c20a44ae37f24d98928118194a07d045131da505a4d52e216ab70f27bd7fd26f

Observation 0c878d87-271b-49f9-a748-107161b0410d · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Learning to Reason at the Frontier of Learnability Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.222795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7555406f84a5cd303fbbc43dcbf7c653be896d304173e8ddfd9b864495f3f561

Observation 95ee720f-f476-4f9e-b569-0f3e85e7c30c · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Learning to Reason at the Frontier of Learnability Rho-1: Not All Tokens Are What You Need

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.218182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2785c57538516fe01401ec8eb354a06760690e86e4c02d61ec2c4072e25b2261

Observation f955939b-2b2a-4736-afb8-5f09842eb4d9 · outbound

This paper cites Qwen2.5 technical report.

Learning to Reason at the Frontier of Learnability Qwen2.5 technical report

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.549825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c966da9253157b9c5771ad9b859bc50f7fb94fdd733d5bc055763888770aac7f

Observation 6a7030aa-b42a-4730-9853-349c3df3b941 · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

Learning to Reason at the Frontier of Learnability MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.213553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:62d37c9a266ff2624d754abd59d2de88b41c0877a9105fce1bf5aaf5886258da

Observation 8393f775-1f3a-43ca-937d-3c5d42a98506 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Learning to Reason at the Frontier of Learnability OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.208213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2b06fc233b4180751c202e5ac7dbf379644e9f5d1d549629cb4873a4d52d2c68

Observation e66f894e-8ce2-45d8-b1a6-37b0f8284d24 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Learning to Reason at the Frontier of Learnability Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.045213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:0d3c02ae11f35db65b095f4006067a181915c466ed26377543bbcde3bf1009a3

Observation f494981d-ec13-4c91-babe-835993ab72ce · outbound

This paper cites Proximal Curriculum for Reinforcement Learning Agents.

Learning to Reason at the Frontier of Learnability Proximal Curriculum for Reinforcement Learning Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.073171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c4a6b623b3adfe8944d32f85dbf3f4e94a580b3e9ffb0420156a73672a54bbc0

Observation ebdc1395-d93e-4c1b-9d96-bd4315a448ef · outbound

This paper cites Automatic goal generation for reinforcement learning agents.

Learning to Reason at the Frontier of Learnability Automatic goal generation for reinforcement learning agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.523015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ad86719dd3b2e850c03fd3cc8379fa19e5c5f78e254219b0c45232eca7a8e931

Observation 1dce26a7-21f9-435d-a4e5-30ef10cb0337 · outbound

This paper cites No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery.

Learning to Reason at the Frontier of Learnability No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.141497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:211a654b7e0fe7a9f71b06735f7620b817e95cc8bd5691a3925d122c62abfa28

Observation 6e35180a-fdf5-4874-95a9-d1a9af850592 · outbound

This paper cites Williams.

Learning to Reason at the Frontier of Learnability Williams

Reference 20

Resolution
metadata mismatch
doi, observed 2026-05-23T02:42:25.514626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b6d5c0ee491fa724d407b6ba443ad03adf41ddb398d80601752a7b80d0c87ce5

Observation dca9b697-bd7c-4ec7-9771-f0cff09a2db2 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Learning to Reason at the Frontier of Learnability Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.179007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:23a3de8d32346f77a90b640e75724e7b9737b055f72ccf6840961c1588b0ac59

Observation aaeb81a3-ee00-4047-b28f-a9f6e50632e0 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Learning to Reason at the Frontier of Learnability OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.089482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f693d48fe51d05562ae18a34186497fee81379b1c9703a303c05d6407deebc68

Observation 10e0113a-ee17-49aa-9fcc-e8b7130c4ad7 · outbound

This paper cites There may not be aha moment in r1-zero-like training — a pilot study.

Learning to Reason at the Frontier of Learnability There may not be aha moment in r1-zero-like training — a pilot study

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.519337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d05746e68a3cc613a3c8e658735fc1c182f765f05d3c6f47ba12c62c2c9258b4

Observation ee5c9610-11d4-4c2e-856c-651230236188 · outbound

This paper cites Numinamath.

Learning to Reason at the Frontier of Learnability Numinamath

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.573705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a7b2b60c2bf956b6c687cab3e224403d997bb9e1be498d851d8c714e2a8d029b

Observation d8c04222-da3e-444d-b7b9-697610f088e4 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Learning to Reason at the Frontier of Learnability Solving Quantitative Reasoning Problems with Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.053039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6089794ffb6cd6d12acc23ecd342195aa24dc6dd018d02ffa93d5465110e7d9b

Observation 23390e2f-76d4-4c61-b3fb-2f047a4552c5 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Teaching Large Language Models to Reason with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.085841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ed87f1c52f29024ab2c5dcd7dcaa2d0c4ff323f63374e0a49c1344f9248af48d

Observation e13af358-449a-4dc1-b110-c6013ee8c5fb · outbound

This paper cites Prioritized Level Replay.

Learning to Reason at the Frontier of Learnability Prioritized Level Replay

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.065432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:9ac9827ccb8531f6bfc503b559e405c2c19f13f145408394cfe9944224c37488

Observation d59497fd-ba5c-400e-83c3-3288b371159f · outbound

This paper cites Learning Montezuma's Revenge from a Single Demonstration.

Learning to Reason at the Frontier of Learnability Learning Montezuma's Revenge from a Single Demonstration

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.057246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2f02067658be3a51e29aafba3f4e2be169a18760b68295c501027369328a876e

Observation b9767396-8989-4e33-86e3-0366a0ea8681 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Learning to Reason at the Frontier of Learnability Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.061849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:4a34d6caa2f69a2e254a1d124c59420df48859955b2377de5264a772e136d52a

Observation 5ba38af3-159d-467c-9c25-3b7f99c5bbf9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learning to Reason at the Frontier of Learnability DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.203304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ef50a6829f7d4ffacc1a271f0db64c56ba7fc2ef9c51a2c4f2bba2230be68c45

Observation a248f80a-a5e1-4951-93b9-13e1d36b7984 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Learning to Reason at the Frontier of Learnability Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.188553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:dc641ee2b351f034b01864bab610f61cc8d0d63932c1ffc4e17fc648610bf6c8

Observation b7293fb5-3be6-4a7c-9258-bd6653cd1ae0 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.183846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1d75c822042427ab58464c185bb7b6c5b22398109ed61573e1152ddde6655959

Observation fccbdeb1-b453-4849-a485-80f6c0eee41e · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Learning to Reason at the Frontier of Learnability The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.511125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1242830ea8e16c7c8b27e42d447ea8f1f581c69e482855caa37d54865e692d1e

Observation 017d8ccd-e4a3-4d4e-a1d6-29eda5c39b62 · outbound

This paper cites Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.

Learning to Reason at the Frontier of Learnability Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.198274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8187e68cc79cee59653f5464ade27166e4b413241ae0d2ebe817a66d323c0dd5

Observation 64e71307-ec78-4c6e-9ae0-e2a5656ffd7d · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.

Learning to Reason at the Frontier of Learnability Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.567335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7bb4f34e963f540dcc2a27fc9c6685d5080f0862d8e564f23556f3bcd64a633e

Observation 2dbf633d-26d8-46a5-8cd3-57a13e2ed0ed · outbound

This paper cites XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX.

Learning to Reason at the Frontier of Learnability XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.048877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:9f58549cff34870dd7ce404f28cbb3794743a1b32bc518fb67854b32e54000da

Observation 8d6e0c89-2840-4822-a6f3-7c0dbb8d85cb · outbound

This paper cites JaxMARL: Multi-Agent RL Environments and Algorithms in JAX.

Learning to Reason at the Frontier of Learnability JaxMARL: Multi-Agent RL Environments and Algorithms in JAX

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:17:02.673620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2333f2364dc2f358aa25d54cd0145af2139a464e7de8e135ad5c2ae3df8d07da

Observation 66e42759-9de4-40d8-8858-fdd92eea52f3 · outbound

This paper cites OpenAI Gym.

Learning to Reason at the Frontier of Learnability OpenAI Gym

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.173933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:5d8ce10dd6f28febf420218d23b2d09b3d9f45b011431a663343a73ff0d06587

Observation 6a88b35b-2f4c-4967-ab08-229c79b51f0a · outbound

This paper cites JAX: composable transformations of Python+NumPy programs.

Learning to Reason at the Frontier of Learnability JAX: composable transformations of Python+NumPy programs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.507646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:36f70e40fbd203b97e6b9cccec0b6c41ad104cd20ef72beb8e9e9f3c4699dd4e

Observation 75b9796e-604a-43ed-a377-36da011eea82 · outbound

This paper cites Measuring short-form factuality in large language models.

Learning to Reason at the Frontier of Learnability Measuring short-form factuality in large language models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.155841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2bb16f087f320f1876e2297e1f720bc9c15904f97d6516b3586d243558174c62

Observation b3d21010-bb44-459f-a500-7b07dff31324 · outbound

This paper cites Evolving Curricula with Regret-Based Environment Design.

Learning to Reason at the Frontier of Learnability Evolving Curricula with Regret-Based Environment Design

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.165070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8eb91976b4e293547b47686b6d0186e8bcb87ab4fc7a1b8b06a222567d69bcc8

Observation e35f324e-bd90-4fd6-bc13-56172bca0ec5 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Learning to Reason at the Frontier of Learnability Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.077279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6b28e2580d4d43603ec9c2ad7e8902a93d1d90da0b4544580e9633d947871bab

Observation ab6cbce9-9fc2-40f7-97e6-c4081d05ce6f · outbound

This paper cites an unresolved cited work.

Learning to Reason at the Frontier of Learnability Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-23T02:47:27.563811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e02166c1bc7ecae64dd90f7ce1f9279530a8e965df673af4dcd0923c2dfc8a27

Observation 3de0a68b-6434-47f1-a568-07ee7e0a358c · outbound

This paper cites Curriculum learning.

Learning to Reason at the Frontier of Learnability Curriculum learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.503948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:83603ed111cc422343dd8097557aa87bc7254a547976f34a50b6248ab7f8566d

Observation e4019767-2d6d-40b9-84de-c43f166aca00 · outbound

This paper cites Learning and development in neural networks: The importance of starting small.

Learning to Reason at the Frontier of Learnability Learning and development in neural networks: The importance of starting small

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.500421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7a42bc1982b33db45649f04875432ee9096fd5e51f356ac1e7af7cb27481cb54

Observation 0d215df8-c324-44a5-adaa-71f935de9a1d · outbound

This paper cites Online batch selection for faster training of neural networks.

Learning to Reason at the Frontier of Learnability Online batch selection for faster training of neural networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.496724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6358cdadc3e0e31cebbe8233e9dc9e6f06aae9fe6f26c81577a4c9e490626573

Observation bda077f8-05ae-47ac-951f-036e3f34c4e6 · outbound

This paper cites Online Batch Selection for Faster Training of Neural Networks.

Learning to Reason at the Frontier of Learnability Online Batch Selection for Faster Training of Neural Networks

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:42:26.116137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:065f0beab089fe2d0d4e198e5bbd2fb22d314407e18807b0b261d49424ef204e

Observation 56f0f289-b948-4ea1-90a1-8512750e7c1d · outbound

This paper cites Ordered SGD: A New Stochastic Optimization Framework for Empirical Risk Minimization.

Learning to Reason at the Frontier of Learnability Ordered SGD: A New Stochastic Optimization Framework for Empirical Risk Minimization

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.193403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6a4fb5cc9b449ee480364ecc5d3f60bf6122b75a1ee050ca11ab167a69ebfb83

Observation 658cc2a6-ce97-4eb2-93cf-fe04e7e4dc14 · outbound

This paper cites Accelerating Deep Learning by Focusing on the Biggest Losers.

Learning to Reason at the Frontier of Learnability Accelerating Deep Learning by Focusing on the Biggest Losers

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.150995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:56b8505b4c4578316e097f6d19806cca566a2464b52fd74c7d9e74948aa8a81d

Observation 8390b0f2-a71e-42e1-9618-87e0c85d49cc · outbound

This paper cites Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks.

Learning to Reason at the Frontier of Learnability Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.169796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:3a39998c6c631eabb867e45236c269a3d3d859d87d068e8a622e74e4315639ca

Observation adf83da6-8a75-4fd5-8bb9-4405c3e7615a · outbound

This paper cites Active learning literature survey.

Learning to Reason at the Frontier of Learnability Active learning literature survey

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.492887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:266c871640bdd2232cf7a14cb532c3e5bc51d51d26d1327010a42a4eef246f6b

Observation afc46c93-d264-4c35-9265-de39385a70cf · outbound

This paper cites Confidence-based active learning.

Learning to Reason at the Frontier of Learnability Confidence-based active learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.489016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:4b932e2c22ce4f4c21167e3c7378a77b8f878a77d0c59145d01d03377d232dc0

Observation 0cd69aea-9ee3-423a-9512-d3105fcdcb0a · outbound

This paper cites Selection via Proxy: Efficient Data Selection for Deep Learning.

Learning to Reason at the Frontier of Learnability Selection via Proxy: Efficient Data Selection for Deep Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.121575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:0fceea5da52a029a16cc9f23ded3e45ca878015df5490cc38cd8ec4e8c98c708

Observation b9113709-88e6-462b-9b33-c67eb55989bb · outbound

This paper cites Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt.

Learning to Reason at the Frontier of Learnability Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.126116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:3a043a6d1721e11fcf6203167cfb4257e5a939d012fb2ebe225b02c4bfcbc416

Observation 898ef982-bd6f-4e4f-b954-3510565f7a36 · outbound

This paper cites An Overview and a Benchmark of Active Learning for Outlier Detection with One-Class Classifiers.

Learning to Reason at the Frontier of Learnability An Overview and a Benchmark of Active Learning for Outlier Detection with One-Class Classifiers

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.111402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:42a491314365576dbae9aab458f9e9783b309e8dd90b7806980a2345671a85a4

Observation f064bdf0-fd28-4c58-9430-a0768496b8c0 · outbound

This paper cites Training deep models faster with robust, approximate importance sampling.

Learning to Reason at the Frontier of Learnability Training deep models faster with robust, approximate importance sampling

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.622517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b634ede117e35001bc0130174e032bfb507d4c6327cd2653c8e180e00ce02d76

Observation 2a4d60cb-1c34-478e-bd5e-6b3a0e26c75d · outbound

This paper cites Not all samples are created equal: Deep learning with importance sampling.

Learning to Reason at the Frontier of Learnability Not all samples are created equal: Deep learning with importance sampling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.618558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:25cddc616b1954a2a1f00e765ea6928f3c6f029fbd5d845eae1d5221298f6146

Observation 341adf48-c1a3-4997-815f-3196b743385b · outbound

This paper cites Self-paced learning for latent variable models.

Learning to Reason at the Frontier of Learnability Self-paced learning for latent variable models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.634546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:fdf7620ea156e1f3e1d2621c33190d5f8174096bf9f12f8e7ce218375a9f54a4

Observation ad35adb3-e42d-4c0e-92c1-9c32b935e279 · outbound

This paper cites Automated Curriculum Learning for Neural Networks.

Learning to Reason at the Frontier of Learnability Automated Curriculum Learning for Neural Networks

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.130839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e9efd309c28e30d9056b81f3b77be23a6ab97d3f20e37160db37250ffe7cc7bc

Observation e480ceff-129f-4204-8390-55f1b6bf2d4f · outbound

This paper cites Teacher-Student Curriculum Learning.

Learning to Reason at the Frontier of Learnability Teacher-Student Curriculum Learning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.097838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:39dd82fd385d29573678e4c21340d91f61c3c705e3841d4f8d4e774f76c1a337

Observation f7a01818-4359-47b6-a735-ded8010df04b · outbound

This paper cites A survey of multi-task deep reinforcement learning.

Learning to Reason at the Frontier of Learnability A survey of multi-task deep reinforcement learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.611853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:0a103a1c1ed7afbdf8f73370740d2e7dd6931cbdd272c0164ede5e5eb4761a9e

Observation 4d81e9e8-2b97-4458-813d-3fd6076f4007 · outbound

This paper cites Automatic curriculum learning through value disagreement.

Learning to Reason at the Frontier of Learnability Automatic curriculum learning through value disagreement

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.604553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:74c9585b010107954bb233c78a5bb3b120e9f381fba7eeeb13774a87c05e5fe7

Observation c3e9b355-ff56-4a18-8a1b-b0961559ae68 · outbound

This paper cites Automatic Curriculum Learning through Value Disagreement.

Learning to Reason at the Frontier of Learnability Automatic Curriculum Learning through Value Disagreement

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.106419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:23a1c2effed07f764ef93a76e22301dd3310ad1a1ec74cbf3b8688ce13aaac1d

Observation 64a38f54-5c17-4084-a444-728db0e96793 · outbound

This paper cites Information-theoretic Task Selection for Meta-Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Information-theoretic Task Selection for Meta-Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.093984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ce28a0e2fb70f1a38de9fb7da469811bbf1dcc4cb038ba12be7026fa0b9aef02

Observation 70dea3b6-cd80-46cf-8035-3d6f2862fc7a · outbound

This paper cites Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.135156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6108188808c244c18266a15d0842bfd966cf82fb66b9e33a63599d643aaf7dae

Observation 1f05f239-4449-48a3-b8f2-60e199228dce · outbound

This paper cites Skew-Fit: State-Covering Self-Supervised Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Skew-Fit: State-Covering Self-Supervised Reinforcement Learning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:42:26.146328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1f95b2833f28e0640c5878518dcb03629f022521f2e76a5b52edc83ff01d9c37

Observation 727e61de-24b1-4793-b80a-fdf8899f22fe · outbound

This paper cites CLIC: Curriculum Learning and Imitation for object Control in non-rewarding environments.

Learning to Reason at the Frontier of Learnability CLIC: Curriculum Learning and Imitation for object Control in non-rewarding environments

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.101950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ac2a6440739082641ab29ba3f1a0e5beab3e0f48c908bd53fb9981c9c3e90010

Observation c34e21d6-43e9-4e82-bc52-0fda7635b0e5 · outbound

This paper cites Goal-GAN: Multimodal Trajectory Prediction Based on Goal Position Estimation.

Learning to Reason at the Frontier of Learnability Goal-GAN: Multimodal Trajectory Prediction Based on Goal Position Estimation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.034216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:55d6902c83f25d29f5769e5eda075497a66676fbdf4ff9f906aa116c8db3509f

Observation 7124d511-8153-4b2c-9e55-ffa1792fb33e · outbound

This paper cites Prioritized Experience Replay.

Learning to Reason at the Frontier of Learnability Prioritized Experience Replay

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.041639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ab232fbf2bb779c898a408b12e1ddf7ed87ca477d8f793f74d976db8b85d9856

Observation cf3a376b-7f5e-449c-bb41-3270b1bba14a · outbound

This paper cites In [48] the authors use the loss from a pre-trained model to estimate the difficulty of new samples for a freshly initialized network learning a new task.

Learning to Reason at the Frontier of Learnability In [48] the authors use the loss from a pre-trained model to estimate the difficulty of new samples for a freshly initialized network learning a new task

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.597336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:36b5e31bd76c6fd62cc2e5287ddd2c531cf2268b24c31431f043d05863f044b0

Observation e4942561-ca24-4a84-8d16-c660bc6881de · outbound

This paper cites LILO can be seen as using return variance—or learnability—as an estimator of entropy or uncertainty.

Learning to Reason at the Frontier of Learnability LILO can be seen as using return variance—or learnability—as an estimator of entropy or uncertainty

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.580531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e3d1fbc3a26f6beae252aeecae8fc657178c368441ac1f12396cae53b49ea9eb

Observation e6b2d2c2-ed7d-4ac1-ad38-a25ab2db34a8 · outbound

This paper cites This allows prioritizing samples that maximize the change in loss—i.e., the model’s learning progress [ 52, 11, 53, 49].

Learning to Reason at the Frontier of Learnability This allows prioritizing samples that maximize the change in loss—i.e., the model’s learning progress [ 52, 11, 53, 49]

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.585243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a57e3c87825a3a01d07ab418cac343dcdabdbb849b4956099acb11fe35cb36cd

Observation 249db8c9-8b4f-4918-ba51-3441497bc233 · outbound

This paper cites Self-paced learning [56] is an early approach that allows the model to determine the pace at which it incorporates harder examples with higher values of U.

Learning to Reason at the Frontier of Learnability Self-paced learning [56] is an early approach that allows the model to determine the pace at which it incorporates harder examples with higher values of U

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.608148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:867a54fba00bb0fa1ec53b0fbce3ef04fa4e56b4f54a1c5f66d3c33cc031211d

Observation c813b16f-bd7c-49aa-a48c-5d1e9b73ad32 · outbound

This paper cites Each bullet point contains a claim and a hyperlink to the section of the paper that proves the claim.

Learning to Reason at the Frontier of Learnability Each bullet point contains a claim and a hyperlink to the section of the paper that proves the claim

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.534779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:91dface5efe6f1a9f5953d503a9695343aa0d98e0a803b7079889c19b55b0bd6

Observation 5eb58b2e-1dfe-408b-a88e-b0a9712c4ceb · outbound

This paper cites Section 7 also contains some limitations.

Learning to Reason at the Frontier of Learnability Section 7 also contains some limitations

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.527386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:90ca7420469733722688d5556b15d48cac1ce1d2a2ad73dc403298832f2bbe96

Observation 78f48f7a-df59-48fa-9422-db9011825e5b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.542651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a48cc18194e6c00b33db722ea32894738150335f3d6c128bc939873f464c0322

Observation 3145bef9-d59f-475a-9721-c673f544bdd3 · outbound

This paper cites The results in 6 were produced using open-source codebases [5] [4] and models, with some small additions of code by us.

Learning to Reason at the Frontier of Learnability The results in 6 were produced using open-source codebases [5] [4] and models, with some small additions of code by us

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.514929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e5c3742a8336573ac96e433ff839136bb2e6272477ec24d8cba7515a023fc759

Observation a8ba68fb-4e0a-47d1-823d-1ead5b3a6ea7 · outbound

This paper cites This is described in Section 5.

Learning to Reason at the Frontier of Learnability This is described in Section 5

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.546614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e98197800325656708ab033504dbb71cb89fec95e0cd86cb2ba54f7812378342

Observation dc7ecaf9-0b14-44d1-9a44-9487efbe03b1 · outbound

This paper cites All the other hyperparameters for training are replicated directly from the VinePPO [ 5] and Oat [4] libraries, and the user is directed to these in Section 5.

Learning to Reason at the Frontier of Learnability All the other hyperparameters for training are replicated directly from the VinePPO [ 5] and Oat [4] libraries, and the user is directed to these in Section 5

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.560504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d5122a18f5057a177571f5c101333de9e1a1c499c3403523c16111938e76f3a9

Observation d9bdbbaf-dec8-49ca-87ef-49ab1f1754ed · outbound

This paper cites We have, however, provided training curves to aid the reader in interpreting the significance of the results.

Learning to Reason at the Frontier of Learnability We have, however, provided training curves to aid the reader in interpreting the significance of the results

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.614972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:94ec67d92d826ed66e0af7cce70257a1a2a1e33de282ad1a4e17f5682f4bb4fa

Observation d8e05a43-969a-4c9a-bbe6-3b08774999e8 · outbound

This paper cites • The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage.

Learning to Reason at the Frontier of Learnability • The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.626385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6a7b6bc2c33e4fd0624f6ae808367a79512fd73821114356bc8b952cb8ecb797

Observation 4937820f-e32d-4279-bc46-91c07d4beddf · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.538684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:69085096cc83a688aee9765eb6a41aec1338df249f0bbc9dcf3d7788428ca50a

Observation a8a36f1f-ba68-4b49-a227-65ae6d3846c6 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.589234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:cccbdff6116aa9a8251d2b1315a9f5ccbcc569631b4694741961403ec315315c

Observation 7e20c96e-748b-44cd-8914-0ed7640fc2a7 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper poses no such risks

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.593649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:046d1ba87db3e68e033694a89af5afb3eb9dd10cc8a62b78b9f51caab6791c51

Observation 6f8bbaa6-8709-40d1-8dab-d2b6437daf55 · outbound

This paper cites The two libraries we used for training (VinePPO and Oat) are both fully open-source.

Learning to Reason at the Frontier of Learnability The two libraries we used for training (VinePPO and Oat) are both fully open-source

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.570800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:01df6eb4803d2af2be48f39db515495f0dd7d0745fd16599aec5d62b5326b150

Observation 990cf32f-c5b3-4c03-90c2-0d0db59eac4a · outbound

This paper cites It very simple, and could be implemented from this paper alone.

Learning to Reason at the Frontier of Learnability It very simple, and could be implemented from this paper alone

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.576905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:43b98f7412747f394f630acc65f6a3a01a6e3d2219f01787aaa440c845272dae

Observation 25126526-67c4-4dd8-9def-b2d074d6cd84 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.600904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7d6347af6827225ed94add1dbbc43efb49e24e75c32ee504661ad15bc9d0e9cc

Observation 8aa1cc51-a145-43a9-813f-4d0b0bb8c35d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.630722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:3100342d2aa644757ae676afb0ffd413dbc72f3a476188e3b91f91e0676d8f5c

Observation a35994fa-622d-45e2-b5de-d565cf4ea1f2 · outbound

This paper cites Answer: [NA] Justification: LLM usage was only used in a standard way for editing.

Learning to Reason at the Frontier of Learnability Answer: [NA] Justification: LLM usage was only used in a standard way for editing

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.531069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:08bef46d3a2bd966757904df01016e4baafbec4dc56181fe025449a37624de66

Pith citing papers

Observation 12e36423-71e4-45c6-ba43-8731c8fc571f · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Learning to Reason at the Frontier of Learnability

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:35:19.260704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T11:35:11.631063Z digest=sha256:87269b7f60031ba9736f858b56b57fadc3ca4149e104e91a2341f68d70c31e29

Observation 2e123e76-4eb5-44c5-8243-1d08c19a53f0 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks Learning to Reason at the Frontier of Learnability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.154486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.154486Z digest=sha256:20ee7c0d87ff58d885314a0d92024281167e6cb51e36210b0d6fa48210e26a03