Pith. sign in

Paper Citation Record · LEDGER

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

As of 17 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2509.23730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.23730 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:44:15.804741Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62e5f660-7ad1-4646-bbbe-0adcc5b60b7c · outbound

This paper cites Ask the Right Questions: Active Question Reformulation with Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Ask the Right Questions: Active Question Reformulation with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.454741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.454741Z digest=sha256:1dd8864f5135e32d6bde3184480e026c293efb3b488f02c006cd5294cb04b46a

Observation a2f5499a-556a-4cb9-b502-68a2b81a64ec · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.604742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.604742Z digest=sha256:d086510e29b17dbdcf222a871d0458c0bb415785b0ffed55e226611f26c63e07

Observation 25a5375b-e003-47f6-9baa-ed17c4240004 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.895720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.895720Z digest=sha256:604042a809d11f96f3865795c3ba53fd9379675f7ce2c015434c7b5e61f0d0b2

Observation 5584dfa7-3459-46be-a689-c1addf41cc4a · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.374747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.374747Z digest=sha256:4019f037e19d2e5cdcad135cce8c88b2e1ee0d258e1489ae0364a769d311356f

Observation c27b8e1e-b6de-48bb-b66c-bb51be424721 · outbound

This paper cites Learning from Peers in Reasoning Models.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Learning from Peers in Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.634746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.634746Z digest=sha256:7f6ca3c8bbfd63e107a8a1757f67ffcd82f3294a24f290a4d8ee65275ae9f304

Observation 8e1918c7-94ee-40ac-9f42-91e812559d95 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.364737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.364737Z digest=sha256:2a644c2c58e615ca9de98cbf0fc0ef776c928e5c33335baa95f7f5047ee3ce67

Observation 20978142-4479-40fd-9c6a-255df20375c9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.514854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.514854Z digest=sha256:51695b4e14fcc479cddc33ca5d9e01ca199dc9ff44457fd2a8111749cf4b4c46

Observation 94a462c1-b579-4624-9661-c2d3ac48cf7a · outbound

This paper cites Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.684741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.684741Z digest=sha256:7e6e731e4f3e5afb7fcc39cb0ad5fd55e5c0761b4e7a99f110b5300e14df0e73

Observation d43ee754-cc80-4702-8db6-99c0c8b8987a · outbound

This paper cites expert_id.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance expert_id

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.804741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.804741Z digest=sha256:9f26991567b576026058548cc24382fc324f4ef280029586164ab624a132cb67

Observation 1f63d06b-0b17-4988-93e5-828fe488091a · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Kickstarting Deep Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.083396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.083396Z digest=sha256:278a7522a9387324e74c3b77743d81c2f58dab7a03f24aad914bfb6cb391b1d1

Observation 42c1ec05-700d-4cb7-896e-f6bbea0b035f · outbound

This paper cites Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.214742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.214742Z digest=sha256:090c6a3daf74cd0c1a16b2dbe3f9202f05b83beb265684e149b3b7dbee00b490

Observation 85a9474b-8731-4776-94e3-93d4f31e777e · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.199364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.199364Z digest=sha256:e349f3a35cdc06bc3f86a1042fcde3117f91a7ff2638d8724309ff56b43540ac

Observation b595ddec-1e1b-43db-a522-d74a07f306cf · outbound

This paper cites Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.954740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.954740Z digest=sha256:26bcdc8d2ce651bb38904fdfbfea299c4ef1482bfefde60534c7b2a99c682e16

Observation ef26b841-6810-43db-a517-287c105c63db · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Training Language Models to Self-Correct via Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.079997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.079997Z digest=sha256:9c7004afe755062c6f9c324f545f6773c80357c25e0a578960acb055aa59ec45

Observation 88d3d7c3-21f4-4525-8edf-18976e896cfc · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.758213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.758213Z digest=sha256:846dc83ae548b7455483f1b928029986cc5c9e3060ace8bb69743bc8844d326e

Observation 638047fa-cab6-415b-9a0d-e9249938123c · outbound

This paper cites Efficient Active Imitation Learning with Random Network Distillation.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Efficient Active Imitation Learning with Random Network Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.324744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.324744Z digest=sha256:ded7d6e9e0603bf0dafa8e1732b3b7b75df3feb9eb8ecd0a53398e863939335a

Observation 96ef28e7-a304-45d7-97b6-74c9564ba4ed · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.782933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.782933Z digest=sha256:2613aabcc36105891c12800f7732d5325636fd91a28ca0396b8b319d5ddc4981

Pith citing papers

No inbound Pith citation observations are available.