Pith. sign in

Paper Citation Record · LEDGER

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2509.23730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.23730 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:44:15.804741Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62e5f660-7ad1-4646-bbbe-0adcc5b60b7c · outbound

This paper cites Ask the Right Questions: Active Question Reformulation with Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Ask the Right Questions: Active Question Reformulation with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.454741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.454741Z digest=sha256:2a837b26d36f06c0b32cc6aca3aac4443397d11aff575d29213a222765cf8206

Observation a2f5499a-556a-4cb9-b502-68a2b81a64ec · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.604742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.604742Z digest=sha256:ec1ad45fa1fd0c4e4fdcd9692a4f2af93367178b67fe8a6aefee4c27803436d0

Observation 25a5375b-e003-47f6-9baa-ed17c4240004 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.895720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.895720Z digest=sha256:6e7efc68ad2ffd8bfb9251c3479b7727f5a556bae04bf0e45506a62d7b7e594e

Observation 5584dfa7-3459-46be-a689-c1addf41cc4a · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.374747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.374747Z digest=sha256:4a998733aaa95995331ee6005cffbda1346bda294adb36ae03ed0e002dfacdce

Observation c27b8e1e-b6de-48bb-b66c-bb51be424721 · outbound

This paper cites Learning from Peers in Reasoning Models.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Learning from Peers in Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.634746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.634746Z digest=sha256:b8a3e65199e7df9632d6d7b41aef50d79866adc3846bd93d0d8c173da361e18c

Observation 8e1918c7-94ee-40ac-9f42-91e812559d95 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.364737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.364737Z digest=sha256:602773ddec43a1f8bd1106c490a26d2137a0c0220490b1ffc0247842f4d1c6e2

Observation 20978142-4479-40fd-9c6a-255df20375c9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.514854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.514854Z digest=sha256:785185ddd64177cbbc7a28290bf1a6e2c36448dcb228a2225b0419acd55bf8a9

Observation 94a462c1-b579-4624-9661-c2d3ac48cf7a · outbound

This paper cites Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.684741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.684741Z digest=sha256:0357adb3b4a194584d6905e9df1a51ea0fd2759263579fcbc618ce8a61c628da

Observation d43ee754-cc80-4702-8db6-99c0c8b8987a · outbound

This paper cites expert_id.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance expert_id

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.804741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.804741Z digest=sha256:c5d8412f0df8ffc51c7e52e7dbf3801234f1e18eb3f037a5900921fb32f5c54d

Observation 1f63d06b-0b17-4988-93e5-828fe488091a · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Kickstarting Deep Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.083396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.083396Z digest=sha256:abde81fe109afb9ccaf14c2578f19cd96d345fa3b7c57f2baf5f26ce11e718e9

Observation 42c1ec05-700d-4cb7-896e-f6bbea0b035f · outbound

This paper cites Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.214742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.214742Z digest=sha256:838ae8eb72603d60ce0990471886ce4f4369ca35efefc2cdec519747ff716a9e

Observation 85a9474b-8731-4776-94e3-93d4f31e777e · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:15.199364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:15.199364Z digest=sha256:46c80fa49b81c644a74d05ed1a5cc291655848b971a378ab8f7a59d161c2d5b3

Observation b595ddec-1e1b-43db-a522-d74a07f306cf · outbound

This paper cites Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.954740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.954740Z digest=sha256:18a7f72d487a7cdb7f6a8508ef36e352b313b7a275aa0727daaba596996c65eb

Observation ef26b841-6810-43db-a517-287c105c63db · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Training Language Models to Self-Correct via Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.079997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.079997Z digest=sha256:a0b7868081cd18ecc63f99fc234d8f3ee68fa592ffbae141a2597329b2a47c00

Observation 88d3d7c3-21f4-4525-8edf-18976e896cfc · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.758213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.758213Z digest=sha256:f564f45bad194a687d841fc5387eb5d7ebf6114498ab78890e13b8a9dbed2c1a

Observation 638047fa-cab6-415b-9a0d-e9249938123c · outbound

This paper cites Efficient Active Imitation Learning with Random Network Distillation.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Efficient Active Imitation Learning with Random Network Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.324744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.324744Z digest=sha256:19ca634306ca020f2d05da297c8d30f1168d4ba7355321750f1da5d958720d16

Observation 96ef28e7-a304-45d7-97b6-74c9564ba4ed · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:14.782933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:14.782933Z digest=sha256:829978b97b5ac8518dd656ef5ca1ea30c58d89b245179384105cb2aa1e8726c7

Pith citing papers

No inbound Pith citation observations are available.