Pith. sign in

Paper Citation Record · LEDGER

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2608.05987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05987 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:53:26.173725Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82d516c7-7361-4464-8e74-9c0c1715ae86 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:24.992869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:24.992869Z digest=sha256:b0ec64ce92a723767dae1572a0d334063dd441b2054169a5b5d478eacc6af273

Observation 9b876981-78b5-410b-af7a-3a85d061131b · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:24.998800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:24.998800Z digest=sha256:7bd5c940e6702afbe6f531025b4ed8589fea30d53bdc33f055a9af688203f882

Observation 942e9632-ec05-48bf-9c02-77f51e92b41e · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.004243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.004243Z digest=sha256:d3b16394c1534f5161ceadab79998f2add9d4de17bc4b6df0ae61bf3881bdaef

Observation 39cb9bad-9b74-4640-8722-3fa2f343bdf6 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.009238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.009238Z digest=sha256:1edcdfcca82dd46c6921d02465713f69c7d009b545bf50316fc5415c299d7ec7

Observation 667a98ec-614e-47d4-8999-26fb44de5b99 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.034126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.034126Z digest=sha256:ecc7f39c41fdb102fef0d4a1ebeb8287c4a38e8273621f7cf3bec58b30cd764c

Observation 7d3d36e5-cd90-433b-bacb-6f821ec383a5 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:29.145173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.085810Z digest=sha256:bb926903d7a7006e4712abe340792a8ee37bcb1a5a1f09e1bb1767cdb695751c

Observation 38f9310c-887e-4e49-82dc-a9f20ac46ec1 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.115986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.115986Z digest=sha256:26d9dd08880bfc856af96c2d4f3416d795f19a0256c72ac9c3a6216934f2c68d

Observation e7c5d4b0-11bd-4d90-8efc-386bb35b325f · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.149185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.149185Z digest=sha256:cdfd0186915a4fca657f3d7ba7f371186332f090f5202ee48a0836923979ecfb

Observation b38d946d-e347-49ed-9f92-f7c6a045cff9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.176942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.176942Z digest=sha256:3c2ebefd1350927863901b828a5c6c472be6dd894d283bbc93f8732cc9d6c8b8

Observation f6aa2346-0797-4f1c-b7eb-0d3b0690e0f3 · outbound

This paper cites The eleventh international conference on learning representations , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning The eleventh international conference on learning representations , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.189302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.189302Z digest=sha256:14abb4e3ebce2832b68a073c8999d1b07066fa5fb523d4f31d467699278cf976

Observation 46fbc3ad-f81c-4ce0-a59e-7e629cc2bcbf · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.218055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.218055Z digest=sha256:e8512b0a9a4217956328359fc2c477d5d1278d9612cb6a5a1d9a5a724773c12a

Observation 93c6b923-0541-4cbf-8dfc-cf2d80c5f27a · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.273982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.273982Z digest=sha256:bbc3fd6f91b970c4f4eb139125d9d4cda395e01a263474e8ab5b1f3f4c71bfc7

Observation 80ccb4e7-5879-4ca6-8af8-86c1b2414914 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.325014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.325014Z digest=sha256:60af0d17af877ac6346f1dccae98da82bfbf4f9ca75e0b0dc542e8f5e3c9a559

Observation 05607002-a70e-44f7-bb3e-313b32897772 · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.341911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.341911Z digest=sha256:5712cadd85cf42ab7abb5592fcf015f50de3d51b1bbba13edb66b9bf1664a069

Observation 9f802145-6ca0-4ede-bc71-aefee01ef16d · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.354446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.354446Z digest=sha256:00d7a6055f1bac7e0b60021d35b948fcc52cec64e0ea541613953daec6c9af90

Observation baa8d1cd-9425-4610-b82a-e6bbd84bca61 · outbound

This paper cites Proceedings of the 2018 conference on empirical methods in natural language processing , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.360089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.360089Z digest=sha256:f8530ecf4431d0e65e9b04caf54da3d66942ed15137cccca34cf2d3cd45609ba

Observation e5af50b8-21fb-4128-b71b-c30f7aa8743c · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 28th International Conference on Computational Linguistics , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.366573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.366573Z digest=sha256:8c7f815752cf44e99b1e14cc7b67b0577478889c4002e0266de175b726a54f85

Observation 38266335-f39a-499d-9c1c-93407a11ae7a · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.376814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.376814Z digest=sha256:c29e2a6cc694c0a1f649a9738e9b3605018cd1df88b9057980174d3e24eb48b2

Observation 47c397c8-2827-4ea1-97a6-b37411405097 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.391568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.391568Z digest=sha256:1e4d2bcd88333e847f9353797fc9966a9795ae18770d4852fd931b3a9733712c

Observation 59f97e98-7067-418a-9b87-e5456bdee721 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.400799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.400799Z digest=sha256:2a564b5ae71a036cc5f71b092c8e6b666663fe98e47d1a4657d6df6a47cd5504

Observation 0799257c-2fd8-465b-8d62-1b94be6143c5 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.407335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.407335Z digest=sha256:725770cd3093e8fcbf0579e00f07b04c08dbffbce0f053bde733a82dc5c3ed63

Observation e19984f0-4bd1-49d2-a2f6-1f328b391069 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.413689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.413689Z digest=sha256:964ca0e0fb115501391b3f787bf5d1a266365594ceb7c999e3549c94f0940c8b

Observation 19a84209-762e-4208-8078-1552acbeca38 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.418872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.418872Z digest=sha256:8221243f68a89a5fe27ab86438980d1a39dd2eec1a2411964ed266c87718aa7f

Observation ed2c257f-72f6-4040-8532-4305bb8a4b74 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.424814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.424814Z digest=sha256:a3090ff454157a4f7daa224fa7aa2b94f7d07c18f30b5af377056d659d3d556d

Observation 435bca4a-1b8b-4bbf-babb-8c9198a3daf4 · outbound

This paper cites an unresolved cited work.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:53:28.960790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.429457Z digest=sha256:4784a1e2e5ce41065774e364499a181035be4c71e7974270a6990ff79ddce62d

Observation 9b77dde9-97cb-4019-9b5b-082451aeb98a · outbound

This paper cites Qwen3 Technical Report.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.434211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.434211Z digest=sha256:790d7519073c42f6645545da09bf589014a58596c48f6bbe2e4b603c9debfe51

Observation d481ff67-37a9-46e6-9107-65384b2c8bc8 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.438567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.438567Z digest=sha256:e4bc0378c58d7ed01d661c0b392cef819ae97d31afd7b2d78e88a909fcff2ea2

Observation 3763a7ad-38f7-4568-ae6b-fe66c08fd348 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.902064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.444028Z digest=sha256:be088f1354e17bc22984136e479f25fc77b9f92436a256623624c379a2764440

Observation c197bb77-bb0a-48d7-9159-3b7a12768644 · outbound

This paper cites Proceedings of the ACM on Web Conference 2025 , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the ACM on Web Conference 2025 , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.850775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.449543Z digest=sha256:91d5593646e3e09426cff5d2150d2e1b5a2784c1ca33143739fa52f7ec067afb

Observation bbbdf02e-35a0-456f-8d4c-fce5ee880c45 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.454538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.454538Z digest=sha256:a09a4bc03ef9561e785ccf752a5099808f3e31739cd82598fda0a10793d52aff

Observation 4862341c-860a-4872-984e-94d3f753b5e5 · outbound

This paper cites arXiv preprint arXiv:2601.16725 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.16725 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.459799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.459799Z digest=sha256:21b371680d887c35b21b5511815184e13c06c254e948612f0b4e33ee2a878b1b

Observation 3057d07f-715f-4706-a0eb-4d67daca02a5 · outbound

This paper cites GPT-4o System Card.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.464858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.464858Z digest=sha256:6042e133c0cf96815ab22a347d7ba8e8a8b10e04baa60de1da8b9454114d50db

Observation 08ba990a-7737-4c8c-828d-c4f80992a5dd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.470259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.470259Z digest=sha256:ec27a19af673f2b7bec1cb2f8abc5673049879f3711cd62dbcd4f0be903f3ede

Observation 1e755faf-f6ac-40fe-96e7-97639f0b215f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.475258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.475258Z digest=sha256:7482e55c52e36fa035b381b050d7f041147f3ca8ca85f591153f64453bea6e0e

Observation dc2c9a41-6e4e-4169-98b9-6e3b1d2631e2 · outbound

This paper cites Agentic Reinforced Policy Optimization.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Agentic Reinforced Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.480329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.480329Z digest=sha256:8682c5319d143415af2748d952261af18e64a4a0665a413c286f54fd079188af

Observation e08336bb-95f8-45eb-ac97-0cc0c28fcf84 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.484747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.484747Z digest=sha256:b69fb2e2ef6707bbb73cf4635e1eca62984a66aee4833a504ff9d7557bf89a08

Observation 572c8c8f-2052-4491-9a37-6cd4287f1e2d · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.511208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.511208Z digest=sha256:3ae3d483f6b63c80c9ed5cbdc1cdbcd0180502a4c16dc483ca572a625f6e127d

Observation eafcfba5-3be3-4aa8-b6e1-d2ba35632447 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.549779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.549779Z digest=sha256:d14ba4189fc54e1913a593d0f998d9f632f43fbd9936081b69aef013c339b382

Observation 0782cfde-7851-4e2e-ac2c-4037575d4ff3 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.564156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.564156Z digest=sha256:70d3295699317d82c0500965bab21637f11499dfa202a16c40835781c6970687

Observation 8afca331-ee0c-4835-bc1b-a594829a9ecc · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.595205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.595205Z digest=sha256:85ce347efd31d054a6d5186c359282ca8948ce97e40d7c86ebabbffd7318dcb1

Observation 6655eeed-dbc9-4b51-b32e-426f2c580a98 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.619108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.623667Z digest=sha256:9f947a4abf9fe9b9ec3f47051d0f868629b125bd0a676c5b8afeb3caca81c1e5

Observation d506fd67-ed4c-452e-9700-251ccf57922a · outbound

This paper cites 2023 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2023 , eprint=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.647136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.647136Z digest=sha256:242889a7b5facc2810e3fbe8c71604be79c89e47053aab9971698da76f06230f

Observation 0694b948-1b45-429c-b819-511adca85e8b · outbound

This paper cites 2011 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2011 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.671107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.671107Z digest=sha256:63d7328db6f779da80b94ad817cfe34ef6f8765e9de0e2d606ed39fe2ab16f19

Observation 9d9835cf-a47d-46c9-a800-00dddecdce5e · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.719544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.719544Z digest=sha256:db9111dcb293919628b40796b9bef49702daf80eebff7021e3dd51081447d33b

Observation 51d3c44d-c768-4c71-8091-f0dc1852d21d · outbound

This paper cites 2019 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2019 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.481273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.744472Z digest=sha256:f2e2c1261339bce1b85e3c086daec4370a5d06a5e23744029f73252773d8ea02

Observation 26c3af54-464c-4908-9818-21de4473cd16 · outbound

This paper cites Mobile-Agent-v3: Fundamental Agents for GUI Automation.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Mobile-Agent-v3: Fundamental Agents for GUI Automation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.765585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.765585Z digest=sha256:62da500ec54e0580550d08a0f7bf612a2c326c585e21c00ccb04f93e9cfa7a91

Observation 69b7dcf3-89c5-4f58-878c-33c07661a583 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.786093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.786093Z digest=sha256:9645605a8fc99358d144133137660d741174ce127a9ca05c92dac7cb3cb1fc16

Observation d0db4c0b-5dc5-4c92-861a-cdd6372cfdf0 · outbound

This paper cites 2024 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2024 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.791433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.791433Z digest=sha256:831e5bad7233c91bb06d6435032092efe1135145d2aa6094ccf255fa259cb546

Observation eee1988e-068f-45b6-b20e-d4002c8314f8 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.796098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.796098Z digest=sha256:d1dae2a1e3782bd795725f78d9b517f74ed40b3f7d3f14b362af3106f16e202d

Observation 97c813f3-7da2-420d-a63b-8389b4d8abb0 · outbound

This paper cites 2023 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2023 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.800734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.800734Z digest=sha256:d838910d5a40f7b4d270fc9d44db278656300ebddbff82e7a4badbd5f6fb08b4

Observation 53474c2f-3fdb-4541-b82d-b9b86cbf9833 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.805370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.805370Z digest=sha256:b992a07465f1be94bbb323b4b222c95c075604afdb99d0b693a60231fe52252f

Observation 3236c731-949b-4c7e-9279-ebe740505a87 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.317587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.814977Z digest=sha256:85c96fdf7b7e3a2e263cd1954b25d4b55e68f8e1e275c5406f2a61a6b056e9ab

Observation d6c29359-0c29-4e01-a06c-fa5bb5030e1f · outbound

This paper cites arXiv preprint arXiv:2602.03048 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2602.03048 , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.824341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.824341Z digest=sha256:29c7da5daa9feeeb0e9b8a53db27a43b84d6fc907ac3381aada42c2cfa9bc834

Observation c9e66127-70dc-4a95-9552-615c82320dc8 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.237592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.837906Z digest=sha256:fd50744c952d9342ee1b9bf369c5135ff5b5159b886864e4e2fbf9c5e95f2f40

Observation 21728072-1869-488f-9096-e1cef7744cfa · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.206314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.854963Z digest=sha256:a9bd245a7622464f66e796fa4c4d0431c91117d262e506129e830fee109ac43e

Observation 34ea40fb-8d47-4cd8-9d2d-bbc6dac40c63 · outbound

This paper cites 2026 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.190864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.868092Z digest=sha256:b423ecd9d8241da8d86feb96f07d71d0ba7a61bac2f8842faaa892bc198bb419

Observation 9808c1ba-c370-49b1-9be9-5c1f14521c1d · outbound

This paper cites 2017 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2017 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.872410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.872410Z digest=sha256:e06e4d8c1bd0f451eef386baca0388108068782464d7feb52eccec6feb78f4f8

Observation af5e4c2e-d7f3-4cad-b2d3-cda22e9a7338 · outbound

This paper cites 2016 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2016 , eprint=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.165787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.877195Z digest=sha256:0887281e462b531d31fee627e2f9996fae141fba74d2311740b9c37df46c600a

Observation 571504e8-9896-4662-9772-a7c3a091c430 · outbound

This paper cites 2019 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2019 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.149599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.882562Z digest=sha256:65c88c924f42d1b8f99d81693b39aeacc4e13843fc5c1f14570fd453b8ad3f75

Observation a7b45576-74ae-44ce-9fa0-e0d633787833 · outbound

This paper cites 2024 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2024 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.134605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.887615Z digest=sha256:6c2c0bf8d9f3812118f8bf76b08df19d824f181e42dd11c7d15d7752e68889ff

Observation d362eef3-4808-431e-8529-c8335ed30d54 · outbound

This paper cites 2025 , eprint=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning 2025 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.892424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.892424Z digest=sha256:fb1d24fdea9845a2f37083223d0e9de70ab50cc7ebe54d46aeed160b895b3e59

Observation 3bc4c876-61cf-4ab1-a78b-c0ae6c7cb9d3 · outbound

This paper cites Journal of the American Statistical Association , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Journal of the American Statistical Association , volume=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.107894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.896700Z digest=sha256:dd9501b796db6f9b09388d2e277a80498fd1863ff23f030a0ce4e578f61393ec

Observation cbb49228-4cbd-44e1-b1c6-5957ad50e682 · outbound

This paper cites The Annals of Mathematical Statistics , volume=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning The Annals of Mathematical Statistics , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.091321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.901640Z digest=sha256:23ca74128867a9e1346333a9376886abe9939255f72179929b9eb88cf2c79aef

Observation 0c69cbd7-d25b-445e-a4ff-0cc00bd9f76b · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.906690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.906690Z digest=sha256:10cf4f720ccdc5ae08af1385dafd8911d466ec83c5638c2ce1ea999717e897f1

Observation f139449b-87b8-44c3-8130-b16a48d37b81 · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:27.087135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.911804Z digest=sha256:4071d41bb8445887464d1958faf7960d57db6d1f499f80041fc5c791445ac66d

Observation 711881f7-a651-4888-8f97-894bf50f5a8a · outbound

This paper cites SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:27.063688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.916475Z digest=sha256:59cddddaf53003ae8026c4cac81c2808faac7b812a1716095f1c08abec700a88

Observation 2822f740-672d-41cc-8588-6f1dfa8c2170 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Self-Distilled Agentic Reinforcement Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.921007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.921007Z digest=sha256:bb2ffecaa21d6ea796b10e419ec693f2e07290f82d0894fef6aaae139438e785

Observation fe8762e6-3bb4-47ce-b6a6-1854bb9d4d11 · outbound

This paper cites Artificial Intelligence , volume =.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Artificial Intelligence , volume =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.926059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.926059Z digest=sha256:62c02f3e2c8348394a4cacecc006de0a00287d308eb3893eb824e237cbc7c84c

Observation 320f4540-0391-40a4-9ae5-48be7875b92c · outbound

This paper cites Journal of Mathematical Analysis and Applications , volume =.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Journal of Mathematical Analysis and Applications , volume =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:53:28.062099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.930651Z digest=sha256:7e3cd5729cf409804a1e3f7e19fb797a5aecf367853b827a4d497f89b6e89d65

Observation 9a0b8028-fd40-4eb3-8fe3-a26b2292cfc6 · outbound

This paper cites arXiv preprint arXiv:2602.07594 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2602.07594 , year=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.935615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.935615Z digest=sha256:c2b82f7f559a9affe44a26e126518fa62412520e5a8ab300983866aef8976516

Observation fe176c02-ad4a-4efa-b45b-4c78bc7b7dff · outbound

This paper cites Look Before You Leap: Autonomous Exploration for LLM Agents.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Look Before You Leap: Autonomous Exploration for LLM Agents

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:26.818924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.940186Z digest=sha256:4f8d05896c4fa3267d1adf942c53dd8f7c990b40f9efc4c3019a8948621351b3

Observation c41e9e68-7cc8-458d-8f78-953acbafccda · outbound

This paper cites arXiv preprint arXiv:2601.14050 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.14050 , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.959586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.959586Z digest=sha256:3e7d324453a51836ccbd23b20143e5f0daee005c263ad52f1c426822fc3a4ff5

Observation a1c9e453-598e-4be7-979c-de5c7e49473d · outbound

This paper cites Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T19:53:26.507641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T19:53:25.999733Z digest=sha256:27c653da33cd3cf84db41fd81510419d1e8a55172de2c659051355cdbfab4df1

Observation 81cb3325-a230-4c30-89de-3c6bf51209f1 · outbound

This paper cites Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.038845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.038845Z digest=sha256:34ced41750cdef1050f30dc426062d01057b115916b07fe91b7a7ffb3d2c92dd

Observation 2972291e-fd12-47ea-83e2-d0fd5007ea01 · outbound

This paper cites Reducing Tool Hallucination via Reliability Alignment.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Reducing Tool Hallucination via Reliability Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.084902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.084902Z digest=sha256:73ffa8a2b5bbf4184e2c80eb46d2c423ce23e4eddf7d181ee22f88961b8bbb56

Observation d28bd039-8e3e-45fc-a621-9b7d51d947a0 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.141889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.141889Z digest=sha256:44428dec9b5a992a71adff83ce24dac13bea5986d106c96602bc23e64f94b61f

Observation 6f511aa7-b0b0-485b-993e-0790a83398b0 · outbound

This paper cites arXiv preprint arXiv:2509.11543 , year=.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2509.11543 , year=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:26.173725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:26.173725Z digest=sha256:f4ca62924cf45b640a742e9ea5ede35b002155a735f7d310f05d82347e50faaa

Pith citing papers

No inbound Pith citation observations are available.