Pith. sign in

Paper Citation Record · LEDGER

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 5 inbound Pith citation observations for arXiv:2607.14777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14777 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:10:21.309226Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:53:25.906690Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:39:07.928445Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a03c48b1-cd15-4d44-8254-c88cd72f9b0b · outbound

This paper cites Frontiers of Computer Science , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Frontiers of Computer Science , year =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.056279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.056279Z digest=sha256:2c4c48bc42f71beda4ce07849722cbd8689939416f679455a3971198a6826768

Observation c72e6d0e-3cab-46ec-b6ee-40506578a36c · outbound

This paper cites Science China Information Sciences , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Science China Information Sciences , year =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.060602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.060602Z digest=sha256:f7b9f9a5c6d9d43feaa4fb56d6eca0528fb97120fe894d5dd129cc1955ad1a78

Observation 8b2c3d6e-284e-48e1-9e3b-15e6679d8e84 · outbound

This paper cites 2023 , url =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2023 , url =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.064601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.064601Z digest=sha256:c800d67aa8eb0d86270947d64419f7673db71008418f63a01047d6d68759ad20

Observation 152a5d64-f8ea-47f1-9d85-d4258411fbe1 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , year =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.067985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.067985Z digest=sha256:88410952096c60417a01d450aaf7168eda872ddbd24084c255120fcf3dd6e6ca

Observation f6da2a51-3815-47c8-995e-353edaf50a2a · outbound

This paper cites and Zhang, Tianjun and Wang, Xin and Gonzalez, Joseph E.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning and Zhang, Tianjun and Wang, Xin and Gonzalez, Joseph E

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.071848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.071848Z digest=sha256:088c0dfedd113021a44a501872a6497f8805bb30a11d385d21308b6482a11611

Observation 735f7248-9106-4fb7-85d5-96a9e5dba52b · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning AgentBench: Evaluating LLMs as Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.075066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.075066Z digest=sha256:cc92d88abc9687dcc72d6551b2f7775d56d3cc60b58637a422a69744ad43d12a

Observation e3544c73-9a32-41bd-83c1-9a1c3fa04c38 · outbound

This paper cites International Conference on Learning Representations , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , year =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.078656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.078656Z digest=sha256:b4411532d3d226ed5af7358b284ecde9ae1b81bf1e8ebaac2aff7d9184a912d7

Observation 1314b44e-ecc8-456a-8ab7-a645069c1df4 · outbound

This paper cites 2022 , url =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2022 , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.082487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.082487Z digest=sha256:8730a0d43f445ae1c3fc79a47952dd00c19dac2028ab22b21855ebd7ca1e3fae

Observation fadb58b4-bc51-45b9-8d8f-3f92ea4a6d6d · outbound

This paper cites 2023 , url =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2023 , url =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.085614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.085614Z digest=sha256:8b1fea298c103c0992c723fc5bf94b0c942636e506e15a1ab5c0774587d501ab

Observation cdb03de1-827e-423b-a58d-9b64f9ff13b0 · outbound

This paper cites an unresolved cited work.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.089481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.089481Z digest=sha256:707078413fc0e35c4c4f171bff49e00cc744b9ce0d96394abab18c2fb644bfd3

Observation c2a3ee49-38d9-4e0c-be0b-8967527162ab · outbound

This paper cites and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.092942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.092942Z digest=sha256:37da58cd90843348ee0826db915e64a7b1f760390e4a663671b35e999137bd12

Observation d635e683-85f3-4f99-a6ab-f7e55e92cf7b · outbound

This paper cites 2024 , url =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2024 , url =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.096970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.096970Z digest=sha256:af897b43a66e3e99628172433c0fb629950eaba11b357155aeefaaaf0cdd4600

Observation fb257b68-31f4-4103-85d1-ef3270f9456f · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.100716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.100716Z digest=sha256:220a60e1ddf60cc0f104ba772bcc32ca2eea1ec8f5a5554ea38732f323cadc04

Observation 03e31e2b-db01-49f5-a0d9-721d3bfd7ac2 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Transactions of the Association for Computational Linguistics , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.104307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.104307Z digest=sha256:942eff7c7fa151f08dad0ea21406ac82044dc74ebd198705f5e51a42f71278e4

Observation aa2cffdd-89c4-4a7b-a48a-739fd4ef3717 · outbound

This paper cites 2017 , doi =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2017 , doi =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.107292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.107292Z digest=sha256:8066411910090944df78254be42fd10710cdd241f868783f806802910697ba84

Observation 4e832734-d4b2-4633-83e7-b6dd533798ca · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.110036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.110036Z digest=sha256:7713898781efa0134d5f54d683705b18fe1a646232e67126f9f04a339bbe011f

Observation e8c86843-8607-424a-8731-9d2d5dd9c9e4 · outbound

This paper cites , booktitle =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning , booktitle =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.113579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.113579Z digest=sha256:be743beeb78980ca2e77a0685693c088c48f8506a4bd46026a214deab3e4c788

Observation 023c511b-a824-4c9a-bcac-9d2c93dbff43 · outbound

This paper cites Constructing A Multi-hop.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Constructing A Multi-hop

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.117305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.117305Z digest=sha256:e8b7619035430414493ab9cf54c149a51cb0788ffb62401983869cd92f66df65

Observation b56077fd-e7f8-424b-be3c-07d6e345a943 · outbound

This paper cites 2022 , doi =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2022 , doi =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.121110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.121110Z digest=sha256:a96eb32d472141586936a42e957cca41f6e36e6baec6f9574bd289e9a75fc393

Observation 6adbdda8-0d87-4ade-b897-d5e465d1cca4 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , pages =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Findings of the Association for Computational Linguistics: EMNLP 2023 , pages =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.125187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.125187Z digest=sha256:ea18fcf6cf615f577fca5edd0b9343b55c1339dc47002c65cbe4f2c9d21bb7b8

Observation 72da754a-404e-497f-b56c-6fda9e9ae218 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , year =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.128541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.128541Z digest=sha256:de66bd647fd99c4acbeae314d670df5bc33cc48e6522236581b6dfe8887e51c7

Observation f3ffc495-9d64-4fc6-84d8-312a3d99acb2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.131519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.131519Z digest=sha256:9624566fa00a8bbe6ef1a6750a1c6967229cd47097dc04625e16d09b34b50db7

Observation be013677-665f-491e-bb22-7ae8ab1adb42 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.134716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.134716Z digest=sha256:e4f4280c7d53bfd3a6c46e6ae01d3c236f7fcc42ff220fb091d04f6e14c5243a

Observation 617abad4-a511-40ae-8edb-cad265016600 · outbound

This paper cites Agent Lightning: Train ANY AI Agents with Reinforcement Learning.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.137852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.137852Z digest=sha256:2c07fc159d194183e9ca56597ff21021a36be3076da1707c7c51453f09c2b3b7

Observation 59c1fd48-65c8-411c-87e9-d3d7961c1f04 · outbound

This paper cites 2017 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2017 , eprint =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.141619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.141619Z digest=sha256:e7d95c26c8faaffbc2fab74a0daac07db3ec5b4ec22ffc4a42b1052be275bf77

Observation 523e20dd-e4d8-4779-a714-f31c5c969389 · outbound

This paper cites International Conference on Learning Representations , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , year =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.145229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.145229Z digest=sha256:37297522ee157bd7e56931f6f508b209e295ec70f0120f87e385c36553636353

Observation 0e4eda5e-ab2f-4102-bd2e-8e91cb38d2a5 · outbound

This paper cites and Gillhofer, Michael and Widrich, Michael and Unterthiner, Thomas and Brandstetter, Johannes and Hochreiter, Sepp , booktitle =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning and Gillhofer, Michael and Widrich, Michael and Unterthiner, Thomas and Brandstetter, Johannes and Hochreiter, Sepp , booktitle =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.148791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.148791Z digest=sha256:8b34502eea50f4dbae947ce676f4b3de59726f2a0be22ffb5f8fc53795ef4d47

Observation 84763708-eda6-4e65-b896-49058a11f9d2 · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning The Landscape of Agentic Reinforcement Learning for

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.152437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.152437Z digest=sha256:31a5255b04011903648931a9a309c98912b0630cc5a7abae564951b5e4cb2226

Observation 80965d57-ab3a-4619-a89a-834df7f72f08 · outbound

This paper cites 2022 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2022 , eprint =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.156035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.156035Z digest=sha256:0e44545d86f38a5c1f77393d22b6c2f0249e62338286551a260adf8fdbd45ac8

Observation 3f07c628-a40b-4ac7-824d-f2385a44f6c0 · outbound

This paper cites International Conference on Learning Representations , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , year =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.159439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.159439Z digest=sha256:b28d02ac99ebba5179d23d8c8ec5a0f520b69c8d0b50629deff126a9babbc0c0

Observation 495aef61-8b5d-4f23-b167-ca8f305b7252 · outbound

This paper cites 2021 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2021 , eprint =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.162828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.162828Z digest=sha256:f3bd5a8db1e5a08e52307f80ae6726368bf4f481abaa2804d8fb2b09c844800f

Observation 0d8485aa-538a-43ba-92d7-ef68f0af247d · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , year =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.166830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.166830Z digest=sha256:2cc6715d2351706681f9a3449a84612682a59e15ba0c6b14e36996264e37bf2c

Observation 58e38d23-177f-4eed-8343-7b3981afaedc · outbound

This paper cites 2024 , url =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2024 , url =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.170574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.170574Z digest=sha256:d890dda8bc650c4890343c91794eb37ac36c6c4645c47ce6e23bc71f1a975e83

Observation decc5674-ad4b-488b-aacf-5882aa543028 · outbound

This paper cites 2023 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2023 , eprint =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.173954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.173954Z digest=sha256:015d77622c9d36e0c71d330c920480a263cfb2408c1dbe9ffa574a1cb81deceb

Observation 082d9a64-5642-4318-b68a-7d791e29e82d · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , year =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.176966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.176966Z digest=sha256:71b4c69c4b650b2d59d3760b623873d8ba5a251d321d30e05bdd0ac7c29cfdb7

Observation 69e6d170-8d97-4f06-88eb-24e6d293c2d4 · outbound

This paper cites 2015 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2015 , eprint =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.180101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.180101Z digest=sha256:abd98c7b21a624ebb925d22724a6a5bebf20d9d40cdfab0717575686f129c41d

Observation cfe4140c-27ae-42bc-a343-271b2e3f906a · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.183105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.183105Z digest=sha256:bbfc7cbb35a13ab1065894a4a1ead4d1c96612c9b3721fd404f69b517c1dd187

Observation 2bab3260-8058-4a74-93df-f92f9a125470 · outbound

This paper cites Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.185956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.185956Z digest=sha256:dd0821c297df820bed451db9aefe9377d9e728f4bc85f44556dbd10c0177db9e

Observation c2aead90-2999-4aa2-9e5e-eb91185f490f · outbound

This paper cites International Conference on Learning Representations , year =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , year =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.189027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.189027Z digest=sha256:b224b7dab2c8d08edd1dce33881ddb3429f34f52b181b4a5637ec0f7d6fb5f6d

Observation 1cc2aa2a-f8fb-40d1-a797-9df933f390df · outbound

This paper cites 2026 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.192232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.192232Z digest=sha256:0e77d9fd8d2a80dc1f706b877e831ef043666edadb8d925f272c6a330324c9dc

Observation 4343d417-ab8b-42df-bab0-c1bf13219df7 · outbound

This paper cites 2026 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.195171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.195171Z digest=sha256:981195ddc9e15daf48c3e647a7dc1082fa99a7d0902195c5227848c2e1707fb8

Observation b6331b0d-33fa-49ec-ba2d-8e864506e8f9 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.198042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.198042Z digest=sha256:4941d676f08f95d8b5ed208d369909b6d71396b7bce4de25fbf4dfbef3824a07

Observation 37b037c6-1def-4f23-af13-511901105b87 · outbound

This paper cites Self-Distilled RLVR.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Self-Distilled RLVR

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.201488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.201488Z digest=sha256:4c212b1d81f141121367a38521711baa663285399a046ba4f225ef2e72fb7df5

Observation 7039d7d7-3d0f-41d0-9b33-6e6203180fe0 · outbound

This paper cites 2026 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.205005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.205005Z digest=sha256:1631fdbaf47862632dfd120f7a441c46a926766f92efb9b425c7a39ba5657666

Observation 7a314ad0-18a8-488a-b9d8-b62c70a61dbb · outbound

This paper cites SOD: Step-wise On-policy Distillation for Small Language Model Agents.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.208505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.208505Z digest=sha256:f775515460b3ea5882f3a5eda662038c0e108b7286aa5bdee78fab97a4f8ac78

Observation c7b042b2-5b7a-49ce-8015-990f5969e8e7 · outbound

This paper cites 2026 , eprint =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.211611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.211611Z digest=sha256:74717f0415ffee329354626ba947af9a01b361c7588c70e1626e7842180f11eb

Observation bf122c46-a460-450d-bb6d-3e530775a299 · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.214520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.214520Z digest=sha256:2216e0ec2d71685e0e941827af15a6b0e016c29bd6ffe40c03fcaa4a0b6be375

Observation c6532049-0386-47a5-b6be-eece2ac59a81 · outbound

This paper cites Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.218621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.218621Z digest=sha256:bae24d68be7e65da3c72e24c1fac93a5e030651ef5b7e404605e9408e568ad25

Observation 27cdb7a9-3d91-4010-90bd-d8a0171ee47d · outbound

This paper cites Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.221920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.221920Z digest=sha256:274c5b0045a0986e86c9e7f72a2d21a1b4f7ffa0b8e4dbdea8d35213a68c9b9e

Observation d5d9fd56-88e4-4bad-a658-4d40cc437e50 · outbound

This paper cites Advances in neural information processing systems , volume=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Advances in neural information processing systems , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.226030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.226030Z digest=sha256:be8542d113891222c976e6b0d0b7bfd258b7296a0f13a83817b8eef929009e8d

Observation 735c0b4b-579e-4418-be6a-02ab0827423d · outbound

This paper cites Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.229971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.229971Z digest=sha256:8eaa0b60435a5b078f0465f6d1714717c5eac824dd55c840b7b6d41a6621c725

Observation ae2fdea5-7636-4866-84c8-4c760c4c676c · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.233653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.233653Z digest=sha256:d7ae4302cbd4bd3cc963a96b714826fd4d60332359eedf048216fb55073dc815

Observation 4bffe0ec-6079-4833-881a-2253446f9b02 · outbound

This paper cites RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.236755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.236755Z digest=sha256:5c7f128bf451f34b9706b615969e391e97f5fdda178c01e867111dade5323c2d

Observation ea0eca4e-ef45-4e0c-b91f-86b8bb820826 · outbound

This paper cites SPARK : Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning SPARK : Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.240246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.240246Z digest=sha256:c837d68652c7d06ffea611994962a5c141895cfb6ce7cded2d97881f34639d66

Observation af3179df-4eb4-4276-ac8f-2bcd261284e6 · outbound

This paper cites OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.243699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.243699Z digest=sha256:2a7da5ab1bfedf31e136bb1cb0500b5851d921d18e8e0ad287cc26af3c76fe5e

Observation ad8dcda3-0ce8-4684-ba55-2e60184c5923 · outbound

This paper cites DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.247710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.247710Z digest=sha256:78c90b5a790ac1ddfe4c0c5d1d9f0870d3545549ede0bbe63b36935bbfad45bd

Observation 9639b6e7-7ac7-43e9-b175-c810cb57eae4 · outbound

This paper cites ATLAS : Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning ATLAS : Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.251713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.251713Z digest=sha256:ccafb27458717dff1c04272d9a9966ee8a45f07e8697d198a187dfd237a5fd32

Observation ef6b86e6-db06-4c4f-900e-fbc8d55f6594 · outbound

This paper cites 2026 , eprint=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2026 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.255768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.255768Z digest=sha256:819a8dc2d31c1e8fb674de54827fcd8eb8dc0655acc4a765042b5e439e336dd2

Observation 35a0467d-beb6-4029-969f-688446922386 · outbound

This paper cites arXiv preprint arXiv:2601.18137 , year=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning arXiv preprint arXiv:2601.18137 , year=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.259062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.259062Z digest=sha256:0ee2d4183dce5d240367a663b24d69424aca7c0837e5d196ca8240ad3aa3c52a

Observation 84746ed0-21e8-473f-a199-62eb1b6413ca · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.262611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.262611Z digest=sha256:cf9fc91480f8b90de0bb3e4cc8be7aaaabee2d46a56acc8b3c1e3d0042e06508

Observation 592a29e5-73d6-4c40-a5c4-d799c0ae4241 · outbound

This paper cites SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.265931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.265931Z digest=sha256:48987adba4cb37e42803f91eabef86fc81ad2b48b5f63b7ddcf6298123f5020c

Observation e75e843f-7333-4a22-a0dc-f831cc737ecb · outbound

This paper cites Qwen2.5 Technical Report.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Qwen2.5 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.268893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.268893Z digest=sha256:22433fba576fcf1ec54008aa752c7f8ace4e74377e1cd26b56675707c59763fa

Observation 6ad18bb9-7b1f-4b0f-a28f-e6014469903d · outbound

This paper cites Qwen3 Technical Report.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Qwen3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.272008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.272008Z digest=sha256:98893341331fe402642b5bbd2f2e954c3afe22fb440fcc7c18381aa6d9c4d131

Observation f0432ea4-cd50-4567-9c96-6c3044859257 · outbound

This paper cites Qwen2.5-VL Technical Report.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Qwen2.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.275184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.275184Z digest=sha256:9b73f0c6b5007f6df8bd1cd138eb4f98933f63150b583d1befa3b5bdbfd707d3

Observation 4e0c3948-de58-4b2e-871b-193bf04ddd72 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning , url =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning , url =

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.279131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.279131Z digest=sha256:65030a70a8dc9e25fc4304724267dfc831f4a56adda9840af3f50a96bcb7f1ed

Observation 6bb645eb-5377-4df9-9c3f-228674e0260e · outbound

This paper cites , title =.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning , title =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.282994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.282994Z digest=sha256:c9a0c859da9be1f76f7257061d6380c5e3f362726ea52ab1d6a683c3bbb55dc1

Observation 83b1f7b7-314b-44b2-9291-4dc98d2fdcae · outbound

This paper cites an unresolved cited work.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.285839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.285839Z digest=sha256:b1e516345a627b89c00d594bd3ae4ec58a70e6be2a7c7eebe0f59f40cdb4ee05

Observation 8f44b57f-241f-41ad-897d-649b179f1636 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.289491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.289491Z digest=sha256:5cc3aa97261410a301e5de9e22b05dcf3d69f2f3c9658e402c7256006c0818f8

Observation a1e6bcff-1dcd-412e-b21a-1df5ac5fae9a · outbound

This paper cites Forty-third International Conference on Machine Learning , year=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Forty-third International Conference on Machine Learning , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.292396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.292396Z digest=sha256:032ddd2a0abd882da02e14279ec6bb3f618174b6f8231d0fb85329bfc22e6f37

Observation ff752704-dcd9-4159-8923-4750dd7bee3d · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.296070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.296070Z digest=sha256:12f1220bc6b6af6fea7cf657e34f0e0fd929692553f370e05518518d040fd615

Observation f65b1c8a-57e0-4454-8180-1a46979f5148 · outbound

This paper cites International Conference on Learning Representations , year=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning International Conference on Learning Representations , year=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.299566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.299566Z digest=sha256:7ec4f2e4d9c8200d6592cb577bba17cebf28d77f62f766541c0a746d2ebb3f17

Observation 82147def-5fc7-4604-aecb-5efb3a05916b · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.302473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.302473Z digest=sha256:6a72e747f465bff23cce8636411594e0829a063c85fb38efd0c69f1577bdd7ca

Observation fd0a1e02-7863-4bd0-aa72-a64d88f4abbf · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T01:10:21.305795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:10:21.305795Z digest=sha256:2d4abf4d41a04159150ed7568db4e519015d6c0609ce0f89cba86544d24ae862

Observation 8aafa1ea-e4aa-4d18-9d27-de137671f082 · outbound

This paper cites 2025 , address=.

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 2025 , address=

Reference 74

Resolution
verified exact
doi, observed 2026-08-02T01:13:58.398667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-02T01:10:21.309226Z digest=sha256:21d8be169d9b6164541d06f68772574325891f5e23c5a3bdab7ca132302e5314

Pith citing papers

Observation 2530b23c-e47a-438c-ba50-134125a446bb · inbound

EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation cites this paper.

EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T15:30:02.812777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:30:02.812777Z digest=sha256:febaa76e651d2af86fdec25a578be36f6606cff06a8cb7a331c7a6df04ed2cdb

Observation ab03c17b-e6c7-45cb-9dca-f1fb3855e2d1 · inbound

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation cites this paper.

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T09:20:20.130244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:20:20.130244Z digest=sha256:0309b75f1a3d6fe5bb0678b821ae88369a9d4c7a3d211428b780212699ff32c5

Observation ae6bd321-e27e-454b-8cf6-342a6b55124e · inbound

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning cites this paper.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.309548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.309548Z digest=sha256:8b858723f01fe85c2ff347d707b7a5fbc7cc755872c35da0cf6e4fb1f97e2164

Observation a54fa22c-d4a9-4e5c-b091-7fcec56b98e6 · inbound

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation cites this paper.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.054402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:39:06.711304Z digest=sha256:10ea99b07a30c03526a99f0192c089fd61ad49514d27a8d6c0269996f58da4fe

Observation 0c69cbd7-d25b-445e-a4ff-0cc00bd9f76b · inbound

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning cites this paper.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T19:53:25.906690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:53:25.906690Z digest=sha256:dad799ffda228f24991847c2b3d4827c5b742d961d0f73747d8c59dd8bc12960