Pith. sign in

Paper Citation Record · LEDGER

Agentic Reinforcement Learning with Self-Distilled Reward Shaping

As of 21 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.03223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03223 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:15:50.932630Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8675798f-b798-4d3f-8c9c-fdb46028c21f · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.861115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.697103Z digest=sha256:be41ca77277996efcb0fefc3f69cd8eb1b6eabc73b59774ab2e11d5889a6ca3d

Observation ff4f27e9-3352-42d5-86aa-a439fd126d03 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.846448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.718401Z digest=sha256:2907ae4586f216972478bff89a3fba02576a365c665d0433da3d59ae09532401

Observation 9adefe45-cc4b-402a-a50d-2d780189b19a · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.724043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.724043Z digest=sha256:acb016b8125bc388ab976b9e3b7c25ff1ffecd32dbb2a9ca465742325fb9a30a

Observation 92a3e15b-408b-41b7-aaf8-4ee8842f5ce9 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.729905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.729905Z digest=sha256:76c388324ee11a278f8c6ef49fd04f664190453bdd1e41acd9d1a57658fe668a

Observation 13156ec5-b935-4634-b9ce-a745c9ddf324 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.830528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.737402Z digest=sha256:e6e9fa563fc2ddfa9d61b6f3eb3c230a70fd2fb707c02a068822ae428f3821fb

Observation b138fd6e-ac23-4e3d-87a5-68837a7b8bbb · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Distilling the Knowledge in a Neural Network

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.742839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.742839Z digest=sha256:bc4b10b68b7d799d1f8f148e7aec12e326fa81319aad464666fdf5ad56c9ec02

Observation 71826fd6-80ac-44a6-9ed0-dd1d21f9e9ef · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.749272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.749272Z digest=sha256:411a3b7efa57f18cc4f69f8da92f9e45b2df0db60c162b7781d5dfd44636b682

Observation baa0c435-e836-4b37-91fd-d4d9b6e538de · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning via Self-Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.754062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.754062Z digest=sha256:684428bea64935f186e1d9d99339d1cc8e3b6d19ea9d06e56f10589f94344583

Observation fab17f2f-af4e-492e-8d1e-40c658fa80bc · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.758886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.758886Z digest=sha256:963cc6de8e23fbccc4ad0d6355bc9301b29cd173f0435f70ddcfbfd23d7bb3cd

Observation 7ebed2cd-4208-4aa5-bc04-7de91ad3f08c · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.804385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.763497Z digest=sha256:f3ab150c617eb2319b2324f74d9234f737f4e25b6190d1a9c5fb16728ee0c2c4

Observation d37c7dd3-f9be-4136-8c78-68c3880e9b1d · outbound

This paper cites Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.767987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.767987Z digest=sha256:69d5039ab8dca445c44fc1494a4c92c8bd9f477c8785dcca4d2dde799f392941

Observation cb700ff0-c4a3-41e1-926a-7527f6a98784 · outbound

This paper cites Naturalquestions:abenchmarkforquestionansweringresearch.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Naturalquestions:abenchmarkforquestionansweringresearch

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.788550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.773613Z digest=sha256:c2770a74501086563ef33dead503f5f159a7d33dbdc2e17a9c384225d8ac500d

Observation 4bd96b65-717d-43ab-8852-d01b4c418612 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.771787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.778337Z digest=sha256:22749ccf99104ae601fd05a5358b6406432765df84ec17b211e1fcb20a6e23d7

Observation e3be0fea-2032-4c84-8e10-cf1b97b8d851 · outbound

This paper cites Self-Distilled Policy Gradient.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Policy Gradient

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.786116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.786116Z digest=sha256:40675dbcb8dcdccf73294c0804cddaf918c1e3503e44906921675ad381fe0a8a

Observation 508bf451-91ee-42c2-ab64-7fdc6610c639 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Agentic Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.790996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.790996Z digest=sha256:57a224ca1313f0ca597ddcfe1735c12e726695a856318fa428bcd511fa9f4b53

Observation 525c4e31-9b8a-4ab2-8e01-d00e5e1783af · outbound

This paper cites Agent Lightning: Train ANY AI Agents with Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.795278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.795278Z digest=sha256:f7e18d36f111e9fed188f314c76ab1a57a11e2612c153cdc51cf2b0c27fc92c1

Observation 5b79254f-3fc5-400c-8af9-09f6132a4b10 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.750157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.800345Z digest=sha256:792da7cebd1f04926f03715464885ee415aae513d1754b6d6585bd39b227dc8d

Observation b605da9c-739d-4e15-9d83-ecdac7487221 · outbound

This paper cites CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:15:51.319505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.804723Z digest=sha256:a0765e6d89a58c6fea31c8c69989fcf5485888150999785b57f740f0c854d729

Observation 1011db5f-b902-4973-b2a2-3ad1da50d45d · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.730575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.809974Z digest=sha256:f83df49ab87a0edebab756423d92c8976a399bae0fadc56c40d84778df88d33a

Observation ecdbda85-cc2f-4523-a3cc-f02fe3097b57 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.713178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.814148Z digest=sha256:8953a5d33c2399fe29484ce613732a2c5bf8601be28423778be68797a442748c

Observation 5dfa644a-6dee-4f8e-a030-92755311a9dc · outbound

This paper cites Qwen2.5 Technical Report.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.818627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.818627Z digest=sha256:4bcfd5f8853683a379dfb1c9589f4eb281570c62e3c13cfc600cc31872c797ea

Observation fd82cd44-8337-4a08-8b01-2c0467438122 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.693931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.823637Z digest=sha256:e9103472d5db3e8a0396d8422fa75c077a5563fb3b75a07ef25f162ffa68442b

Observation 035ec4cb-5dc6-43fc-bbd5-e08891a2cefe · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.827697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.827697Z digest=sha256:d36f8e867e2b5d7097d43dfe40e16c00ce1fa4ec68c9a1ee3df2d86788957a5e

Observation c6d0de95-b5ef-4323-9231-c65a20275e07 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.837192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.837192Z digest=sha256:042fd262d404ea13ba2152316dd8b3a6050f0803ba0f0fdd19a91938c2f544f5

Observation 81c9f18d-99bd-4bcc-adde-d2766986be77 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.843229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.843229Z digest=sha256:7d5e2dc66db2bec4dbaf6b069c5d49d64e90c9b8641ae03873839070e81a8721

Observation ff8d01c5-2bd5-42d1-9e20-50c0d9be4d9a · outbound

This paper cites Learning by Distilling Context.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Learning by Distilling Context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.848819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.848819Z digest=sha256:69ce08d83f117146cd7b747c84b1f7e739379e3365f707de4a3e3fadc814e37c

Observation f9fb4306-7ed0-4c05-affa-c6e28e54c521 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.853031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.853031Z digest=sha256:883d64bb1913efeaee72ea979196e6c0b630e263d3bb9093d0cfb9902c86bb46

Observation 5cceeed0-e3aa-4b6f-bc35-7df47754fbee · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.636021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.862402Z digest=sha256:981b750c89ce52e1f41c048dfde25fdea99018ce0fd2b78170dd3d98ce67a800

Observation 38069836-2da3-4d8a-a6ed-bf13e7a964df · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.866749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.866749Z digest=sha256:3612b38994b118fa068050ddc3f9d00099828edeffdf1e1eca8f13058244b9e8

Observation 332f1b5f-613f-4323-b26f-5a8389fb2d4c · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.870908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.870908Z digest=sha256:176ba519bdd2a2235e07f96ea45bfe4b3cb6769720721eb878c5b4924d3985c8

Observation b747ed9c-5279-4141-b5e1-09f56823bc86 · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:15:51.619372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.875491Z digest=sha256:bb3fee5fe054d5eed9b0287a33dc4bdac2643e86a3444bc76976508ea5b8df02

Observation 5e49090d-7305-422a-9c4d-dbb8011c04bb · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.886657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.886657Z digest=sha256:abe833dad2f314555524d99d19e1c8dfc2cead962943470e0459d1e52a0de156

Observation f8857167-3609-4823-94d3-4c04bdb123df · outbound

This paper cites KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.891014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.891014Z digest=sha256:9269b19a62fd20593467b792fba0f39de0fb8b678d1063099aa10ee256e5b0d3

Observation 580e16ac-39d3-4085-96fb-0f8cca1a2119 · outbound

This paper cites TIP: Token Importance in On-Policy Distillation.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping TIP: Token Importance in On-Policy Distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.895596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.895596Z digest=sha256:338ab35aa8bfd53233f54b9656597b64a1e449e706c6281a4cf87c9852e32be9

Observation 62ec70ad-3164-4a3f-86b3-0c610941cba0 · outbound

This paper cites Qwen3 Technical Report.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Qwen3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.901475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.901475Z digest=sha256:fa2ab90e2b77e04fb1d159b20d7f39370c04958fef2d7e94ffeebf914778e776

Observation e73c43fc-446a-46e9-bf20-de450cbf756f · outbound

This paper cites Self-Distilled RLVR.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled RLVR

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.907312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.907312Z digest=sha256:b555e36323985c6662e1a1a4794ef19af17467c2d238711daa2eb4d0aea11e6e

Observation fd9b2f95-9564-4cf8-badd-dd4a3e8e1e6f · outbound

This paper cites an unresolved cited work.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.913796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.913796Z digest=sha256:8a3358565b85822d2d543ba82751f5b67d5c6ea6be1d53224fb543e25ba59571

Observation cce0ed66-ec96-4eae-b737-28ed9bc82c38 · outbound

This paper cites WebShop: Towards scalable real-world web interaction with grounded language agents.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping WebShop: Towards scalable real-world web interaction with grounded language agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.590013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.921335Z digest=sha256:88ad52f84399a433f404dbeb1ce7505f3187b8bd73c39aea0144e5b2688abed3

Observation b3da775b-c531-4833-81db-c3208aff0302 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping On-Policy Context Distillation for Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.925900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.925900Z digest=sha256:4012096cdd79c0694e1ac21fbc9a8d0fa4faa45f5ad3cf77e164e32d4607062b

Observation e5ec1a51-7670-44cb-ae15-62522eb0765f · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-08-05T23:15:50.932630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.932630Z digest=sha256:2c116395843f93d78797dbd84387c0ca36fac5aea604a44eb73a29ccb2994abd

Observation 24d858c6-4fac-43ac-9b1d-6f58d24bf818 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.832775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.832775Z digest=sha256:44b7ea708a7c036c2976879b82247af928fb30e9754092e1074a0ad8e97bc078

Observation 73286c70-c049-4d27-ad0f-0f428c56a001 · outbound

This paper cites Transactions of the Association for Computational Linguistics10 (2022), 539–554.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping Transactions of the Association for Computational Linguistics10 (2022), 539–554

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:15:51.654412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T23:15:50.857722Z digest=sha256:e5ee5ebd4cb9b461504467ead00813d1a402a506107ed120b4933ef1b51b1aa5

Observation d9033018-4208-4139-a8c5-f5fe7d1a2488 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Agentic Reinforcement Learning with Self-Distilled Reward Shaping RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T23:15:50.881548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:15:50.881548Z digest=sha256:1a3e67de1c43a6e5858517bf601c8019cda2726d7618f862f62efda64a36d0dc

Pith citing papers

No inbound Pith citation observations are available.