Pith. sign in

Paper Citation Record · LEDGER

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

As of 17 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.09555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09555 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:14:31.352194Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 507b7d8a-38a2-42b3-b2b6-0acf7a468d1a · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.976180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.976180Z digest=sha256:c54779fc510eec43bee52d5f558c48cbf46d4878770e8064b93458ca33a8c5cb

Observation ec0e4c63-4d3c-49c9-9246-db7a4424b54a · outbound

This paper cites 2022 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2022 , url =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.980248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.980248Z digest=sha256:37a4932b61f5569d49bb00427042c6d0c948cbe1d0a342d9100324947d10fc61

Observation f414b360-34f2-4511-a4ab-4ad9c7342747 · outbound

This paper cites 2017 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2017 , eprint =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.983543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.983543Z digest=sha256:5e0728523a27274e7fe293196821e405fcabab4d6650466a1aa3342f09160b3c

Observation 9f3f7cbb-3c51-4871-9536-6cd3e0f3067a · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.987646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.987646Z digest=sha256:142de1158df1be570e1be88de67b2b988c9316ad49545d0008d46b1bacf2ab75

Observation 3efc714c-d80d-41c1-91f0-89e9c2ece75e · outbound

This paper cites 2025 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2025 , url =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.157380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:30.996013Z digest=sha256:e4818c0441fea5b63b474b0276d518e1fbd00c70fffabece507cd62d6705000b

Observation b4e1fdf8-05e5-4121-bce1-fb5acde55492 · outbound

This paper cites 2025 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2025 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:30.999591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:30.999591Z digest=sha256:043cf2457b41b0741f02f3715d5db4b1eb2020fd0093b2424269909533bea38c

Observation 92734b7a-db07-4d6d-9659-552b77650e4b · outbound

This paper cites 2023 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2023 , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.003152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.003152Z digest=sha256:a1b2bbbfa744e86887d5acfa93cf9830552aad5906169a385018153cea7112be

Observation 3bece669-2d40-414f-bb49-228eee50de32 · outbound

This paper cites International Conference on Learning Representations , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations , year =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.006428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.006428Z digest=sha256:b501d65c26c47974b7859530ee5db3fa0c67f3a92166fb2743c7b300ebd6f384

Observation 99017870-c196-4656-8f59-c359729bf9e7 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents International Conference on Learning Representations (ICLR) , year =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.129503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.010594Z digest=sha256:1002aa300c8ba80d2d343b026412756958f72dd329435b8ff554760af7bbca2e

Observation e40c3af5-00e0-48ec-bc7a-d4b0a9fe2f33 · outbound

This paper cites Second Conference on Language Modeling , year =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Second Conference on Language Modeling , year =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.118597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.015953Z digest=sha256:d054cdc5a29ac2853a3660c5ac50c9d8ef4c0d1ddf6ef6c1491c3ce5af20d995

Observation 6dd801f0-3f9e-4817-8d7e-3e1cfa13d9cb · outbound

This paper cites 2026 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , eprint =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.109156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.020099Z digest=sha256:4b3e08b6f8bd6aaa742768a278b53ab131470f48e5c670aabd45b3e4c4532024

Observation 735f970c-752e-4b13-9537-801c6497fb0e · outbound

This paper cites 2026 , eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , eprint =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.054296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.054296Z digest=sha256:1e696c940e68c86d43db6bd5fa20db14e161d842601371a450cc2c16b84efc96

Observation d178ae83-6349-47c7-9064-aaad55c6bdf3 · outbound

This paper cites GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.058179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.058179Z digest=sha256:440d2367e36d684e95e34e652d404f15d8e9b3361b9e6ca092e15844fcabf24d

Observation 77e13218-f9ee-42fb-94a6-63aad1b95416 · outbound

This paper cites 2023 , url =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2023 , url =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.092993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.073931Z digest=sha256:3d517012f2a1fa511d4b4b3c92f5e956505156981c6dde9b2eb987c42f66e0fb

Observation ea3cc6d0-a8f2-415a-8700-49c02817f0e3 · outbound

This paper cites 2026 , month = may, eprint =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents 2026 , month = may, eprint =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.082071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.078010Z digest=sha256:7c137089510836aa6ac0c83e3b6c3f3ddcbfdeae94bbc297bd0d2de02d970b1c

Observation afb1506b-0360-48a2-b815-4fcbb778c4b2 · outbound

This paper cites Proceedings of the 43rd International Conference on Machine Learning , volume =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Proceedings of the 43rd International Conference on Machine Learning , volume =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.071353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.093811Z digest=sha256:23c7928b01ca774d52dad431d271e3fca0c137d7f0a321d4484074a3d73fb394

Observation 9625b209-5e08-4a6b-9e98-aa3ff18ffc87 · outbound

This paper cites Proceedings of the 43rd International Conference on Machine Learning , volume =.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Proceedings of the 43rd International Conference on Machine Learning , volume =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.060992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.099335Z digest=sha256:05a04caf2d64ed943f0e246401089637d7a0acab35a3463521691408e5faadfc

Observation ec4464c1-269d-46b5-87f9-ef545c6507cb · outbound

This paper cites Advances in neural information processing systems , volume=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Advances in neural information processing systems , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.125179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.125179Z digest=sha256:327e2167a7864760f4e4f15170ad90daa8d2127a179d019ff234988e0830a522

Observation c3291273-6664-455b-8734-e00f2be9b645 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.139246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.139246Z digest=sha256:fb436ebde22d153da71f8f65340d1a7c20ac2290f4b5eae04adbece4d0355b15

Observation 2fadadc6-3f06-4d56-bdb6-f4c7f8fafa4a · outbound

This paper cites Forty-third International Conference on Machine Learning , year=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Forty-third International Conference on Machine Learning , year=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.026672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.142247Z digest=sha256:e62274fbdee33dc9b070e454cabe645854ef97fb8e7591bca87d9327cd22a9df

Observation 94c8467a-5574-4b64-b110-113fa04877b1 · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.014141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.145122Z digest=sha256:a21f3fef39733171e48384f1ce4db243dd9d9cae5713d118fff7755d9368dab9

Observation e3d8046d-db4a-4506-925d-baf8001abceb · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:32.003465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.162221Z digest=sha256:1aa9d77f173e8e0fe89605eebc23293c9058aa4ca36ec454ba6d102dc3b91b5a

Observation c2e3fe46-1ea6-4bd7-a256-7a6315a69c4d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.992407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.173862Z digest=sha256:ce68668efaa3089aabad248771655244481b57a29ae729d92bc64e03a4312936

Observation 0fb8bede-4e9f-4bf5-80ef-f88c3eefe78c · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.177934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.177934Z digest=sha256:fa61608786c9051a0aeef33591ad581a699ca3262db46f90bb37791789f5547b

Observation b1f247c1-3cb1-4910-83e1-40bf05a79537 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.973251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.183855Z digest=sha256:910e293a71f75d5402923da31e8010c4ee1a253c6eb4ddc8fe1a3be4e163b5e2

Observation 7903429c-8806-47f5-980f-186ef96cb48c · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.186850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.186850Z digest=sha256:5174b29c41afc27a0a8bec5b8603e2b94eeb3ae0017632419e29020684f4e1e3

Observation 4593b1e1-e865-42c5-80b1-38575037517d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.189966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.189966Z digest=sha256:df9cbecbfbfbeb558ef40db0d3b61ab222073258a24eed1addc3c36bf4e7b63f

Observation 9f45d545-e925-4af2-bfa7-27ec2a10c16d · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.193019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.193019Z digest=sha256:85a316c981450d282d91d72eb033648640b9684219e24771c5495880368ca620

Observation 71d40a3b-ce9c-42d6-80d4-a5be85594e34 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.953664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.196233Z digest=sha256:538d3273c3120f9daa93f22848469e8aeb9d1290b7ae8a4290c3f1b8b2e52a53

Observation 6fa0560f-0a04-4218-8b47-06a2655faad2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.199401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.199401Z digest=sha256:cea758339dd2d050aa10cae9294fe5204d10e58cfca68ed0007bfef91df96ff0

Observation ba50a996-3829-4f31-aa03-7ad9e382394e · outbound

This paper cites SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.203282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.203282Z digest=sha256:f27e1876eeed773e0b677493aed284f769cbd3e7d3dfcb745210c3baa6f5165b

Observation 1a5adf60-4924-4522-bbdc-8762f73a4f17 · outbound

This paper cites u botter, J.; L \.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents u botter, J.; L \

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:31.941921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.207650Z digest=sha256:6e473623e4dbf21b77d47345784539b7c15f7055bdac6be7b227f85a4d8902cb

Observation 332ac8d8-a995-4f21-9355-a4f861712fe7 · outbound

This paper cites SoK: Agentic Skills -- Beyond Tool Use in LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SoK: Agentic Skills -- Beyond Tool Use in LLM Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.212116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.212116Z digest=sha256:fcd7c58d7aff3edf3b3c15511f642a7a7568e1e2fe198f7f3db35f0f8b4ee5b9

Observation 2aab8971-1b3e-4fcf-8d37-b6971394b381 · outbound

This paper cites EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:14:31.556314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.217336Z digest=sha256:df3232a0bcb70009668057c29d2c16cdfee7e39d75283081e78db54e579877a5

Observation ebec9a3e-062a-4e55-a20f-6ed588e00157 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.929641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.221025Z digest=sha256:e526545f5d4dd634dd75ffa3ab558ff12109cbbe603f8bf92975eda6724b5507

Observation 7609e894-bf27-43ad-9b10-a3b69843474f · outbound

This paper cites P.; Li, L.; and Li, Y.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents P.; Li, L.; and Li, Y

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:14:31.918665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.225101Z digest=sha256:3ffc997d5e93a65e5ec0c2ec4f4edf684c2890f5f90e927b38bf978e6b9f3971

Observation 503145b6-beff-4139-ba22-c6a09da9dca3 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.229838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.229838Z digest=sha256:296c82552c1241e430f26dba9d7c08d7b70f3ecca03a34c6f61cb07e09978110

Observation 6a4ff8dd-f86f-4355-890f-0676403ec2fb · outbound

This paper cites SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.233000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.233000Z digest=sha256:8d89cb994a6ac92ef0148ce04fb61254b26651a6baa06b4b505e23d02bb62f0c

Observation b2730c46-5290-4f27-89a8-d2e32580b505 · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.236567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.236567Z digest=sha256:2151740e5e403909314b0297d95776fe1cb4c287edb040c7c8f77f0e8ef539ac

Observation f0c112e7-467d-4bfd-ba97-f50e376d4d6d · outbound

This paper cites How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.241692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.241692Z digest=sha256:df8296d03f1d3d1dfc13b41749bb6565204d81504b5c42a497bc0e8b6dcabc93

Observation 6bc7b0b5-192d-4e4c-bbdf-54a7ffb1cbc1 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Self-Distilled Agentic Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.246337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.246337Z digest=sha256:4af1bf5f0eb74953299cc7acf1856196820c439d30da7700581bf74cfc79732e

Observation 5921ee81-82c8-4cdc-b881-177b858037a4 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.251390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.251390Z digest=sha256:b9d75d9d4d3556b7ff30184385e004806db3b0b8d23f4774f56f6c02104ae55d

Observation acaa1adb-a251-4f1e-9c7a-a594570efc28 · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.255877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.255877Z digest=sha256:e74577a06ab7d40951ee7102f4e7dee4e1355d79e23ad813403bcc6b0c3acb34

Observation 2fd1d420-36c9-458a-b1cc-3269af44cfce · outbound

This paper cites Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.259533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.259533Z digest=sha256:cd45205314e093d4f092455a5c1d98c04b19ff487edb529c526694ca795b4bb2

Observation cc3f5913-16a2-4ea7-bcb1-91dc3ab8e900 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.264307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.264307Z digest=sha256:821aa1648acd76942773d1983aa3fa3d72bd79542024001d859ed38af58a42ec

Observation 9135da0c-edf4-4732-9c47-80ad2c89eb9a · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.268830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.268830Z digest=sha256:770b938e199e7c7813224cc57f38cc0093b107ca5b22223acc029af3425318dc

Observation 1943922d-9eaf-45a5-892e-17d2320a0630 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.273003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.273003Z digest=sha256:3d28048508899bf2d1247465ff73e824613d31f76f033fa619957f879f5678ac

Observation af4d37ac-b3d8-42e5-b92e-e11cc2b8acb3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.276375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.276375Z digest=sha256:bd811f53a322b9a128c9a430498ea0986b13f0f7cbcefde07e56aae5dadb75d0

Observation f55933d3-ef4f-4477-af85-e6428f2678e7 · outbound

This paper cites Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.279791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.279791Z digest=sha256:11923e30878b029dccc8d57e1f3d7c5a2c95454b53a0ef1d9251145e908558e0

Observation 79d3ff01-bf7c-4ee0-8f10-d36bb43efd3a · outbound

This paper cites Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.282967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.282967Z digest=sha256:d752abacf6ec5ade7a19dcdd29f16db582bc7843225b3853109e11624e810c19

Observation 5925ede0-620e-408d-85f7-5ffcc4a6cd17 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.900861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.285811Z digest=sha256:ecf8f86442ff400a943530211afb517c75de2c9121820ed9ecfebfb1acef8f01

Observation 03d315e8-672a-4aaf-9c97-5582db8c0e50 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.289006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.289006Z digest=sha256:55a74a056be182a12fbc208948cd2e07748d0e33072cdac1ee1bc36aa11c821f

Observation 65f63379-d7dc-45ae-a4dd-85fe6f5ba98a · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.293116Z digest=sha256:1fcb96cd5b689e161bc451b411ee0affadb552ead76f4b1f7271483d1c02166f

Observation 4a6e1d17-bafa-4b86-896e-574f0074307b · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.296243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.296243Z digest=sha256:091ef2457a6329b0ea88da4430ec7d3c3618e93bc91c0d020c837cf03e8d4ba3

Observation 696b4123-bccf-47e3-9afe-e2687846f7ec · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.883537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.299306Z digest=sha256:f9013845cf537fc2c2ba3e522ddc60da8fd3b3f217cd93f5277101d908a565e5

Observation a6f46b7f-6602-40b4-bc82-886745447927 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.302274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.302274Z digest=sha256:af619e836c0eda6fd9572e6981d19403354db058ece2a5623b7a67aff2bbf21a

Observation f746a172-24c2-4525-ab6c-2d609408231b · outbound

This paper cites Qwen3 Technical Report.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Qwen3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.305540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.305540Z digest=sha256:066afc812b532b77a6d640fdda592469828e5a3bf8ef3b58e3ddb7a3ab1a636a

Observation c4a4b54e-0ca0-4b12-9818-4f47142fabfa · outbound

This paper cites Qwen2.5 Technical Report.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Qwen2.5 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.309205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.309205Z digest=sha256:4a00bc240eae3a8f1414b362d02f1df696faef89163236733453f1d8c22b929c

Observation 53025679-a086-473e-8a08-d6103d765c5c · outbound

This paper cites Self-Distilled RLVR.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Self-Distilled RLVR

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.312419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.312419Z digest=sha256:07526c702ef19212aae2196d27193fc33fa417e79258f4febd26835d1e72c81a

Observation f5ef13e7-1098-48c1-96ef-4159d503cbcd · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.316254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.316254Z digest=sha256:0f7bd316e1dc000e95b6e7a889ba8a8a31b1d42de6b175c3ec8b3727c19e83ae

Observation 406e0b03-90a7-46b6-93ee-5ce3060a1db3 · outbound

This paper cites OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.319381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.319381Z digest=sha256:5d975b8b93a496d356e44a426d1a193d118e0e4143671c01c30e0d00cf7bec7f

Observation f08f7ea4-21a0-4269-b632-0ab9f03f23e1 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.872955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.323263Z digest=sha256:e1e28c3f3a2236727938428d4107872e675acb5c31935310dbbb8f387c89913e

Observation 88209d9f-c0b7-47df-a118-3fe92ac2f159 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.860620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.328269Z digest=sha256:e6da3182245af23ef23431b2c454c8daad6be0e522e7ac024b12b152d4ccb16c

Observation a9b215d6-3d46-4fa7-bae0-17f50f5065f8 · outbound

This paper cites SkillEvolver: Skill Learning as a Meta-Skill.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SkillEvolver: Skill Learning as a Meta-Skill

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.332594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.332594Z digest=sha256:1e705dedda458780625177cc6fa00befe7db16085acebecd6205c7fc8ae5c4dc

Observation 0be07b05-d407-4886-b3e9-ad59563e7207 · outbound

This paper cites OPSDL: On-Policy Self-Distillation for Long-Context Language Models.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.336027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.336027Z digest=sha256:d02eab9717ede80623899ce2e7e0b84013e9f833a7a7d8f0cc9d6603074d5a2f

Observation 09ad5cc2-8a49-43e4-9f4d-fdf361eb19f7 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.339899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.339899Z digest=sha256:028fe98272c435c91ac0f04129265e0a1ecc9e5bd44cf6ec17a2c255a3f61b3c

Observation 2a11cf51-f63a-403c-aa6c-549e0cf300b3 · outbound

This paper cites an unresolved cited work.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:14:31.848387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T15:14:31.343797Z digest=sha256:8edde5000878a64082b14f9344e92f2933addb8ff48aad09d23df8435d0f0243

Observation 2782d5f9-46e7-437d-90e8-8929a3f50b8b · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.348326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.348326Z digest=sha256:055d018242bff0f108822735a9fd0f7761eee83a3a4b7635ba482364e3f388fc

Observation 599e9b4b-7b41-4dfb-b197-fbe4ed3e0544 · outbound

This paper cites SOD: Step-wise On-policy Distillation for Small Language Model Agents.

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T15:14:31.352194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:14:31.352194Z digest=sha256:4d7d36f3cf55e3cb72554e338f6dfca923c2af375fe8d582d6fb414b9a9e0fbf

Pith citing papers

No inbound Pith citation observations are available.