Pith. sign in

Paper Citation Record · LEDGER

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems

As of 4 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2604.10029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10029 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:24:28.048403Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:44:32.617211Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact19
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d53063b9-aecc-4e31-be81-8f8160b9b60f · outbound

This paper cites Self-attentive sequential recommenda- tion.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Self-attentive sequential recommenda- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.807390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:fb408df0c39673aa2c2268c5d3ce84067c23cb37b0e028cce7a3d743760524ad

Observation 004673d4-d277-4971-b29c-fba74226ca7d · outbound

This paper cites Lightgcn: Simplifying and powering graph convolution network for recommenda- tion.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Lightgcn: Simplifying and powering graph convolution network for recommenda- tion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.802483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:f85f58447b9ea80fb6e557622289ae0c4735e4f962f13e14d423a39c9d21ed96

Observation e049c78f-75df-454b-b40a-8083d12d1d15 · outbound

This paper cites Efficient bi- level optimization for recommendation denoising.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Efficient bi- level optimization for recommendation denoising

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.800251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:50b67187be7f77653d2a24f055303a779be7b15feb3cb0459f68215b2283d7f7

Observation 265aa792-383e-4cc8-8233-2462ee8e66bb · outbound

This paper cites Agentic feedback loop modeling improves recommendation and user simulation.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Agentic feedback loop modeling improves recommendation and user simulation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.804911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:895a2af138230b09d2b98d6464647c76d9a9170b5d59edbf0ef171ae8cb7a265

Observation f6a94efa-f04f-4547-884e-ed9b78454830 · outbound

This paper cites iagent: Llm agent as a shield between user and recommender systems.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems iagent: Llm agent as a shield between user and recommender systems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.810037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:8717278ec129d2e4e268d87a05eeea2024da449851c3ffa39c30935ff259efd8

Observation a19e5a89-2daa-4e0d-9bc9-d48983e9e6ad · outbound

This paper cites On generative agents in recommendation.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems On generative agents in recommendation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.812149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:44c139e2f831ba4d767dd05744148b233ef5fe26e6fe7d43ef830fae2ad218c1

Observation a2014003-e152-4844-8471-6910c3a0d9d7 · outbound

This paper cites Recommender ai agent: Integrating large language models for interactive recommen- dations.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Recommender ai agent: Integrating large language models for interactive recommen- dations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.819617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:c9f983bde0655a67336a794ee3a955f9eff4e0c9a27faafc069c464a91a9e7cc

Observation 424c98d2-d34e-440c-91a8-6febe9ff8d0f · outbound

This paper cites Re- flexion: Language agents with verbal reinforcement learning.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Re- flexion: Language agents with verbal reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.824593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:27907d043b366280fcdc16d7d1f4dcd01cd06fcabd08b64cfd07fd1a6b454c8a

Observation bbca6b1a-1446-430e-a2b5-bf9631f2e591 · outbound

This paper cites Entropy guided diversification and preference elicitation in agentic recommendation systems.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Entropy guided diversification and preference elicitation in agentic recommendation systems

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.127976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:c393ceb7986fba0495c9f3165ec8f1768912823209dc1857e27a04445a09888f

Observation 1e6a731c-e948-40f1-926a-1d0f766b3e70 · outbound

This paper cites Memorybank: En- hancing large language models with long-term memory.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Memorybank: En- hancing large language models with long-term memory

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.827397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:8c8968cfcca7ebe53e1d660b0e0135593b90b552fc5704e9c8f3633f361ecc9a

Observation e05d95aa-3ffd-4d06-a64a-c59c7d9b43f1 · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems A-MEM: Agentic Memory for LLM Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.212756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:e6b1ad26c8caf6dde00ed24b6c1ad87a1a12562ce680da8d6432b7c618477718

Observation 92f4dbd6-ee38-4495-ae89-95f6a3f1f76b · outbound

This paper cites Tallrec: An effective and efficient tuning framework to align large language model with recommendation.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Tallrec: An effective and efficient tuning framework to align large language model with recommendation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.787562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:8c6786dd7860ad8ecd2fefcbf0134d51068067013de8dc766d7cd6d331d596a3

Observation c5bd9325-2b47-4bde-89ee-3988d8ff067d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.225700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:7e32bc9078f01d214de4b5671cdee9ecca2fa67539dfbc5ed41a807c07e897a0

Observation 027e70ce-192f-4647-becc-138042797b73 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Proximal Policy Optimization Algorithms

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.179684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:30898c94aee99ed4afea3e20a909ef59fcffcb1d7002e017883fb42e10137f5b

Observation 2035d37c-4579-4a40-a1e1-93e3616bfd03 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.792482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:53d5a788f5d4fe447525e18e611244ee52959d065638388f837c712fdbbe92cf

Observation a91d93f1-b305-491a-b8e3-0a7b3f18dc95 · outbound

This paper cites Amem4rec: Leveraging cross-user similarity for memory evolution in agentic llm recom- menders.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Amem4rec: Leveraging cross-user similarity for memory evolution in agentic llm recom- menders

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.121845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:687bc3c2d354414f03af8585703aa1623b9f03dac2cb808b08378ebe5b6f799f

Observation 5baec4f5-324a-4310-8ac3-69bae702ccd3 · outbound

This paper cites A survey on agent-as-a-judge.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems A survey on agent-as-a-judge

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.220165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:c1e59e40b9d32dd66622e8b7133bee5e156cb72e2a0f070103efefee0b9603c6

Observation 9a9f5a69-edd7-4e7b-86e0-fae84021fd1f · outbound

This paper cites RecoWorld: Building Simulated Environments for Agentic Recommender Systems.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems RecoWorld: Building Simulated Environments for Agentic Recommender Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-28T02:22:23.314678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:3e644e56476f7ab217e8a69c5e0d853b337925130dfc32879f612a4bc2c64e28

Observation 9ae63de2-cb16-4583-bad8-6ddbaa637a2f · outbound

This paper cites RuleAgent: Discovering Rules for Recommendation Denoising with Autonomous Language Agents.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems RuleAgent: Discovering Rules for Recommendation Denoising with Autonomous Language Agents

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.140206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:504dd5c3f046d3467ddceeee04b1f7cac82bd888ce11433134057f335c55b172

Observation 4f33a700-ae2c-4547-8713-6ffdaa565abe · outbound

This paper cites Macrec: A multi- agent collaboration framework for recommendation.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Macrec: A multi- agent collaboration framework for recommendation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.780685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:979ce34913dc7dd46e4bcb83b550328ac0b5e5858f3ed17bb3b40268a5640e65

Observation 9664e998-438a-4047-8696-33e059526f36 · outbound

This paper cites arXiv preprint arXiv:2602.02482 , year=.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems arXiv preprint arXiv:2602.02482 , year=

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:02.173063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:19f371e40324ecf69d4dd078073f67a3916ea1e4da2f613984616b1aeda85b0a

Observation 8ccd2e18-9a25-4da2-8a73-d799961d75c3 · outbound

This paper cites Treerl: Llm reinforcement learning with on-policy tree search.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Treerl: Llm reinforcement learning with on-policy tree search

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.785196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:c12577e65ba46b0c1c887f1d59b44c3a6796b7ab0eb6439b4be55345996d12d1

Observation aa9dd6c5-6bd4-4fc7-855d-5344e529ec31 · outbound

This paper cites Supervised pretraining can learn in-context reinforcement learning.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Supervised pretraining can learn in-context reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.773319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:bd1da7fd1332446ad0e7960830c9da597fcdf26ab798208876722cbfc773d661

Observation 3be3a04a-107c-4a9a-886c-3834f5d9d2c6 · outbound

This paper cites Reward Is Enough: LLMs Are In-Context Reinforcement Learners.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Reward Is Enough: LLMs Are In-Context Reinforcement Learners

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.070609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:61c093b15239de77f704b51257440bb537102995ac25489f2418e51e07f84368

Observation 3348998f-c5a2-4a65-bca7-2bc632d57b7f · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:ca24dba79913455b732498afc5049f5d2289994a0d61f1de1c9fcacdc84874f4

Observation 491749b6-12b7-4845-92e4-515c786d01c9 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Self-Distillation Enables Continual Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:52.723793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:7a7ab6dde781be7e0303f832ff5686e32c31b7c1ddcd01bfb5d6fab5748fbf2b

Observation c9b4f882-5512-42be-ad90-1f98b2eec84c · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Reinforcement Learning via Self-Distillation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:29e0a4da7889626395596e25879889658b82620b3e1275bee857246b6a6ba82d

Observation 887d32f4-fa2f-46ef-a8cb-213c931246c8 · outbound

This paper cites Self-Distilled RLVR.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Self-Distilled RLVR

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.092723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:a63780ebe0706642f3c82a12be9bbf4c5af70bce51389388dd6aa2afb0266992

Observation 20a7b5c6-b756-45ab-a252-800e25f92d69 · outbound

This paper cites Second workshop on infor- mation heterogeneity and fusion in recommender systems (hetrec2011).

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Second workshop on infor- mation heterogeneity and fusion in recommender systems (hetrec2011)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.775887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:3a4ca39a09b7ff261db70d00a283534ccf4c86b8f4acf87875bdea54ab321535

Observation 94b0ae1d-bbed-426b-ac67-6784f6ac5fdb · outbound

This paper cites The movielens datasets: History and context.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems The movielens datasets: History and context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.778186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:ebb2b6bb3473fab3731c0c40e16323c5798f770dc2717b506871e181724664bb

Observation e6248ad7-e440-4414-bfd1-9e47d9fef17b · outbound

This paper cites Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.195798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:f66ba7584c60bad7be670bd214e57acfb6fdf55ce3d1c6ef42274f09d8daa0c7

Observation 25c7840d-4141-441b-a28f-73e075306bd7 · outbound

This paper cites Let me do it for you: Towards llm empowered recommendation via tool learning.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Let me do it for you: Towards llm empowered recommendation via tool learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.782938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:c262c918e3f23c548dd952d9fc5d9095bb1bafcf421d35870718c0a5433808db

Observation 708f2f17-0135-4b9d-8ab7-89c46c7bac65 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Lora: Low-rank adaptation of large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.789867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:5c36af045c8a6e26e903088188e43a00b52b807f7ae06189a817c228f8e8896e

Observation e39cb58a-cef5-4ab6-9188-47a402cb0ef6 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:839c4654fa2e76bb7f5f7fe926d9453e3135b2b4f10f33794f10ff1c21a26a97

Observation 6dd8baba-c873-45ae-bac6-a8e256312811 · outbound

This paper cites Negotiating the shared agency between humans & ai in the recommender system.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Negotiating the shared agency between humans & ai in the recommender system

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.795096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:72123947f7c759892c42a00e10e95d4730cfa18bbf17bc5ee468b47a0debd423

Observation bc2b7259-1584-4f12-afc5-494c5223dbe1 · outbound

This paper cites User behavior simulation with large lan- guage model-based agents.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems User behavior simulation with large lan- guage model-based agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.830002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:f81ddf815bf98cc00ff0303bef3b4b934e3470fca38e75d9ced0484594485b7a

Observation 2153ce36-707b-4465-9fa8-4d405eba717b · outbound

This paper cites Recmind: Large language model powered agent for recommendation.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Recmind: Large language model powered agent for recommendation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.817428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:467ed3c03011772dd75f08529f5f1cb9bfc313c9965fe6473e43e3293844077f

Observation 3d8a3fc9-d1e3-44e3-8ab7-a1ebcbdf9dca · outbound

This paper cites Id-free not risk-free: Llm-powered agents unveil risks in id- free recommender systems.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Id-free not risk-free: Llm-powered agents unveil risks in id- free recommender systems

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.814834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:a4f5cba28cf9fcd9e36bb5f16817e0556090480b5afeff965c6a840383ba553b

Observation 8b804b5b-c66a-4187-bc16-60ab154bb62a · outbound

This paper cites MemRec: Collaborative Memory-Augmented Agentic Recommender System.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems MemRec: Collaborative Memory-Augmented Agentic Recommender System

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.160365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:f1b71a948fae4a8b16b69c5f374545c6b582c798052240fd68e04f6d38899465

Observation 8bc6d331-a311-48a7-bd55-f96ea70c9114 · outbound

This paper cites Agentcf: Collaborative learning with autonomous language agents for recommender systems.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Agentcf: Collaborative learning with autonomous language agents for recommender systems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.821919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:be32308c9e99887d391ae94a3cc75200e6e29bbe171560eb75a3f32b9b602260

Observation 67d6453c-12cb-4168-a772-dfbe4c7c8125 · outbound

This paper cites Agentcf++: Memory-enhanced llm-based agents for popularity-aware cross-domain recommendations.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Agentcf++: Memory-enhanced llm-based agents for popularity-aware cross-domain recommendations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:31:55.798096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:bb54d97c540e68eb352cca66758e227871795b81f359a5b812b460d1d1733bf7

Observation c6c01178-2781-4371-a4e7-aa3f7f6f528f · outbound

This paper cites Multi-agent collaborative filtering: Orchestrating users and items for agentic rec- ommendations.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Multi-agent collaborative filtering: Orchestrating users and items for agentic rec- ommendations

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.232542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:d6a1593250fc479b18e93ec6a09814937aa76b77265659100d40caafc701e3eb

Observation 22b2884f-a43b-4317-89cc-049ce1ff0cfb · outbound

This paper cites Recnet: Self-evolving preference propagation for agentic recommender systems.arXiv preprint arXiv:2601.21609.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems Recnet: Self-evolving preference propagation for agentic recommender systems.arXiv preprint arXiv:2601.21609

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.204406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:d6a6e8714c3c7053e0fb403da9ba3c6a8872d17a3547eb692b34927993872719

Observation aa7799e6-5e29-4d17-b41e-e2cf8e40e283 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:56:02.187439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:24:28.048403Z digest=sha256:29f3e6fe2799821ccaaa89faa0b2b342afc668f31be6b45f68c0e9c49c25e2e7

Pith citing papers

Observation 5ba483d8-be47-44fc-a746-f91a8801b4f8 · inbound

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning cites this paper.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.617211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.617211Z digest=sha256:40ee6fec1a31ac6360a274cc88f14e831b1f9c39ec2392bf599cc0cbc8e964e7