Pith. sign in

Paper Citation Record · LEDGER

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2509.08721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08721 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:11:38.252026Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T07:40:35.113858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T07:41:14.649759Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04965135-5b5a-4c16-b64c-7e8a866b9a66 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.023666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.023666Z digest=sha256:47640082942877ddfa6a807f8e1ac2362ef1e8bbd2a81e5da8d96c07ed0f45ba

Observation 28f31821-5267-445e-be6d-d62de77433d4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.029405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.029405Z digest=sha256:1cd6b806f36bcb51771648a2d0bec6f8800d900da62c2a390d75af286ff37272

Observation 93dd3fd5-e281-4625-bf24-5b7927551f3a · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.035959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.035959Z digest=sha256:3a653bc9f101bfe87836cdb7f6fb7676f6fe0f1af027921dcff7a9050989604b

Observation e6c60293-83df-4036-8035-987740f8026c · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.042144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.042144Z digest=sha256:63de1065a30403c1ef2e021fe49a9adbf4a4ba19c792328ad25d9ed0b39a8b84

Observation d6810192-2a98-4fcb-ada7-0e1d88d533be · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.047737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.047737Z digest=sha256:06dfee6df370dc593ec2ef601cbec7d1ddae66c263e5ff5ae4a9bfd2e1d796e1

Observation 0dc77fef-8e5d-4786-a540-594480ad10e0 · outbound

This paper cites Introducing rl swarm’s new backend: Genrl.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Introducing rl swarm’s new backend: Genrl

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.271650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.053251Z digest=sha256:36401076e2fa7de321ead0081ee8f87b83268b3d2a6dedd7d892809fb2247fd5

Observation 269bfd1b-dbed-46e9-81af-3aacf9f19b3d · outbound

This paper cites Gensyn rl swarm.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Gensyn rl swarm

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.255993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.058693Z digest=sha256:81b59b1ad3f5dba11a615ccab827b9784e1eaf11c4d5b9392d643b94eac5b3ef

Observation d71e7bb5-96ed-4c46-9e88-5fcde5ef24ec · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.064488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.064488Z digest=sha256:d5004c640fa23d0c2bc37a70719c4aa53c2d3510c46ab3898ed494bb7fe6bc2b

Observation a46e0d65-e37e-4aa3-8edc-0ecfa851d540 · outbound

This paper cites Bowman, Tim Rockt\" a schel, and Ethan Perez.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Bowman, Tim Rockt\" a schel, and Ethan Perez

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.239713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.071245Z digest=sha256:01a52e52edc3f16677c5ff323547b68541d6d5a98ed3b720f9ce97c10d325749

Observation 0c69c874-c4a3-4b7e-9290-54a0a856b72d · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.076287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.076287Z digest=sha256:2f546f69d97ec29200c91974c3bded2c3fc617468d0ced7a51d7a0aab296c2df

Observation eae964b9-d577-476e-948f-b091cd0b449e · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Coderl: Mastering code generation through pretrained models and deep reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.082881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.082881Z digest=sha256:d8ceb7f42eadba80e2426a68b917683b1fb103efc16bafbf130fac982815967f

Observation 704f084e-e67f-4ce9-a5ef-2fca2056a526 · outbound

This paper cites Camel: Communicative agents for "mind" exploration of large language model society.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Camel: Communicative agents for "mind" exploration of large language model society

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.223623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.087984Z digest=sha256:d95e4e993c7c1b20b101e62d2dae035b079f96c0dec040e494df715d10ae04ed

Observation 5be3c361-0515-4534-9677-6fc7350b755b · outbound

This paper cites Improving Multi-Agent Debate with Sparse Communication Topology.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Improving Multi-Agent Debate with Sparse Communication Topology

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.093375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.093375Z digest=sha256:2330e6b530b6fa58cfc8c6b2ec048bdd4eb6dc1e2bb534b86f4df21b0538fd08

Observation ef425cb0-e61f-4267-b996-c693242287e0 · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.098724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.098724Z digest=sha256:613f0791eea04a038b4c0afeb364faee444c3e9bd4457ea8b6dfd5ee3953baae

Observation 42e39dd7-55d2-400f-aab7-650c5e6df32d · outbound

This paper cites MARFT: Multi-Agent Reinforcement Fine-Tuning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.103829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.103829Z digest=sha256:321844b48cadb85bb20d9ff4644ebdb64b182ac757b6bb1f17f6a6be5331e3da

Observation b55754cc-3787-4499-a329-4b63d9320414 · outbound

This paper cites Llm collaboration with multi-agent reinforcement learning, 2025.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Llm collaboration with multi-agent reinforcement learning, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.109179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.109179Z digest=sha256:c3e280e11107b41f897347d8d60c295d61da863d74f509f9da0c4516543f984c

Observation fa691390-2f9f-4629-b252-83b5634efa0d · outbound

This paper cites Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Coevolving with the other you: Fine-tuning llm with sequential cooperative multi-agent reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.206463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.116384Z digest=sha256:e833f986f6534ac0b568502accc533c04eb88810943e65eabfc14fb608f59465

Observation ba52a9bb-f820-45e3-a236-89c70f7539ee · outbound

This paper cites Magistral.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Magistral

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.122268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.122268Z digest=sha256:d6eba957db54a106142340bffa81bcd67e0060fa51f07ffbd53660804fa1adde

Observation 3afbe756-0010-4c42-9b30-c292b97b8845 · outbound

This paper cites an unresolved cited work.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.127762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.127762Z digest=sha256:8e2f52ad7be12cad19acffac67fe50c65388adbccd173a52c8beabd5ad9df7d5

Observation d70bd771-4080-42a6-aafc-d4ffcc97d08f · outbound

This paper cites A Survey of Small Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing A Survey of Small Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.132668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.132668Z digest=sha256:fdd11c0d850320588054c48c50bba60ff32eabdcf3d0ba16595f887bd7e59cfa

Observation a6475cc0-fe77-43e3-9e92-4992aa9c8d6d · outbound

This paper cites Aligning language models to follow instructions.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Aligning language models to follow instructions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.137824Z digest=sha256:a22d608f40a4567a23eef345030e2cb6307363893c1208a4f9142a3c235bb2bc

Observation b1967b71-42e4-4d3f-8704-939d8f2cd846 · outbound

This paper cites Learning to reason with llms.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Learning to reason with llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.160352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.143047Z digest=sha256:6520d79a76a077e1199f23984d1abcdedb1c90fe01a2b128ed9487ef413dd673

Observation 827aaa37-c583-4756-b3cb-817dc88ce216 · outbound

This paper cites Training language models to follow instructions with human feedback.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.147977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.147977Z digest=sha256:276100382edd618634fb27a990d5b103fc75e32f8c7f96901740ad2038a9c9ba

Observation ad40bdd9-db83-4635-ab1b-fb9e691d24f2 · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.152836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.152836Z digest=sha256:56a118d5b34a337b6456db36d1e81303ba71acaa507907d1f463ee700135aa10

Observation d045c633-4a48-4240-b5f8-3e00400c50c2 · outbound

This paper cites Red Teaming Language Models with Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Red Teaming Language Models with Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.158686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.158686Z digest=sha256:1b2e6bdd00131c98403f733c7f711e27e923000c1e839470bca9e49018a3fe14

Observation 5a97264f-265c-4f92-852d-c6e3ed522a5a · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Qwen2.5: A party of foundation models, September 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.164045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.164045Z digest=sha256:29dd76d69e7e27a1f47c5a65fcb18d6a0a72ade80057516d9adad58a164a9e95

Observation 8febbaa1-a426-41cb-9ff9-5e68eb4c2459 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.169427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.169427Z digest=sha256:ee95fdacbb719548a72c0708d57597bfd1f3eca59002ec9267586424fb6c5091

Observation 9d47a817-b740-4b5b-9a05-0c50a2b5f7ed · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.175033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.175033Z digest=sha256:258891099a9be1977f9c84db30fa256e622cfce32f13675b084d026de3ecbd2a

Observation 21f9b2df-4f6a-43d8-9b18-9758b989a4a4 · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.180073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.180073Z digest=sha256:f4d34e26ce12c800ab7a1e094ea01cc0a090c25927e18a3f0df652c539b7987d

Observation b6b384e7-e0b4-4da2-ae8b-8f5aa93894fc · outbound

This paper cites Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.192344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.192344Z digest=sha256:2e968afd8a82109d0aa8e9118ed656f10c10c807de30642145fa93b1a26dc897

Observation 1add689b-884f-42f3-88f8-c32f9affdb43 · outbound

This paper cites Fine-tuning language models for factuality.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Fine-tuning language models for factuality

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:11:39.119425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.197585Z digest=sha256:c2ba0dd5e4861072be440286bbb0f3d2c3aed2b98ba2c6cdf724da49757a8261

Observation 875b28fa-e440-4d6d-bbee-76b50c82aa86 · outbound

This paper cites LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.202370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.202370Z digest=sha256:6cfa662ff647f7acc9a7b21c062b61e6b47178f5445e51d30fe746c802d3fc98

Observation 83f1c6ee-957a-4375-85b5-6dba009facdf · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.207162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.207162Z digest=sha256:6161aa9019060f1d3553aee6607b0e2c5c628ee67ae36e90c8ee0ca1a39f2b10

Observation 9fb83e6f-e9ff-4ebf-83be-ad000124da5c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.212687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.212687Z digest=sha256:63cc1f227b550daca56039af157d86e72685161159a5ad859ff2e0d3ea79ad00

Observation 43ee3926-cce7-4f2f-82a2-6936c1ba95ef · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.218726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.218726Z digest=sha256:e07897efd360543a7e8fb713ffdf8a74daecad569973acc7d9f1571090a20b89

Observation 151856ee-cfed-4593-950e-6a8c2fb8bf71 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.223789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.223789Z digest=sha256:aa2b6d3c7a0e06d48b47fe54ba55c619a8a046660a6784054256f4a7fbe68881

Observation 2d5ad3ed-b20d-409a-9eed-2dff7c548699 · outbound

This paper cites SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.229066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.229066Z digest=sha256:513a60ef0bec8e6121b04080362b7aa9706125903fddec6d2002fef1608ee1e5

Observation 324bde60-eac4-4c01-8fcc-467217f76eda · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Fine-Tuning Language Models from Human Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.234186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.234186Z digest=sha256:195170055f93e6d042d98a33736d776c0fd69ceadecc521e0c3fca70f36dee8b

Observation cc2580f9-6482-49b6-8f05-4750c704c76d · outbound

This paper cites @esa (Ref.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing @esa (Ref

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.239822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.239822Z digest=sha256:e1eb06c8c01da5dd1094cccad2c218d3522220d6861e4faae0a26dfc864aeb17

Observation c599f75f-d57a-4a96-a53b-4c927e5bde66 · outbound

This paper cites an unresolved cited work.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:38.245234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:38.245234Z digest=sha256:173178c5d8b583d18301ca7d338fc998cae963958011ba6ad8b2896ea60b64d5

Observation a5f4ac86-69ee-4962-8a27-59534893ae92 · outbound

This paper cites HDEE: Heterogeneous Domain Expert Ensemble.

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing HDEE: Heterogeneous Domain Expert Ensemble

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:11:38.305753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T16:11:38.252026Z digest=sha256:0fc6f775458aa71bca0bbdfe373a13734f73de83a8191030a0c3c07065f0d4dc

Pith citing papers

Observation eedf2711-0005-4d73-b17b-6169657a0b6f · inbound

Backdoor Attacks on Decentralised Post-Training cites this paper.

Backdoor Attacks on Decentralised Post-Training Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:23:26.221641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:23:22.658691Z digest=sha256:75de8e8ac96c23ac3ebf1c32636274812385f5e21112243c3f4c4a8dc6462caf

Observation 2d9cc47d-bec0-4e20-8335-f2074985d647 · inbound

F-TIS: Harnessing Diverse Models in Collaborative GRPO cites this paper.

F-TIS: Harnessing Diverse Models in Collaborative GRPO Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:41:14.654011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T07:40:35.113858Z digest=sha256:5f797cd343266b96751944c9287d4b4830a07674a166a079967a6b49499b90e3