Pith. sign in

Paper Citation Record · LEDGER

SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2503.15478.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15478 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:31.534331Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.712764Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1ea8e37-40a5-45a3-83d0-29338277be84 · inbound

MARFT: Multi-Agent Reinforcement Fine-Tuning cites this paper.

MARFT: Multi-Agent Reinforcement Fine-Tuning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.534331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.534331Z digest=sha256:4e4d0835e2ea539f3210087985a8337e58be38c12f112200a677455a8ca0d394

Observation 39e00457-fde6-4e61-8f4e-5047a1735199 · inbound

mrCAD: Multimodal Refinement of Computer-aided Designs cites this paper.

mrCAD: Multimodal Refinement of Computer-aided Designs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:39:22.099377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:39:22.099377Z digest=sha256:91914183bcbeebcac308f8e156372151c75a3d6b113aaf24aaf3b25b0a9d163c

Observation 0ffbc057-70a4-4ef2-8c40-058532d0b2b9 · inbound

GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection cites this paper.

GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:32:23.086502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:32:23.086502Z digest=sha256:1bafcb513d73e1c7f05ed29b1e72988af7cd10e824ce0559354d934538bfe89d

Observation 4ce1d644-c834-4bdc-b210-d44b430e6f02 · inbound

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning cites this paper.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.744073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.744073Z digest=sha256:f150599b0370a14adecd21d0523735dde083cd8cb6e0ac37eff0eececb9cb89f

Observation 14a9e272-3755-4f82-8ebb-4c5d0b603de9 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.255061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.255061Z digest=sha256:0e093ae9dff4cfdaedacf16c4c2772e63f874bb97f893b551fa12a3618a7f4fb

Observation b169f929-fb11-4c83-8be7-a284473d8032 · inbound

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay cites this paper.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.796768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.796768Z digest=sha256:cc8b3b32cc1b4bf66e463e414d7b4013fdb7f983a711099388dec835fe27ea3c

Observation 73ce5631-804b-42e7-a0dc-7ac9d5f4e945 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:31.351435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:31.351435Z digest=sha256:da9eec16b3ecb832be7dd9e96cbef1b92bffa2458be4caedf81f2ce1784c77c4

Observation b6dca655-ce83-4790-b963-216ffd7ff111 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:59.729856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:59.729856Z digest=sha256:4d40ecc19d2a1112a1ae00e3ead340788a1c655dc875cd1eb29678e8e776a0a2

Observation 1454204d-6f50-4207-9c68-4aac41fc2c92 · inbound

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation cites this paper.

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:35.042447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:35.042447Z digest=sha256:e540f7b43903a8428b9c0c8166e051717c0dccfca4a779a66f36575c42efe3d1

Observation 4f807b02-017d-4bed-a1be-c32a4eab9262 · inbound

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence cites this paper.

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:52.419144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:52.419144Z digest=sha256:94e1d0bc9abe98e0afce5648baac29d5a2afa878a63e0a7b790ba694ce03c509

Observation 125d7087-241d-4a98-9ab3-b3ac2c9313c8 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.135410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.135410Z digest=sha256:5d8ded66542ae379fb7a38c9114fb4547e93c11cff3c4256db7d52b7e08cecf4

Observation 17e33e8c-9b01-4cc9-a34c-5da8a28ac835 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:58.657810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:58.657810Z digest=sha256:8be5d1226c817668afbf8c65b795a4c656ff83fb1bf8b9c38726d1e3a81ab000

Observation 3d37a5d8-7487-4031-94a2-c63c0b08b245 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.314881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.314881Z digest=sha256:a3bc4e5b55646d82b762307a875a78fa58ef114dca29e687e9ce94ab17c91725

Observation 7625a69a-a765-439d-8c35-df5ee4ddd1c3 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:39.160869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:39.160869Z digest=sha256:8623ba85f2f5729a0d6fc3f39bc6d0e41434c43d33ed61455e2fdcf6068f47e3

Observation 1f0dc4f4-f5e2-4c7e-91bb-808ff5e823e3 · inbound

Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching cites this paper.

Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:45.270627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:45.270627Z digest=sha256:2e1c69e880ac88ef62a20c0db3cfdad72bec987e703ab2ae8f04016798186397

Observation 633bb555-fc2c-4003-9188-8f313ed67ce6 · inbound

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs cites this paper.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:13:00.021058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:13:00.021058Z digest=sha256:745f65606af3dba193318f6bcf9b82eab56578d917f2303aed3255699edc5364

Observation dd58bca8-26d4-4d0c-b6ab-e0f0866f15b4 · inbound

General Modular Harness for LLM Agents in Multi-Turn Gaming Environments cites this paper.

General Modular Harness for LLM Agents in Multi-Turn Gaming Environments SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:22.785205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:22.785205Z digest=sha256:5770af5bc05f898da700a036cdb102844ef14965bed048a55998a2b3ff529e51

Observation c738fb90-9627-439f-a3ee-726a5cca853c · inbound

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL cites this paper.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.363452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.363452Z digest=sha256:dad144bcc1863e6151c7ae9dab38188225b15a832533bea2b4c1c259d1bcf4bf

Observation aa65f488-02e3-49d7-8feb-539207c4b1eb · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.728989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:4c91f06c9110ea98b62049d0825c34a0daa5ecab1bdcb70db45495f6db90510d

Observation df998f68-4a2b-44f7-8386-e11527426948 · inbound

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs cites this paper.

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.226541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T05:29:46.136115Z digest=sha256:9a44923dddf436a410bbc1125602de11c90517d59b98784f3e53fc9bc561a72f

Observation b182a79a-c432-42dd-ab04-819021cdafe4 · inbound

InteractComp: Evaluating Search Agents With Ambiguous Queries cites this paper.

InteractComp: Evaluating Search Agents With Ambiguous Queries SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:45:00.354038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:45:00.354038Z digest=sha256:7746d53a0e7635af4ddefeebadd131cc44cab00bb9b37b823fb0eb89c75f0568

Observation 43c7aad6-49c0-49e6-906c-976e2360c91e · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:05:39.034862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:27639b0d85630bfef39ef825f341880310f389a4db71d3649e2327bfd3aeae55

Observation 1a6179a0-be0e-4fec-a489-4edb3cfde94d · inbound

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction cites this paper.

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T14:34:51.341451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:34:51.341451Z digest=sha256:e9e88fdb5519cee29d7f260b9a930d71e223c6bacc74ef876cb79073df369ec3

Observation d137f4c8-ea46-414d-83e9-9e000525dd33 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.300750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:99047255de8a7e6a28aac5ed88fa5c2e6fd08e50a74b8befe1c65ad0acff65a7

Observation 1a126558-2aa2-41fa-822f-bb3b714ebddb · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:22:31.136212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:89537fc625766808cb2d944097e38dcffc23bca4dad20a8d6794b0f67c950723

Observation 04ab4af9-3fd6-40cc-b505-75a93ffb9a05 · inbound

Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration cites this paper.

Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.529146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T19:51:50.503945Z digest=sha256:64982f7ac26bdeaa9889ef88e66ca07a79eb08884f4bb97cf5993e8cc92e6113

Observation f3d42dab-7dce-4859-ac9e-c1ed570faaeb · inbound

ActivityEditor: Learning to Synthesize Physically Valid Human Mobility cites this paper.

ActivityEditor: Learning to Synthesize Physically Valid Human Mobility SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:50.000162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:42:36.437661Z digest=sha256:e3600b1dffa33a62900de93c51093d1504d3dc5e377fc1f5ac7b9c883d9d222b

Observation 26c3d4ed-4455-4ae1-9410-e2bf0424da5c · inbound

PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent cites this paper.

PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:06:11.230632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:17:59.156121Z digest=sha256:a0269c7c4ae3cca21f4056201c873684de70251a152a27b417920aafdaff200f

Observation a0fdffef-ee8e-439a-865e-3d2bcf9dfcf6 · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.471183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:17b72ffb64c2e82e0e125dada2cea2bb7affae544283468cd04e23c96adcce21

Observation 11058383-d232-4920-a4e6-e116d4043f25 · inbound

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators cites this paper.

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.049413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T00:51:57.796883Z digest=sha256:c6f65fdb4683d90aa5a313596893b53ce8955ecf83f53ac082c76f074c0de4ea

Observation 3390848c-dd81-46ca-98fc-9a18b501fb71 · inbound

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators cites this paper.

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T14:37:24.690196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:37:24.690196Z digest=sha256:37ac9f7e2a30ce1c54c0861da86ae370a75fde1a244c02c0228741f961c210ac

Observation b9ad27e3-ca85-4d65-96da-dcefaabc01cb · inbound

Step Rejection Fine-Tuning: A Practical Distillation Recipe cites this paper.

Step Rejection Fine-Tuning: A Practical Distillation Recipe SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:06:35.455960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T03:44:33.525966Z digest=sha256:04ff7a971c393c11bdc7b99ceb8692f2e7ee060b5f48fa1e460bfd456360f839

Observation 754fb7ba-0268-41af-9679-6dd5bcb6f288 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.469797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T10:41:25.205368Z digest=sha256:656aeb5311667e38d9725960b6386332780f07f9dfe65808c31635d8c8c3be35

Observation 7e2d00a5-471e-438c-94c9-204a55df7b01 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:32.483381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:32.483381Z digest=sha256:c4544071cb6c849afe52605b863c3f85957aa22514626b345b03ecee2c7a02fa

Observation f5193e78-721e-4d63-849f-01c0f30ae694 · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.521037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T05:31:37.171422Z digest=sha256:1a96c44cd5e7449c7a3dab3f71d27b3a8ea999432217838c6237020ac582b8be

Observation 8d6e6666-2b71-4cb6-b8a8-6c70dc737ad5 · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.714454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T17:31:11.074231Z digest=sha256:b6f3ef9be2aac5dccbe552e8a0bf2a977ac3cc4a341c7603cf9f9c67897e70b7

Observation 2e6c0734-d37c-44a7-ab34-e99e658384e5 · inbound

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use cites this paper.

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:36:55.011277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:36:55.011277Z digest=sha256:8fa54e55f237ec38a87f71ebd45cc574432bf4bc6609f4f58f4abc22089ae20e

Observation fb439d78-5096-4625-ab82-618defb3f50a · inbound

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning cites this paper.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.431790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.431790Z digest=sha256:1a75e4647f9fb14dea2d8881e678ccc3e35a58df2fc99de3ea1a957406006950

Observation 0188d5a1-8c76-42b0-9bc7-0a7d4b590917 · inbound

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning cites this paper.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.291734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.291734Z digest=sha256:bb2d06c7782ee2df440bec444255a2c37f8446531aff01f94281ff1368112b15

Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.690152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.690152Z digest=sha256:9cc292e24f30bfb99010d734ac18b7697300eda671cd7435d694eeb22c190ffe

Observation 2c457f6f-b581-4af4-ac77-d68b18bbca7f · inbound

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL cites this paper.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.693596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.693596Z digest=sha256:f56f6de8a5f91c64710160173914e179676e597c54e10b57319d6fda2c809ca0