Pith. sign in

Paper Citation Record · LEDGER

ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2402.19446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19446 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:39:37.983121Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.549262Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2372b857-3786-4fdf-a12a-237a24cce147 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.455408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:68d7ef68a7e84f5839d45b84a291ce185ccd8a101f29aebd3d87cea0f4bb1cc7

Observation 8b5a8aa2-19e3-4392-a33e-eb36ad3f05c5 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.461524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:9eea2a505f863303a27f943c2622336cc48ea4eeef8262e0bcff38b324d7d96a

Observation 11484c73-dde9-4c93-aa03-456ef3ded975 · inbound

Process Reward Models for LLM Agents: Practical Framework and Directions cites this paper.

Process Reward Models for LLM Agents: Practical Framework and Directions ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.983121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.983121Z digest=sha256:dc84bb88923e2f402f9db0ec21b9b26642fa85a4e4a8b6b112b261e1691262ff

Observation e5c53f99-5b8a-4238-9ef7-6295b93fd574 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:59.670045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:59.670045Z digest=sha256:550c4ce6fc9ff91fdcf82500595c3ef7709f6c6360615f40667a24055d90dc85

Observation ac711189-9f6d-43d7-9184-727b8818fe3f · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.408981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.408981Z digest=sha256:caf0883922b9fffef21cadf5b884f5ec72e5eaae1bd32a4d9ae3822a7aaf260f

Observation b389c14e-e26e-4d02-9c71-a9b0eadee153 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.985006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.985006Z digest=sha256:c8d85068d854679dfde8dd43c04e7f03d4782f5d20881eaeffde9f5cb1a1e804

Observation fe81ee85-bce5-4ea0-9666-744b80f639c5 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:58.654310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:58.654310Z digest=sha256:561b52c1dd0f537840b9ee60ec1bd03a6f056e6518e7712e7a339b44d879547f

Observation f04a9702-4b69-44c0-ac3b-86915a75c762 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.549942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.549942Z digest=sha256:b8ebd4f6c20d86927bb3e42e58252df673d7c54c422579f964e30bc03b739c10

Observation 630fbc79-01bf-4030-8d21-19eaea262789 · inbound

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback cites this paper.

Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:30.382349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:30.382349Z digest=sha256:bf8169385cb8f3ba4d28cfd2587bdb734049131e7928697f7886cdf75f7e74e6

Observation e863c26d-93fc-486a-968c-e70c501dd59d · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.356815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.356815Z digest=sha256:80415a9320200087e8d53023f9e39b0201494f2c143d23895260c5411f3fb1ea

Observation 42837901-c731-4b89-8ada-6a3b37bb2e11 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:39.085675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:39.085675Z digest=sha256:07976a443b1fa041cf25ddaadefa27e16909dc820ba53d249baaa7aaaf46fadb

Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · inbound

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison cites this paper.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.960285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.960285Z digest=sha256:dff4c262a54f1af621d71d4ff8504966f41c31ce276c215c0e6f67d21ed710a3

Observation 6de5f99d-b149-450f-a171-fb3dcc223bb0 · inbound

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning cites this paper.

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:04.049831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:14:04.049831Z digest=sha256:4d70c9bf308ac112e2004a4092679eaa5e854e2308c04ee1c84eba5b55b5fef8

Observation b1682c18-d074-4e3e-ac43-0408273afa93 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:51.017069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:51.017069Z digest=sha256:dad99a5097c63138b5a31fbea2812596ae2b3417dfaf7547f95334e53927b34d

Observation 535dcb56-38ea-4662-8ec2-dd8827cf14c3 · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.947042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.947042Z digest=sha256:e7894cad60341d31fe16f7be19ef6945b9e2f67d207bddf0e578f20ac75d31f5

Observation bc9e16b3-d475-4011-8cee-c3b82b527ec1 · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.745515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.745515Z digest=sha256:ef3c6e52a0c9dd33f8e22c1dc478a5a92babc4c26df688960a4716cd140a44da

Observation 95e6a819-f23e-479c-a18a-da58ad239fed · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.014508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:05873ab61c96429748752f72824ddb4480288b5971c4961f0951fb57ba46b801

Observation 46186c15-5bed-4455-bc3b-dccfdcf18c18 · inbound

From History to State: Constant-Context Skill Learning for LLM Agents cites this paper.

From History to State: Constant-Context Skill Learning for LLM Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:05.529590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:50:43.547830Z digest=sha256:cb2d70bc3bd72379cab3b1f651daf9117f1309bcb873933a7d9fa4131a201d1d

Observation e1984de5-e929-4a2b-94e2-83467b69203e · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:12.171337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:364da64f50e8655f6e3b00d36d35fe5b5f46d71c074222a8fdcb9cbac3110ad4

Observation 830e555e-e74d-40ef-84fb-4290fa4652ad · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.475555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:5c9512bc414cb6f820572373375424461c5fe990ce09d8a5110e3c086535f614

Observation 761b8ce8-90d4-4dfd-b99f-1e1288696ce0 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.153287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.153287Z digest=sha256:91f2ba0220a38730ac037b3004779323a042c9edce043fbcbcdeb66a1d7912fb

Observation 4ac2a3c3-7399-4a83-aacd-b7eedd783c24 · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.483436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:98f2e4bc3511b64cb48d18c939e1af53e6fec9d360a2dd082a239e61f43b7b3c

Observation 300a6290-65b5-4edd-8337-474164b0d212 · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.495072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:31:37.171422Z digest=sha256:443b88a863016c3e46157c3c29a5e35781677130e046c9df2eff4947fb97584d

Observation 1d4d9da5-573f-441a-a56d-71f2d1ae5eec · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.708669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:31:11.074231Z digest=sha256:25da3a7a84e9b0f772cba5ff8295d48b2616ac1dd017396cf11451e6fb107067

Observation a3d0c357-7b0e-47d6-ba3f-8aa4ff6a0155 · inbound

ECHO: Terminal Agents Learn World Models for Free cites this paper.

ECHO: Terminal Agents Learn World Models for Free ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:04:46.491807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T14:57:03.095107Z digest=sha256:88e2d7dfa057d3a149ad5618e8de78191880a23b10f88fe7ded9cebceec4e303

Observation ecc25546-622a-4b51-95e1-5097794fe2d9 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.296998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:51100a55d1bb7c72464c9c5d84da173d1b2eee79e2c6eb39b14599e897c8756c

Observation e73b26fe-012c-41a7-8cf0-a8486627ae88 · inbound

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training cites this paper.

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:23.950721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:33:00.408984Z digest=sha256:fc451ca4a717ba901af2e819b2e9c1d7501ba7113ffd3d6745e51f49581265ba

Observation d7de7e4d-e539-468c-b87b-9ffc62e5ab04 · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.932501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:fe4ea1a9e27259f16749796e62b626c350c5959827bc0b7ce723be3ba06ffb3f

Observation e4197b60-6ae7-4740-8e5a-28a05289d8a5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.601656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:fb20931d8922e7c2842d5b183294a1fb13920cc7f33ed62e3d11e416a49afecd

Observation 0e1bac64-edce-4f51-a4c1-d4b607991b0a · inbound

Diagnosing Task Insensitivity in Language Agents cites this paper.

Diagnosing Task Insensitivity in Language Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:51.551580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T04:58:11.929896Z digest=sha256:f048ae56d01d5e8684130d31974b944745982b66a4cdaf78c86efdcd2bfa75f3

Observation 6d680464-8def-4b55-8d3b-4975436da8a8 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 123

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:50d1992fc9b707286a8f988ae91f5216deb47dd3bd2d57fd3611e0fa96bad5bd

Observation ef2ce5bc-eb45-405c-aec1-288c1d48ea6b · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:35.709395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:35.709395Z digest=sha256:6dfd950aaa218606ed43049fc70585fda17a5914208f2247ba324d16ff834385

Observation 7d9fba77-4255-4205-b1ba-41b57be0a765 · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:23.548585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:23.548585Z digest=sha256:da04e3304ee51ae0c7be2e9a87c4f77f29e283c503c6714006401ff9b068ce28

Observation 1d592169-b77b-4224-8570-132ba46210a4 · inbound

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning cites this paper.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.221262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.221262Z digest=sha256:8044973a062b342e4ad3b3c4f7a151b4019c838804f748987566889407b65e51

Observation 028e7780-fa07-48ae-a71e-1085e1c681df · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.223723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.223723Z digest=sha256:6b152fff8f335abe1b26b3d2f68cc3189b4944ea1ff4905b2dc6127aa467761f

Observation 55ea2611-4fe6-4363-b476-5a5e49756d56 · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:15.761164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:15.761164Z digest=sha256:5a1ec514043e857d9f4e5d8f4966508c7e5210f1472e29d7943e9f2825098f3e