Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:04.438910Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.00539.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:04.438910Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:42:59.909229Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T10:48:12.945063Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80ad7cea-d893-42d4-8fb8-8d1dc96e9c22 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9620f307-85c9-438c-ac40-08993cec4f70 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Large Language Model Agent: A Survey on Methodology, Applications and Challenges
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08dc02f7-5467-446a-b50b-0c95824aed5b · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Cognitive architec- tures for language agents.Transactions on Machine Learning Research, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26122e5f-d53f-416d-afdd-296c44456f03 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c161d582-e1ed-4dfe-80a8-3d1af98e887f · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b3b0c7d-8cac-47ed-820c-658eff6bda9e · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation ScienceWorld: Is your Agent Smarter than a 5th Grader?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbe73be-554d-47b1-a061-01188b55cced · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a042abfe-c539-4c27-a6c1-e58591cbdc71 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75ecb96-8246-4666-9bf4-f38c158d1f4f · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Evaluating language model agency through negotiations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad64771c-0f22-46dd-96ec-5627d4093b27 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b68e55-dc53-4d56-9ebf-81317dd519d4 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a26fead-1325-49cd-8a15-53c0823dec83 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0282fc0-9feb-4ac3-a30a-0f6c441ec3e6 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64dda7ae-cab9-4447-b0e6-144f1c6b5289 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acfa2ab7-d82e-4c26-a9d1-1892540ca541 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Selfgoal: Your language agents already know how to achieve high-level goals
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 60fadc7d-c7a7-4280-a143-c016ad995460 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f102a7-b5ef-4855-98d4-bbc843c64d73 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Webshop: Towards scalable real-world web interaction with grounded language agents, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8385a51d-7781-4cde-b779-bdf200c5f991 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Deal or No Deal? End-to-End Learning for Negotiation Dialogues
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05755954-fad1-44ff-9e02-6556f7349060 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Glee: A unified framework and benchmark for language-based economic environments.arXiv preprint arXiv:2410.05254, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b389c14e-e26e-4d02-9c71-a9b0eadee153 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125d7087-241d-4a98-9ab3-b3ac2c9313c8 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b7ff850-89b1-49fd-98db-56015d0e1fab · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11531a67-1dfa-4331-ad07-68535afe723d · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Williams
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f549245-1f64-40f6-a229-01c395381c2a · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Hierarchical grouping to optimize an objective function.Journal of the American statistical association, 58(301):236–244, 1963
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c052aa-133b-453d-be09-6464489841e1 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b6e2d0-369f-4eb5-a5e5-42983e1c28e9 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e1c9e0-5ab3-441d-a0f4-7b0c98d53ea7 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation AvalonBench: Evaluating LLMs Playing the Game of Avalon
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2858b76-b913-4b2c-b2fd-2b24a6e29b11 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Self-playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bbd1ebb1-33ab-4190-9d6e-d788077e060c · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation GameEval: Evaluating LLMs on Conversational Games
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c288d72-2f51-4edd-9cb1-e6f6d4a8f2c4 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Least squares quantization in pcm.IEEE transactions on information theory, 28(2):129–137, 1982
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb4b8319-ba75-4af0-8f3e-c49a25c75d6b · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation A density-based algorithm for discovering clusters in large spatial databases with noise
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc0febd7-dfd3-4257-b303-d6d35b337bc1 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6ad255-2e81-4393-a49a-7658872f2539 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Direct preference optimization: Your language model is secretly a reward model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25cb6302-961a-4c3d-8c72-d1278b151cd9 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation KTO: Model Alignment as Prospect Theoretic Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6833c2ab-c860-44eb-bb1d-7b5757caf38d · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d10df5-5527-4490-9578-e69f4bc0dc85 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1674156-629d-4af9-9f73-e7dcc6888ff6 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Methods of Hierarchical Clustering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00333b33-ec03-47de-b745-e3d1bd363535 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Silhouettes: a graphical aid to the interpretation and validation of cluster analysis.Journal of computational and applied mathematics, 20:53–65, 1987
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa50cd92-6e53-479a-ae57-4bd157a20113 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation A dendrite method for cluster analysis.Communications in Statistics-theory and Methods, 3(1):1–27, 1974
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca82fdc-9a2f-471e-bf1b-6e772530ab90 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation A cluster separation measure.IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 2009
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3285c6a9-3962-4f20-a7e0-e68f3db952a8 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3cee0ba5-e240-456c-a8e8-3aaa2c5bf7a5 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Training agents by reinforcing reasoning, 2025
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f040559d-b41b-48ae-92a5-ffc0336ca356 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Glee: A unified framework and benchmark for language-based economic environments, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3865c1e6-fa12-4026-a4bd-f08b25df12c6 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Llama 3 model card
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 459696a0-c5ae-4e1e-9e10-1c9e624a0749 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Text and code embeddings by contrastive pre-training, 2022
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee2879c-ef7a-4d54-9b9e-abc74446ea9f · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Gpt-4 technical report, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae6b85f-644d-4cc7-a598-cc367b24b11c · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Introducing claude 2.1, Nov 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 869581f2-df74-4623-807b-0e43b3fb49e7 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7612ea95-a64f-4ab3-8ca7-c8bef3c4fd24 · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Qwen2.5: A party of foundation models, September 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91bc5c2e-1b1a-40d5-9ef2-e5bd3ba6696e · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fddb8384-c1c3-43e1-8a42-afdff67bd4ba · outbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3128c10-4099-4755-8cb2-2203ceb52e11 · inbound
Unsupervised Learning for the Elementary Shortest Path Problem ARIA: Training Language Agents with Intention-Driven Reward Aggregation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d425eb3-e2b6-43f7-87c1-d24c0a9c6d16 · inbound
Latent Action Reparameterization for Efficient Agent Inference ARIA: Training Language Agents with Intention-Driven Reward Aggregation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.