Pith. sign in

Paper Citation Record · LEDGER

ARIA: Training Language Agents with Intention-Driven Reward Aggregation

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.00539.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00539 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:04.438910Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:42:59.909229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:48:12.945063Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80ad7cea-d893-42d4-8fb8-8d1dc96e9c22 · outbound

This paper cites A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.443310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.443310Z digest=sha256:67f5522bc688414165cd41344c3ca6855ab3da801c3deba914d5b6db5eb18226

Observation 9620f307-85c9-438c-ac40-08993cec4f70 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.562789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.562789Z digest=sha256:fc4429470d2945df9852b7eb651e9c221e2b9bfa490e4af3816f45c18e9930b2

Observation 08dc02f7-5467-446a-b50b-0c95824aed5b · outbound

This paper cites Cognitive architec- tures for language agents.Transactions on Machine Learning Research, 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Cognitive architec- tures for language agents.Transactions on Machine Learning Research, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.722060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.722060Z digest=sha256:b9b7c17a09d731eb3cdfa3a738c19f7cae2f21141c18a3cea812f7a9ce0851e7

Observation 26122e5f-d53f-416d-afdd-296c44456f03 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.885758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.885758Z digest=sha256:caa67525c2b73791fadef8dc57e63dc95b30925ce0e73079e9fe4b4974ccdc94

Observation c161d582-e1ed-4dfe-80a8-3d1af98e887f · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.066328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.066328Z digest=sha256:dbc3901aa956ed5bcec7c68bc9f6889cee099c1150ee1b2b8fc57ae280f997d5

Observation 8b3b0c7d-8cac-47ed-820c-658eff6bda9e · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.239397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.239397Z digest=sha256:43eb39e67e4890ba67c96e29b113cad727cd687bb3e3601300762f0fd8647b5f

Observation 5cbe73be-554d-47b1-a061-01188b55cced · outbound

This paper cites TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:05.090357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:08:59.403455Z digest=sha256:d989b0089246ad75f3115e58b1da810d00c98528d750dcbe041426fb6b0ace7a

Observation a042abfe-c539-4c27-a6c1-e58591cbdc71 · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.541246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.541246Z digest=sha256:d6f6e6a2c763beab8971a9b371449dcf45e8b8cfe8420e2338c8e7dbb95c2fbd

Observation b75ecb96-8246-4666-9bf4-f38c158d1f4f · outbound

This paper cites Evaluating language model agency through negotiations.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Evaluating language model agency through negotiations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.754365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.754365Z digest=sha256:8a07611cd676ed026d7d2faceedd3e1fc6b3ff0b4004d951dcd9290c22d46fd1

Observation ad64771c-0f22-46dd-96ec-5627d4093b27 · outbound

This paper cites How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.948196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.948196Z digest=sha256:da1fbe32504eec2991fba4bb6eba7a6025a5bc7471642253ed5c413f7d39a680

Observation 46b68e55-dc53-4d56-9ebf-81317dd519d4 · outbound

This paper cites CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.107183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.107183Z digest=sha256:19ec5c58c828b72a9e5859070afc47792adfc09f56cd9532abfd552a6a0f39c1

Observation 4a26fead-1325-49cd-8a15-53c0823dec83 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.264859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.264859Z digest=sha256:b23a7a15eb4aab8a895bed234c033be7269fafa086ff5896e96a53b45c14c6aa

Observation b0282fc0-9feb-4ac3-a30a-0f6c441ec3e6 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.382912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.382912Z digest=sha256:60aa9089ba18db268ce77e7995b9755a9faba9dfb8ce4b11c1439f49ae837c13

Observation 64dda7ae-cab9-4447-b0e6-144f1c6b5289 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.454738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.454738Z digest=sha256:e1eee0a5a42f18e062de80de8297f8e1386e2bf6ce11b76bda1d4f1f190f8937

Observation acfa2ab7-d82e-4c26-a9d1-1892540ca541 · outbound

This paper cites Selfgoal: Your language agents already know how to achieve high-level goals.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Selfgoal: Your language agents already know how to achieve high-level goals

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:08.621937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:00.548686Z digest=sha256:53cab84ec23b9ed6727cd21b6fba9922c244807885ac6c6a9ce5969a918e3408

Observation 60fadc7d-c7a7-4280-a143-c016ad995460 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.612952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.612952Z digest=sha256:e19f0b55c05e39af8912a59d594f5d182006453219e8377282a964baf236043d

Observation f8f102a7-b5ef-4855-98d4-bbc843c64d73 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents, 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Webshop: Towards scalable real-world web interaction with grounded language agents, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.701429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.701429Z digest=sha256:2238bfd3518d6da00e3f1210c390813ba730c317d9bb67ab4e8ffb82aea8d253

Observation 8385a51d-7781-4cde-b779-bdf200c5f991 · outbound

This paper cites Deal or No Deal? End-to-End Learning for Negotiation Dialogues.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Deal or No Deal? End-to-End Learning for Negotiation Dialogues

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.787910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.787910Z digest=sha256:13ac04cd39d00494e55dc1d8725258a5ef1e4c1a17540cd1fdcb7eac70927aec

Observation 05755954-fad1-44ff-9e02-6556f7349060 · outbound

This paper cites Glee: A unified framework and benchmark for language-based economic environments.arXiv preprint arXiv:2410.05254, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Glee: A unified framework and benchmark for language-based economic environments.arXiv preprint arXiv:2410.05254, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.857814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.857814Z digest=sha256:570bf39fd21193885d92423ec9c59e73041e42f7fb6afa63ac0001d009c85038

Observation b389c14e-e26e-4d02-9c71-a9b0eadee153 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.985006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.985006Z digest=sha256:c8d85068d854679dfde8dd43c04e7f03d4782f5d20881eaeffde9f5cb1a1e804

Observation 125d7087-241d-4a98-9ab3-b3ac2c9313c8 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.135410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.135410Z digest=sha256:baca21a23b80db70863397a5c62ae9b7c05c77b722ad67a866e125c7830895b0

Observation 7b7ff850-89b1-49fd-98db-56015d0e1fab · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.250499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.250499Z digest=sha256:73274e3ddba690de8d14fa9d61a2b208510a9fbe3837d97978e150b654c6d54f

Observation 11531a67-1dfa-4331-ad07-68535afe723d · outbound

This paper cites Williams.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Williams

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.348373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.348373Z digest=sha256:05953c45d003fab9a508148a13540eba0a08516ebf3a5328367979100ca51571

Observation 2f549245-1f64-40f6-a229-01c395381c2a · outbound

This paper cites Hierarchical grouping to optimize an objective function.Journal of the American statistical association, 58(301):236–244, 1963.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Hierarchical grouping to optimize an objective function.Journal of the American statistical association, 58(301):236–244, 1963

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.453841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.453841Z digest=sha256:8bbb5757576dc4cdb1c34a21eeb733dbcd934b074c2eba98629626e0971fb1b3

Observation 23c052aa-133b-453d-be09-6464489841e1 · outbound

This paper cites SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.633394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.633394Z digest=sha256:e8c4e07600cfb211adf54cbc3b52736f88f32a986c02732d87499105f81c9a28

Observation 29b6e2d0-369f-4eb5-a5e5-42983e1c28e9 · outbound

This paper cites Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.819815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.819815Z digest=sha256:c5a3cf5fea1096fe3609edca543273a65ee5e6ebdbe78969705107c8e1319969

Observation f7e1c9e0-5ab3-441d-a0f4-7b0c98d53ea7 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.061227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.061227Z digest=sha256:b80550b5f4493a3f4cc5581c51e7fcd7134128665e0158f9414ad9e4fe1ac552

Observation d2858b76-b913-4b2c-b2fd-2b24a6e29b11 · outbound

This paper cites Self-playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Self-playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:08.356809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:02.194205Z digest=sha256:e7d2a270abc18d4a52f772176e6ac59bd61f681a0ad814453f948ea319642fdc

Observation bbd1ebb1-33ab-4190-9d6e-d788077e060c · outbound

This paper cites GameEval: Evaluating LLMs on Conversational Games.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation GameEval: Evaluating LLMs on Conversational Games

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.360560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.360560Z digest=sha256:2e24060adebe5184a20630a591e9dd0694d676ad629a8a56c4aac8fa4a38e349

Observation 0c288d72-2f51-4edd-9cb1-e6f6d4a8f2c4 · outbound

This paper cites Least squares quantization in pcm.IEEE transactions on information theory, 28(2):129–137, 1982.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Least squares quantization in pcm.IEEE transactions on information theory, 28(2):129–137, 1982

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.463723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.463723Z digest=sha256:e8c1e9ed51bec84f76728824be297f05c93cf87cab058002632c56a8afc28415

Observation bb4b8319-ba75-4af0-8f3e-c49a25c75d6b · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:08.105891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:02.575342Z digest=sha256:de0bd71c3706fe76cca6d8f2d64f71090e75ba5dda8969f111fe7ce7e3ffadf8

Observation dc0febd7-dfd3-4257-b303-d6d35b337bc1 · outbound

This paper cites PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.673624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.673624Z digest=sha256:497e552b7b53fabda941ae8c2ed411d0885c232d3c94988eaa0f36f1b9b6e721

Observation 5d6ad255-2e81-4393-a49a-7658872f2539 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.783375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.783375Z digest=sha256:39ba13f635d53d39e6991a31b51976decbeee9ee06f350eb298fdba8647552d5

Observation 25cb6302-961a-4c3d-8c72-d1278b151cd9 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation KTO: Model Alignment as Prospect Theoretic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.856184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.856184Z digest=sha256:f17bd968a1ea7506edbd50a28a8795562147c47e85628c053b4015bc7facacc8

Observation 6833c2ab-c860-44eb-bb1d-7b5757caf38d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.951993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.951993Z digest=sha256:5de298e43f52615c4e818a121d259b85ae8fcdd7f10f1d6e1fcb95809ad89aef

Observation c1d10df5-5527-4490-9578-e69f4bc0dc85 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.054149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.054149Z digest=sha256:192c1df1ed8427b0767c020b9b31885138d6587d73eba0a0a20bf80ff703bcf4

Observation d1674156-629d-4af9-9f73-e7dcc6888ff6 · outbound

This paper cites Methods of Hierarchical Clustering.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Methods of Hierarchical Clustering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.131933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.131933Z digest=sha256:84aaadf8c5ce5815e25754f71c44f3dec146108f7c2a06f1d275f295095cf071

Observation 00333b33-ec03-47de-b745-e3d1bd363535 · outbound

This paper cites Silhouettes: a graphical aid to the interpretation and validation of cluster analysis.Journal of computational and applied mathematics, 20:53–65, 1987.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Silhouettes: a graphical aid to the interpretation and validation of cluster analysis.Journal of computational and applied mathematics, 20:53–65, 1987

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.225444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.225444Z digest=sha256:5b09c9a94c539994b5ec1296cc6e6133fd070d8c23e90cbfcac362384abdfb9c

Observation aa50cd92-6e53-479a-ae57-4bd157a20113 · outbound

This paper cites A dendrite method for cluster analysis.Communications in Statistics-theory and Methods, 3(1):1–27, 1974.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A dendrite method for cluster analysis.Communications in Statistics-theory and Methods, 3(1):1–27, 1974

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.304018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.304018Z digest=sha256:a0b18810a82369312a8b12a0f9e90d62ebfac109dce8fea774fc6ac4be84982f

Observation bca82fdc-9a2f-471e-bf1b-6e772530ab90 · outbound

This paper cites A cluster separation measure.IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 2009.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A cluster separation measure.IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 2009

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:07.852623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:03.440574Z digest=sha256:37df57b2ded51d5af79e7a7ffb9b6fd59304f953971214d8ab76b90bf73c6a03

Observation 3285c6a9-3962-4f20-a7e0-e68f3db952a8 · outbound

This paper cites an unresolved cited work.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:07.565203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:03.516997Z digest=sha256:ef699606830e51f2fb3a36207812ba486a988ce4067079fb52384109861c8783

Observation 3cee0ba5-e240-456c-a8e8-3aaa2c5bf7a5 · outbound

This paper cites Training agents by reinforcing reasoning, 2025.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Training agents by reinforcing reasoning, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:07.345382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:03.619683Z digest=sha256:5bdc9d8674cdfbb4ce1591ee8cdf02b01d5a0f412685c81ce36e345d57f0d679

Observation f040559d-b41b-48ae-92a5-ffc0336ca356 · outbound

This paper cites Glee: A unified framework and benchmark for language-based economic environments, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Glee: A unified framework and benchmark for language-based economic environments, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:07.099292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:03.705946Z digest=sha256:1dadcd25e9b57edf666a930a46692fdd194850df681e708d43ab52c301f53919

Observation 3865c1e6-fa12-4026-a4bd-f08b25df12c6 · outbound

This paper cites Llama 3 model card.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Llama 3 model card

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:06.818197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:03.782318Z digest=sha256:2e2560b409cac67348efd783842f49e1ef3debdf86e3c7a7dc4dc9cebcf9d59f

Observation 459696a0-c5ae-4e1e-9e10-1c9e624a0749 · outbound

This paper cites Text and code embeddings by contrastive pre-training, 2022.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Text and code embeddings by contrastive pre-training, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.880519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.880519Z digest=sha256:b1534da1c73dcb7c856eeb84b7edf1f0a661523484933e170daa564c2aa3309d

Observation 2ee2879c-ef7a-4d54-9b9e-abc74446ea9f · outbound

This paper cites Gpt-4 technical report, 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Gpt-4 technical report, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:04.011359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:04.011359Z digest=sha256:64f7eb238a334d2509821e3051bf248a72d7a9025ce3ef3bc7dea3756196081a

Observation aae6b85f-644d-4cc7-a598-cc367b24b11c · outbound

This paper cites Introducing claude 2.1, Nov 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Introducing claude 2.1, Nov 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:06.679584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:04.104372Z digest=sha256:b73122bedb9d5dfdcb0e0017d8a589ca9d91fc21f8fd6f377c9abeef774a7b06

Observation 869581f2-df74-4623-807b-0e43b3fb49e7 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:06.305099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:04.178352Z digest=sha256:4d312ddb7fb59fec4fb30dbf8c4d23884e9758411937310e8a96f903f7f56026

Observation 7612ea95-a64f-4ab3-8ca7-c8bef3c4fd24 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Qwen2.5: A party of foundation models, September 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:05.910016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:04.271744Z digest=sha256:b3d2249a21cf00e650542c732dc6897172a6c384c17e8456e454790465888f57

Observation 91bc5c2e-1b1a-40d5-9ef2-e5bd3ba6696e · outbound

This paper cites an unresolved cited work.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:05.625633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:04.344497Z digest=sha256:935e0278027daf1a1024fb4d86a323ee6d293255a46e8d1365cf61a8202b6f8a

Observation fddb8384-c1c3-43e1-8a42-afdff67bd4ba · outbound

This paper cites an unresolved cited work.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:05.376816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:04.438910Z digest=sha256:0ad49dd393b8e34ab6f3f5d900837dd0a3fc2c8c639f6bde94039263137ab23a

Pith citing papers

Observation c3128c10-4099-4755-8cb2-2203ceb52e11 · inbound

Unsupervised Learning for the Elementary Shortest Path Problem cites this paper.

Unsupervised Learning for the Elementary Shortest Path Problem ARIA: Training Language Agents with Intention-Driven Reward Aggregation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:42:59.909229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:42:59.909229Z digest=sha256:59c6995e4684da2e7bedb677f998101c51113ece492cf1f62c62d609ee07dd2b

Observation 6d425eb3-e2b6-43f7-87c1-d24c0a9c6d16 · inbound

Latent Action Reparameterization for Efficient Agent Inference cites this paper.

Latent Action Reparameterization for Efficient Agent Inference ARIA: Training Language Agents with Intention-Driven Reward Aggregation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:48:12.946615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:45:20.306945Z digest=sha256:929dbe548a66b9ad336943114643b4fd5da26d7b536f56f31cd0d04c21eaa754