Pith. sign in

Paper Citation Record · LEDGER

AgentRM: Enhancing Agent Generalization with Reward Modeling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.18407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18407 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.996079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.639558Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 75960073-c17a-4b22-a1eb-b9bd92624d67 · inbound

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning cites this paper.

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.996079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.996079Z digest=sha256:351b014e839c0b061881dd501ac96ed9488e7c9f6f1e66839c014fc72e4ed0db

Observation cf9fdf4f-a4ba-4e57-b9c9-f3686dbf2bf3 · inbound

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents cites this paper.

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:15.862250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:34:15.862250Z digest=sha256:109d09b2b12b13237e7efb99424636433441e6ea83af4e22622878b7b3c3cdb6

Observation b0bc5883-362f-47e2-b5bd-e51e120a841e · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:22.078905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:22.078905Z digest=sha256:cf4c9bef28b9360ff9afc6e7d40c287ee7ef9391aae4d282e17af7a28bdbfec9

Observation cc3c6853-25fa-4ff0-a0d5-6bd34bfbdd3f · inbound

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction cites this paper.

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:40:44.032643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:38:37.818623Z digest=sha256:d797842f05947319b86eb2957c9b8e4a7a1e8f8cf85f5e6d1f27777bf83f5c96

Observation 7d999d4c-8de5-471b-ae07-425c2ea28736 · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.641467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:30acd51a3a0d9d52157a071da87cbab0c858bbc3e07c3313a463ebd9066e749f