Pith. sign in

Paper Citation Record · LEDGER

AgentRM: Enhancing Agent Generalization with Reward Modeling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.18407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18407 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.996079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.639558Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 75960073-c17a-4b22-a1eb-b9bd92624d67 · inbound

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning cites this paper.

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.996079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.996079Z digest=sha256:7e4e7b4e024466dc4e9e1e0c4949a7cf5071208049d2357f40482bb991c95f2a

Observation cf9fdf4f-a4ba-4e57-b9c9-f3686dbf2bf3 · inbound

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents cites this paper.

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:15.862250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:34:15.862250Z digest=sha256:51f53c6c56b48612f57b6ad89d833dc0ada1d6f5936c4275ca29aa560a9e9961

Observation b0bc5883-362f-47e2-b5bd-e51e120a841e · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:22.078905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:22.078905Z digest=sha256:38f099daecbe4caa83343469d0df02745ea481e535a23dd393793d829a506412

Observation cc3c6853-25fa-4ff0-a0d5-6bd34bfbdd3f · inbound

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction cites this paper.

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:40:44.032643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:38:37.818623Z digest=sha256:a825edf8d86bfd8bdbf1d7aaa5596f9f309de46554d5c2694cf985f71c09024e

Observation 7d999d4c-8de5-471b-ae07-425c2ea28736 · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents AgentRM: Enhancing Agent Generalization with Reward Modeling

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.641467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:922065a6d64f65330060d53f1430d015c818c2f33d18d7a79b338002c863f8a1