Pith. sign in

Paper Citation Record · LEDGER

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.21252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21252 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:46.166368Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f06f09b6-722d-4be3-b938-e09a910cd995 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.369605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.369605Z digest=sha256:24fd2c0ea89252d1020b238233fea1f4cac303689aa3061fcc63621730040ff9

Observation 9725d029-a434-408e-b4f7-1a3a251fada7 · outbound

This paper cites GPT-4 Technical Report.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.438405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.438405Z digest=sha256:6c2d0a0cfd5725d18650068d60660a74c34c0c83ec5f8af66042519b3248c953

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.685574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.685574Z digest=sha256:fa0618d5f44ba0e72e99679cc65171bf3d37cb09115bf199b313e8889c1c6d03

Observation 69c1f3e3-825d-45de-a713-bd71074f7788 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:48.332392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:41.826052Z digest=sha256:60f7d289d07fb4adccea25cb779d089c2582aba1038f4aadff302f164e4289ee

Observation 17a169ef-baee-4141-a20b-587776fef169 · outbound

This paper cites Large Language Models for Planning: A Comprehensive and Systematic Survey.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Large Language Models for Planning: A Comprehensive and Systematic Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:41.978125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:41.978125Z digest=sha256:512e9aa70fa7a8fc42686f81739c2f48281481fcfd6f2755f5f43bb46669bc35

Observation cfbb0d56-e472-4344-ab30-6a7a09779e7c · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents FireAct: Toward Language Agent Fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.146724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.146724Z digest=sha256:e6be0af250174d3ea1e11329c3ba8fd02035de4f27f88170af7cab90772e4088

Observation 2d722460-63ac-4ef0-bdcd-672c1891413e · outbound

This paper cites Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.313802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.313802Z digest=sha256:bfff44a6c7d9829fde326bcacf037971751cb93aba074abc4d715f20f4c5a96b

Observation 8b15e6fa-9b30-40d4-8df5-08b7af498f81 · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.406234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.406234Z digest=sha256:82f1285d70bf4fef4df4b54696befa3d6b105e5b36f175811cff0e0d4a672a59

Observation bdb8f448-6613-409d-8391-bf6d51358543 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:48.031390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.497421Z digest=sha256:afe9c6fd49bf8ae6671250276cdc91802f2485824d328d5bfff6d1409c1c6174

Observation 40593145-71eb-4183-9f85-5f38e66d7558 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.561879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.561879Z digest=sha256:997788ecef1ea046c4028828b4afd76948ff01f3b60f5126f6a3f57d0d6c8fb2

Observation e650d70f-b58b-4084-ab85-5bdf84073b81 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.726733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.625120Z digest=sha256:da879cd4637276c5239afa405f0dbfe0596c58fdcdff1acae923b81a06e61281

Observation c4f224b8-47a4-4c60-867b-e6dd639c5596 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.536393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.687111Z digest=sha256:c15a909a0baa46e4df580bf49d3dea27b13db63a14e5522136907af52dbfa0fc

Observation d349826a-2549-4668-ab70-27a71dc6a5f6 · outbound

This paper cites The Llama 3 Herd of Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.774037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.774037Z digest=sha256:3c30af80e2c504adc0cb08fdec22a139748ef10f90a68f9ec7fa8d7892670856

Observation 12488757-e78c-4a47-bf9a-8b41e6bbb716 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.296199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:42.844408Z digest=sha256:ba4509e10659ca79f2535b0c4ef4f41fdf35827b316d662cb6af9a1f814dba78

Observation af810d31-6b93-466a-aae4-81e72e25c7fa · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.927894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.927894Z digest=sha256:3d47b59334810696caf733ef9d2e7aae776cdc7e92107e2de6d6ffdd4980e540

Observation 11592a68-2150-4cfa-9ff9-147ac89d1dbf · outbound

This paper cites GPT-4o System Card.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.013554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.013554Z digest=sha256:41b485b57d0a8bfed40b5f0f06a66edb5dcd9385d7fd7f0fa3acfb4dd6f097fa

Observation 563cd57d-4319-48dc-a29a-5bdf892a5138 · outbound

This paper cites RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.134536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.134536Z digest=sha256:8b043ba2ead2158ddaac21fc647ff7bc3d31be8624260c1ca8ae68ff11f304a9

Observation 55131c7c-5202-4653-8ad2-711d937d028a · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.222340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.222340Z digest=sha256:0058e551c46b979b553b5297072c5cfa6ddf52897b6dc2622de747003190ef08

Observation 168d3974-78aa-4789-8369-2f46f9d735fc · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.283968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.283968Z digest=sha256:7f3422eacae416c10639682ca7b13380747e9f2e3c06d955af6729f87a8c9dc1

Observation 2d3bb893-9a25-4e4d-848c-60ef6b79823b · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.360429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.360429Z digest=sha256:fe0c389fc6a406dae9e1652405be8d265e96cf6c67bda8de780f0b4e826c44c4

Observation 6c9e2b1f-c263-474e-aa11-0fce5b7db5a9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.438047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.438047Z digest=sha256:225885d18962ac9a2decb951c97293771b7de892de8514f1f08b7a6612eb0816

Observation c4bef8c0-b8c5-474b-8a7d-b23863924a88 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.531337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.531337Z digest=sha256:4c085d42fc6827720b36099c0311e4136f7286be27385d4982b1df707f7e3a5b

Observation 2c681772-ba8e-4653-83e6-f73694c759db · outbound

This paper cites QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.620059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.620059Z digest=sha256:dcbd0fea046f4ac2f62fb5ecb8be1b337be6233fee64259aee2f5d35f13fecfb

Observation 027d2cb7-8b0b-4b7a-a572-cc99be388e3e · outbound

This paper cites Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.709254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.709254Z digest=sha256:89ab191cfc2bbe676d3ff9df27de1d544f2863c7733950cbe50a9428e1bff1e3

Observation 0cb8f075-81df-44af-a819-b23870384605 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.779856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.779856Z digest=sha256:95fcb27685c4e3ca2848800681643f40d5b163a41bcf993ffea1d58eea7475a7

Observation 36cc2361-8787-4bb2-b086-cdebe9bba711 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.891326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.891326Z digest=sha256:1cad8201c39bf58c13ccf4642046a7922dbbf34aec898bcf3f0f68813a53ede1

Observation de0dac99-505e-4fda-9f2a-d93267561f42 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:43.969795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:43.969795Z digest=sha256:dfad6c06e2cfc4b08933fe15a35d771f523f59d766dc59dea462c7efdb090e7c

Observation 3dc86184-7470-46bd-ae30-9af0b0e30b48 · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.079391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.079391Z digest=sha256:3bb0c659cd8086c9ff35c5432a55f13b641252d973240be9320ab5038b157d03

Observation 3bd24895-869d-4623-ae1a-74c79815fa6f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.193779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.193779Z digest=sha256:bc4e73b76f4965399fecb17de7530cd95aa8f19ddb812d3d0bda52968e085e01

Observation eaf61ff5-95e3-42b3-8bb5-622ab65d8b6c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.304809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.304809Z digest=sha256:6d365560a806a13f96967d96833b35b50d72911ad3cf8f7bc92d70597efa16a2

Observation c46d6f02-f238-456e-8f5c-127d2a1518c0 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.366489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.366489Z digest=sha256:101f52af898b7089400fe2bb43688def7f105f52ecfb04dae25c3ccfa67fb083

Observation 9bd1a4cf-3b2f-4787-ad6c-73ff556a3023 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.432191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.432191Z digest=sha256:ab23509b82cb727a773870d382c5ab9c1c0a06b79096f44c25f17f0fe03cac36

Observation 97b2a971-9d40-41bc-86c7-83e970d60b63 · outbound

This paper cites VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.501851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.501851Z digest=sha256:50eca96a8011cbbc32278181631a8f3a8b27faf56dbca59cc3340790ebe0b251

Observation 8e0486c0-ab10-4aaf-a464-eb7c2fc8e9d9 · outbound

This paper cites TravelPlanner: A Benchmark for Real-World Planning with Language Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents TravelPlanner: A Benchmark for Real-World Planning with Language Agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.505963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.505963Z digest=sha256:962fe283e2cd4d7cd6478baf5fe716b9c336393bf7681e82e69773b24083c644

Observation eac92648-2a35-4996-8114-fa90e82c1a78 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:44.585127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:44.585127Z digest=sha256:05fa0a6d5324e2d74e90868159937519694517d1ffa4f68cf08272da67d4dfc6

Observation a412a3c8-72dc-466b-9fde-208523838d65 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:47.087466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:44.721148Z digest=sha256:b3e6491584340edc1f0fc9a9e93a0593d55a337bea86d42b6e67354eeec3430e

Observation 5d209ea6-adfe-47c5-bf66-a0419f304ee3 · outbound

This paper cites an unresolved cited work.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:36:46.860630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:36:44.901469Z digest=sha256:894fcb18a635a819e1aaa9ca014d45ce478cae42ae43e6fc23965f01a6759dcd

Observation 7f309774-b59f-4f7f-be72-d7885afdbca8 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.082095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.082095Z digest=sha256:669a7aaef1c0944b9de0c93e2d34075d0469bdf6f7c7c5d84a516354647c3e78

Observation 72928ad8-48a3-482c-9f98-3429532b762e · outbound

This paper cites Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.285843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.285843Z digest=sha256:df66d80a594c62916d3f1aa4a9278d986d4f99618fb288e9ead6fab6bcca84b6

Observation b81944de-802d-48d4-8d7c-31505eeed8c7 · outbound

This paper cites MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.463037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.463037Z digest=sha256:418759bfef957da4ea03609a9a8b3a8542e28d9bcca34b1f1e12c6fd8faf5af3

Observation 8a5f0f52-67b2-4554-a6b8-26260ebf08cb · outbound

This paper cites Attacking Vision-Language Computer Agents via Pop-ups.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Attacking Vision-Language Computer Agents via Pop-ups

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.668547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.668547Z digest=sha256:b3bc5f0691ff50d3877261bfcdb68eb680f607668e744bed9d4bcfd056f15f74

Observation 9f1ab5c3-d448-454f-9bf4-2ccd1f9dac4a · outbound

This paper cites WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.814121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.814121Z digest=sha256:bff5408acadc7ecea9a5568dd44ac63b2b4143cd89c2df16d1394d4a80dff977

Observation be8ff1ed-8400-433b-aa1e-e13b8a79b760 · outbound

This paper cites RMB: Comprehensively Benchmarking Reward Models in LLM Alignment.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.894072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.894072Z digest=sha256:9923a84a503c8bc5f16b696b470608c2b9f4388b554cd42e82af87f54aaaa1ca

Observation 59d6978a-d62d-4dbf-9edc-cd76423a9675 · outbound

This paper cites Multimodal Situational Safety.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Multimodal Situational Safety

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.933900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.933900Z digest=sha256:b663da94193976548f83c02f5fd6b8063a071a36c02416b0fa9308567d140ce6

Observation 83c0accd-e77a-4a31-b5f9-5fa06e6e5ad1 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.975203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.975203Z digest=sha256:ffa78e58f1ec8e1a91421af85433da556164905acf8188bb12d3485abe94510e

Observation 3493dd4f-8f64-4b55-97a7-1aaf7321bd13 · outbound

This paper cites online" 'onlinestring :=.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents online" 'onlinestring :=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:46.052241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:46.052241Z digest=sha256:45393ad25008931a42511c008a36cbb87c21b218d82fbee33d614d95a893b5d2

Observation b1d9ba60-f5b1-496e-bbd3-e045859f0fc6 · outbound

This paper cites write newline.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:46.166368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:46.166368Z digest=sha256:a7332d5c5d67360ac2f72fed9c7635ee3d725dfced54f93bf23239f306abc29a

Pith citing papers

No inbound Pith citation observations are available.