Pith. sign in

Paper Citation Record · LEDGER

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

As of 15 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2505.20726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20726 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:54:10.639156Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:36:21.457436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:36:21.667275Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a672600d-eea8-4f3d-a679-ce150f1dfb45 · outbound

This paper cites From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:54.825537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:54.825537Z digest=sha256:0c91ffd93b84138996fc019a858461a63a21e81a436495d9ee7eb841c818d8e5

Observation 1a0a896e-7dc2-4001-a7ac-35e96e09321f · outbound

This paper cites An Embodied Generalist Agent in 3D World.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making An Embodied Generalist Agent in 3D World

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:54.957407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:54.957407Z digest=sha256:0fc4558afad65206610c755bd4e88c1548df792c6058ba03936c9aad7c55d288

Observation 2bce66b4-e12c-4561-96fb-633b72938e40 · outbound

This paper cites Rearrangement: A Challenge for Embodied AI.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Rearrangement: A Challenge for Embodied AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.061063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.061063Z digest=sha256:e9255acd0b0ca581c300e6d8a931619ed6f6fa8d627500136321f88ce5d12cb4

Observation 15921749-94dc-40da-880f-362e189737d1 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.137137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.137137Z digest=sha256:15e0675f0c123ae9449ec5d7b275634d66fe0526ddc226b80545b2febc37eb66

Observation bba39573-5795-4d24-8a53-c3c84d3a8ba0 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.233055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.233055Z digest=sha256:f3aa4ce997905951bc331c641a4a61cfc6929e4a11c319f4f25d412153d1a242

Observation 5f10fd22-c294-4a2a-88ce-f1e5cf49199f · outbound

This paper cites Palm-e: An embodied multimodal language model.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Palm-e: An embodied multimodal language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.328013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.328013Z digest=sha256:66cb3c551f0beee048a5a6e7a9f0b47158df7d5f4ad1d551e8d2449097c9ff89

Observation 09681592-73b2-42d2-b461-4ec0bcecad87 · outbound

This paper cites Large language models as gen- eralizable policies for embodied tasks.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Large language models as gen- eralizable policies for embodied tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.428981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.428981Z digest=sha256:786c3a8c76460ebfeed4443c840944bbd385289e54c237aafc9ba4b749ca27b7

Observation 2e61dc64-8296-43f2-8a98-8aa3e3fab4a4 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.547151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.547151Z digest=sha256:2ba441159cc3657e9a17e774baf5801b6b37198ce8b5ae4b4fc9acf6395c6f3d

Observation 84fa2242-9b03-4602-95ac-7cedd8676442 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Gemma: Open Models Based on Gemini Research and Technology

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.659073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.659073Z digest=sha256:28cece8bd73deb781b851f4463ebabf04b8f33a0a2704e6da335312c65ea808d

Observation c8d01f6a-31c6-4a0a-a29e-ca21030d8ab8 · outbound

This paper cites Qwen Technical Report.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Qwen Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.744626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.744626Z digest=sha256:75e5fd4531d51379b52a66561189ba753bc6afd8c92d7d3a4e82cbe7ca594a3b

Observation 03a7cf90-24c6-465b-b207-a1ac23c24956 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Physically grounded vision-language models for robotic manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.818724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.818724Z digest=sha256:9cc150b2210c9e64f3a73ec2133c66f21caf74e4fb9d69e663af70433392008f

Observation d8505a54-c944-4812-b1ab-56058bdd0763 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.891849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.891849Z digest=sha256:5db89ebbb606188bbd6445e29e2025446c307d56bfb1ad85628e2b9a45770ff3

Observation cd42b183-0e0d-4c4f-807a-e76342afb5fd · outbound

This paper cites Multi-skill Mobile Manipulation for Object Rearrangement.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Multi-skill Mobile Manipulation for Object Rearrangement

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.986927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.986927Z digest=sha256:e0966adec305da985211a97a09320db3b7b92b59176d39ac644b194c9b0aee53

Observation 9c452002-c1d8-4c61-81e2-0d5069078971 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.087735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.087735Z digest=sha256:4ace5475f2cb13c531ffbd6621da7a3e1c5bf7d367a56d8c030aaa5736e4713a

Observation cdf95c37-3141-4763-9c69-7a7712077c8f · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.192260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.192260Z digest=sha256:75bee9180629a46f2e312573f3cefb7201229d59d34b3e43c79dea0507232e6d

Observation a31e1bfd-41f2-48ea-a237-abce7c0f9499 · outbound

This paper cites M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:54:11.656017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:56.263491Z digest=sha256:c0979173fb3fd73204a88223ccae35af07c8c9cb75bd38fcef65cc2599e9130d

Observation fefe6b08-2850-4476-8406-fcf0e1e875a4 · outbound

This paper cites Embodied agent interface: Benchmarking llms for embodied decision making.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Embodied agent interface: Benchmarking llms for embodied decision making

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.353324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.353324Z digest=sha256:3b862838c807048f20bb7f8017dbba0227382b2de038738e9003e9202cc650aa

Observation 9a2f7942-6d4d-4a4a-8a23-45cfbbc8e047 · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.447497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.447497Z digest=sha256:1f0509b6499947bc53f1d611c147a79a146d803b6efd2c0ba27053a97f75e9db

Observation a852e5a6-9764-4616-a394-990e93212934 · outbound

This paper cites LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.545921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.545921Z digest=sha256:ba68ced9268241e2181915a9b093468b3f8d1e2da8ea72b62656e7c8ee77d65f

Observation 9c824d0b-2e4d-43c7-b157-8e79aefbaee2 · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.725406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.725406Z digest=sha256:3879920535273cb5d07d42fc8398a938a430b0af44167395311d0545e1b42232

Observation 19189071-c053-470b-8c64-f823e3ed0860 · outbound

This paper cites EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.823209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.823209Z digest=sha256:86d344e27c340292856eef438a07ab49e89bfc3d3a7d881decf6e6a94ff9e1e0

Observation 6b626a01-d55b-48c7-93c5-59eb075399df · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Habitat 2.0: Training home assistants to rearrange their habitat

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.922046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.922046Z digest=sha256:21bbd5ce5485ee2a813ec28a97cfc335aeceb9f70be8a9ffa855acfe294e7d44

Observation 15c6aa96-a5b1-4838-a960-763c259df222 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.005818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.005818Z digest=sha256:2943acf04ef31a25295ac5d5d2bb6fc0c42099a68ff42080d8661809207608ea

Observation 80dae249-2d41-4925-af64-4a7145de8b2d · outbound

This paper cites Sun rgb-d: A rgb-d scene understand- ing benchmark suite.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Sun rgb-d: A rgb-d scene understand- ing benchmark suite

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.095946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.095946Z digest=sha256:79161092a8f2c899a2b393e7e162b9684e31005b22af4985c3c56de2735f535d

Observation cde03339-dc71-4810-ba10-3274b15fde30 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.168766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.168766Z digest=sha256:a055c220001cfd1b4f36989753da88f2ea3e4e83768723a429c1494789e5b990

Observation 589b5135-de90-4db7-8927-50b5f19f43a7 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making React: Synergizing reasoning and acting in language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.240312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.240312Z digest=sha256:d23b6b80e56685d2d8373e8a140fa4bb697d7612a835963beca0d5f241fa71a1

Observation 87457872-452d-4681-a83a-e7d9e985fc23 · outbound

This paper cites AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.326324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.326324Z digest=sha256:6d6eeed0a4d874caafaedc42d8885b7237c5623aff06a63cf7b6aa1edd9d49b4

Observation 26a37aa7-7d09-465c-99a3-a8bb426f167b · outbound

This paper cites Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.380409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.380409Z digest=sha256:a912255ad4a30ca5af317a591e047873880e524f28901ea7adc6e3f1e09dbc7f

Observation 20ed65eb-b100-49cb-a67f-ed491cfd3b57 · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Leveraging procedural generation to benchmark reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.458580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.458580Z digest=sha256:48dfe30ec1e0ca03cbfcc7c9ff594e969b82ccef5e6ce6c15c23af53cb0395e3

Observation d8cb284d-c321-4eff-9481-1f696a007576 · outbound

This paper cites RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.551508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.551508Z digest=sha256:cef0556235176d225cb31322029ead845f3f90402cbb94fdf291c05f7cfad07c

Observation eb3a3191-31d9-40eb-979f-18817f69b652 · outbound

This paper cites Adaptive Procedural Task Generation for Hard-Exploration Problems.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Adaptive Procedural Task Generation for Hard-Exploration Problems

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:54:11.407001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:57.649094Z digest=sha256:5851463db303fdeb709f94a6e8dc8730493cca23acb20c58a10b1dd5564fd99f

Observation 60942058-9243-4d47-908c-449eb87b0c8a · outbound

This paper cites GenSim: Generating Robotic Simulation Tasks via Large Language Models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making GenSim: Generating Robotic Simulation Tasks via Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.740352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.740352Z digest=sha256:dcae630a7d993a297e0e17b96abeb3820427b15c3716ea573da492ac9c389d76

Observation 9cbbf7f5-23a6-4329-86b8-4accebf656b3 · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.835282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.835282Z digest=sha256:18451f2032b134c535a8aa599fc0d49ce612d31761e8979f302c5da611fef08f

Observation 491d1c65-0b12-4e12-b7e2-732400a61c2d · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.932255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.932255Z digest=sha256:06250e1043d2c9d190d79a2ac69c8d9e53a8fb95954955333a186cb28cffcac7

Observation 8e9a6e81-93f2-4cbf-b8d8-6067f5312c51 · outbound

This paper cites Habitat rearrangement challenge 2022.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Habitat rearrangement challenge 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.573015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:58.002577Z digest=sha256:1e00e5dc8007fd61d0456f0f6088eb0c15f4a03e03b3a2c60840f291f27db271

Observation 65561e2a-1be2-4230-8bd2-485387ff7555 · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.124688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.124688Z digest=sha256:0902e02effd968ae6520f00f3da8d2be3a67b9dc6830e17a0c8b24f1b8c762ae

Observation 860b9867-74b1-4221-abff-77adbc88d483 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.184489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.184489Z digest=sha256:8896dc62c1c01d702556317e4875c31165a2514929440a8f4e0826d39d49c49d

Observation a28e3a21-dbaa-4511-bd18-61b1703b31b5 · outbound

This paper cites {\lambda}: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making {\lambda}: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.275583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.275583Z digest=sha256:f23ba6fc9144a17b7cf0ecf27c09fb2d6590c9135593bf8ceccc24ecd1a0cad0

Observation 022070a7-067d-4dba-85cd-84da8c35bcd0 · outbound

This paper cites GPT-4 Technical Report.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making GPT-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.405553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.405553Z digest=sha256:b2a179415bd438230450e7ca569206aff40382bc0579025dbdbc26d30114b2f6

Observation 157df454-c573-47a6-8b78-4a8671dc9d89 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.489673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.489673Z digest=sha256:71be0e79228ffa059fab380296ba0d55cc447919d677ad61e46cb34a7bfe996a

Observation 2bba0d62-520d-49ed-a488-3294bd16f4fa · outbound

This paper cites About claude models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making About claude models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.408060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:58.621363Z digest=sha256:8427bfc578981893b61d1b0f425307f99ff8bcd187b2def69e259552f395321a

Observation b1723ad4-33a5-4287-b4c8-dbe1d937df6d · outbound

This paper cites Sapien: A simulated part-based interactive environ- ment.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Sapien: A simulated part-based interactive environ- ment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.722494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.722494Z digest=sha256:1d8a692dab2154443faf1ac893b966fe5fc008a624731f2d038aba655406d35b

Observation 31cfd8f5-151f-467f-84ef-371c1cbc174a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.811914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.811914Z digest=sha256:70c3d3e2f35de27eea912c631f188b0dad94f568a338fe68182dfbd84d1625b2

Observation 98e7bdec-d034-4e89-b96f-203611b907ca · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Finetuned Language Models Are Zero-Shot Learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.927954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.927954Z digest=sha256:4752ff447dd8814edd0377f8e212870fcacef685b446b8fd620bf47dac0a1694

Observation d0e6ad65-1dd7-4515-8481-bc19822d0031 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making The flan collection: Designing data and methods for effective instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.022673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.022673Z digest=sha256:947f5bbfb23932438ca77fd848bee7a0fa24a403ebfc26ed3193890ebd20cc06

Observation 08a4195a-2f23-4da7-8006-6e5654bf458c · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.239608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:59.119195Z digest=sha256:3f989d985522c372a9e1419e7739189bcbfaf38c94b21869171eba42c7ca98ba

Observation 08b3811a-e0fc-4c3e-961b-604bfd8aff1b · outbound

This paper cites Cogagent: A visual language model for gui agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Cogagent: A visual language model for gui agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.194279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.194279Z digest=sha256:14a9701fc8f8517cf1d0ad8c4c0ce3eed01a6200f61048c10c1272fe7c68c065

Observation 69e68284-c532-4ca1-8d46-b976065c6844 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Autonomous Evaluation and Refinement of Digital Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.280411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.280411Z digest=sha256:adec50d2a6546f1b34d3200b019ac0db40365296e846392dd27713754487190d

Observation d478c978-2271-4eef-acb5-47450fea440c · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making ReFT: Reasoning with Reinforced Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.379335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.379335Z digest=sha256:8eebcefe58ba12feba184171b0bd3efe28f68078f60a128ad4e5f7e0b90d3074

Observation b0817a52-26e0-459f-93af-865c92ee0d83 · outbound

This paper cites Grounding multimodal llms to embodied agents that ask for help with reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Grounding multimodal llms to embodied agents that ask for help with reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.449785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.449785Z digest=sha256:3baef80213f08a4092149a865c41c04605aad699408a26325a075c54c030bda3

Observation 507ccd29-01c6-4931-9078-8b810bcaa289 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.540056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.540056Z digest=sha256:ebc896d5c0137e2b6146df51a79ff9b0588adba487e3daa5ceb9c3ddc14fe9da

Observation 739da9cf-ded9-4927-93c8-05911f78a8c3 · outbound

This paper cites Grounding multimodal large language models in actions.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Grounding multimodal large language models in actions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.060631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:59.614383Z digest=sha256:885d27daa4105ce44063d8ff8d571ef21c62e7c26c22d188a644e21f95be65ef

Observation e8dd4e21-f13b-468b-b1b5-90ce35bf3098 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.697480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.697480Z digest=sha256:ec9c8a78694e6c66255daad0748fa982a5c6b7aa6b506c647869bae0f517b00c

Observation c3f27528-3245-4652-bff7-d2160513b6c1 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.798027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.798027Z digest=sha256:71f852c06866dc04ab09041048a4adbd9854a4a89fd17ba9944287019cba9f25

Observation cc9773c0-28d8-4aa8-9f2b-d683421c85af · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Fine-Tuning Language Models from Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.863345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.863345Z digest=sha256:ebc9a08e0306b8ec8cdaebcaf660e384cd92e94933c817d79658384167a7394a

Observation 3cfe34e8-1326-4792-b080-df0193139256 · outbound

This paper cites VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:54:10.983636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:53:59.934887Z digest=sha256:08f6a703a4c2cf9258585ded7ae96e9b72de2db0c98338748ccc1f5dd3187fe7

Observation 9837c052-b1f1-4fb9-a027-bd72f5b4a1be · outbound

This paper cites Training language models to follow instructions with human feedback.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Training language models to follow instructions with human feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:00.020348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:00.020348Z digest=sha256:78ad6a8bda3a5497ad4bb9a5044fdf8738b23ed28b446b00e5971fc9c638736c

Observation aaa36dec-fb2f-4a78-b795-9f6d5efe4387 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Grounding large language models in interactive environments with online reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:18.904679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.341575Z digest=sha256:f7a8f6c81fbe0e67d0056ffa6a1460481fa1670f437a65f429176663d328375f

Observation 97f0bd1b-2be7-4665-bbc7-87b794e22a3a · outbound

This paper cites Amazon mechanical turk.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Amazon mechanical turk

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:18.764475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.415231Z digest=sha256:1ce39b322a59f978dd117dc68984b6260dc5a8fa2960e5f6238e039f798def32

Observation 17e23274-fe3a-44e7-9c82-fd81388d39ad · outbound

This paper cites center" region. For the surrounding eight regions, we designate their directions based on the surface’s heading: the direction aligned with the heading is labeled.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making center" region. For the surrounding eight regions, we designate their directions based on the surface’s heading: the direction aligned with the heading is labeled

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:18.618071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.501803Z digest=sha256:38bf2389fdaf708d0f3fa7f48285942e70f7a485956f2f91d2f66ba4f25ef5fb

Observation fa5d6184-56d2-4c9a-98bb-686c0cf15a7d · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:18.421170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.614772Z digest=sha256:1b6225706403fb2a38ad8d3713428f2680c48a5ab6cc373599aa77e00dd8d5d0

Observation bfb85de4-ccee-40dd-973b-7b0456d5cd73 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:18.276703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.709463Z digest=sha256:e8d9b7fd3860f41ce6f15c608c795314effe20d8cf1a3eb6bfe414d3615c0e31

Observation dd4dc1dd-4ff8-4f29-b379-85e0078a52b9 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:18.127181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.837553Z digest=sha256:2e4979e5f72fb138dbeabf167e8046eb2c6fe6649cab9f7d45969ed65bd9c3ec

Observation 63d23804-baa1-48c3-a0ee-babf8dd229f7 · outbound

This paper cites Feasible.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Feasible

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.962332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.916794Z digest=sha256:d4cd1eb4d33ca19969db3bd51a5cd04e553d5ffd43ab9c0e9070081110e91f94

Observation 56623e55-9369-4f8b-accb-fe24a4275571 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:17.791018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:00.997321Z digest=sha256:39aa50ac1adfdfc0dcdb6d21df8c641f66ebf3834f9102bc3fe0e144d8852ff5

Observation caa15747-9161-4e46-9752-15250857f940 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.651016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:01.084572Z digest=sha256:36f4901dedc1f8c0c188afa75de113891351d2a5520d7ac5460ea16298707bcf

Observation 76a29f80-09f8-458f-a15f-0d45c2f7eb10 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.486413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:01.745645Z digest=sha256:605547456f868beb92658b0d8fece04c9fadd8d966f0bb834eb97a6a90ee54c9

Observation 3dbe886e-a1d0-4a99-ab0f-03b037400551 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.289911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:02.163867Z digest=sha256:cec8a0cdf1aa55319b79299e886ac23f806a9b72419b33ba5557f657d88ee3c0

Observation 4ede6929-eb16-406b-a8e7-a0ab6766d809 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.066562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:02.521528Z digest=sha256:033111932abec8a94f51da6d49e27d3fc5505f2325e25aff49a53f52d816af14

Observation cb745ab7-32d5-4804-8ffc-bec17e621db8 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.819053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:02.955476Z digest=sha256:48d73b39743a9514b8a34a051d06fd0ed1a9bf04a272e059fcfbb181556b2940

Observation bb20f84f-0ff2-45db-9f6c-3acd49494a80 · outbound

This paper cites Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.624389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:03.173448Z digest=sha256:9f9bfc67da309095380640b14e1ee9dc607cd3524da1c9f1c70e57470eac1fee

Observation f6e1e9f0-d471-4676-99fa-b272fa9438e3 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.431290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:03.392298Z digest=sha256:036ee89219a8fe91a2031ae6dcdf41bb6009a3a75b1ffd6e1a5f4a154c1232a6

Observation c5c672cc-6e64-4da0-84a7-4cd28d662672 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.184493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:03.786523Z digest=sha256:1c04084bc1f338411211d36a3b7d4303f753ec834d94bba4eca69f53785cce66

Observation da61c604-ee1d-4f79-b563-07b1714ac0da · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.007876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:04.211528Z digest=sha256:8c248bc13f6dc108f98440b12ea86d749725dbaca5c1486fd1ee6d0553c36d7c

Observation fba5372f-70b9-4ad6-a192-4c4d43f0e07f · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.764957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:04.629175Z digest=sha256:c2f01a31ff5cf7f28177ce74bd17f99db0c757f0df8df4921071d38e59f670c8

Observation aa90358d-d8ee-4b60-ade0-f37056f8815a · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 106

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.589581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:05.152138Z digest=sha256:d5282d8324255ed42ba444846685897fc3d951a8693ca0562c558f51879dc9e7

Observation 0cbccca1-3d35-4a1f-acd1-2a7aecbb5df7 · outbound

This paper cites This is important because the receptacles may not be intuitive.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making This is important because the receptacles may not be intuitive

Reference 109

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.421053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:05.467669Z digest=sha256:6e7fac7539b65f2614184055d9ef59b9f02daed02f6d001a5e6d0433e771f98b

Observation f7fbb0c8-4d20-4986-a7f9-e411d2e87971 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 110

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.201698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:05.580311Z digest=sha256:8da4fef1d8400ec16e6d1904a25638b60ec6227e95e1ddfc2b0b50569958c72c

Observation 12af8a4e-7f72-4ffb-9cc5-6baf120c27a2 · outbound

This paper cites show_receptacle.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptacle

Reference 114

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.919640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:05.985117Z digest=sha256:6d8dbb5f39cbd9827c4681fcd3fd99f98c6b9e22341f1b776ad9c4ef5a903824

Observation ead19c69-2f99-4516-a993-f5b52ddd15b4 · outbound

This paper cites show_receptacle.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptacle

Reference 118

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.697302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:06.393888Z digest=sha256:68c1c21fb0d6884f152807b95982f612d4ff32ac7a47668cd1fbd3208002c190

Observation 7a66cf5d-d623-4805-b8fe-4835337b3ccc · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 122

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.517228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:06.638527Z digest=sha256:65b25529d32fadd59b63cff747077dc800208f1f9c32f43a219485e422567d88

Observation 9f298546-1cb7-4b93-9abc-c22bebdfecbe · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 126

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.336685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:06.755301Z digest=sha256:4cf38821e1cb4728fc06d0b8cb5d45b73c205b57a0968a871c99b2ed2d7a5efb

Observation e43a41f9-68d3-4081-a995-b1111a8aabc4 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 127

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:17.472081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:06.851144Z digest=sha256:26b56ef353816ab37abc2aa7b7873a96f9e800125df2c86ce456cf181c3cd05b

Observation eec26299-e387-427c-b8b7-1a9b33773c66 · outbound

This paper cites If the platform has no objects, a 3x3 grid will be marked on the platform to help you place objects.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making If the platform has no objects, a 3x3 grid will be marked on the platform to help you place objects

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.316512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:06.945591Z digest=sha256:6bee40c11d3b6be66c5f78c3872a907b30e990fa7cb7bd383690129c9b446e37

Observation b476bee6-c400-453d-be23-6de13f74f7eb · outbound

This paper cites show_receptable_of_object_x_of_current_platform.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptable_of_object_x_of_current_platform

Reference 129

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.135743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:07.075792Z digest=sha256:7ce85bdffcbd74a36ca265a64b0c19a03c2c26494ae1631c0981f579d95761ea

Observation 8dfe0851-fa84-4ce0-b662-4b3df8ab90e7 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 133

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.194035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:07.659565Z digest=sha256:1d88c8e061d60da35946bb3787556c920f1e48ae7ab6640e549f0f5c8cd21f0f

Observation 66eefccd-e5da-4293-8e24-044ce22ff8f5 · outbound

This paper cites This is important because the receptacles may not be intuitive.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making This is important because the receptacles may not be intuitive

Reference 136

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.660557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:08.078388Z digest=sha256:747b6846d751d3412203fb195e544f3302c774d043358d5958b3c457c8367aa2

Observation 5da1660c-31f1-466e-9a6d-ece6b36564c0 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 137

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.014392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:08.206444Z digest=sha256:40c78ad8e42389dd37d30fb2b7bfdf59e875fa9067aecd3971188a99c34099ae

Observation a54dc73d-f3d1-4427-ac0e-21d48fce2c89 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 141

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.632092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:08.812446Z digest=sha256:125b1f48c5294bb59ad22f9d449dc3d19c3c667b8d420b08943b903dad6355a3

Observation b5b23e3f-e249-41a7-81d3-809e929a45e9 · outbound

This paper cites Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful

Reference 143

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.491374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:09.077039Z digest=sha256:cbb493d815069474561e994022d888034d62fd0c9c127865f4d76cf2d6924dc2

Observation 68d9d04a-736e-43af-bb73-40c35be53013 · outbound

This paper cites show_receptacles.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptacles

Reference 145

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.355455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:09.396524Z digest=sha256:25eb72bc2160cffd06645891b622c4f7a6341bfb37d7e9be5dcd12db778cd5d9

Observation 1d4bb4c1-eb42-43b2-895a-96a2038cdf3f · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 149

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.132558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:10.097488Z digest=sha256:08fa801cf053aba341f429dddf8c881a22e7c4b39173a874bd184dd69909b3dd

Observation 0b37ff0c-3379-4de1-ab10-80215710c272 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 150

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:16.981330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:10.216972Z digest=sha256:071c337a63374e00f509466bd370239eea7f694a00e8be4cc01a9de391c900d1

Observation 51d0f704-bee8-46a3-97cf-f6242392fab8 · outbound

This paper cites Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful

Reference 151

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.837465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:10.327005Z digest=sha256:df832193371eee883c05ea7d052e36df5a9f8df913e794f8819621976b4b0c58

Observation 539127ba-b0df-447d-be4f-332f0db22ebe · outbound

This paper cites This is important because the regions may not be intuitive.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making This is important because the regions may not be intuitive

Reference 152

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.851308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:10.473014Z digest=sha256:03c0388a0a5266575f9004bcf671495b2d3ab7ffd178a3fbdd3b11328a0c3980

Observation 36847eec-c898-443d-a95f-98b92e2b42ac · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 153

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:11.932375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:54:10.639156Z digest=sha256:fba783d16779d4c5d52adcc31f4cbfc3b16f49059c6eb1ae04ef29a0db756850

Pith citing papers

Observation 62e2116f-9236-46de-9cc1-b77ce111f607 · inbound

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation cites this paper.

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:12.440926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:12.440926Z digest=sha256:d4e35379a3e6a9791d5459d36b809a14fa86f63a0983578fa86f2097dc8de433

Observation 7b68a88f-8926-40c1-8f47-cbfa9a74a8f3 · inbound

Compiling and Benchmarking Task-State Horizons for Embodied Agents cites this paper.

Compiling and Benchmarking Task-State Horizons for Embodied Agents ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:36:21.671674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:36:21.457436Z digest=sha256:09ae16d7c0ebc7c4a6d998ed49c1decaad3581ab56f4ba6e157cf364621329f3