Pith. sign in

Paper Citation Record · LEDGER

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

As of 19 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2505.20726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20726 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:54:10.639156Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:36:21.457436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:36:21.667275Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a672600d-eea8-4f3d-a679-ce150f1dfb45 · outbound

This paper cites From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:54.825537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:54.825537Z digest=sha256:1a580d6396d3e32f33dbc007745c514864b23a1586650c9509ceeb8270c7dc9f

Observation 1a0a896e-7dc2-4001-a7ac-35e96e09321f · outbound

This paper cites An Embodied Generalist Agent in 3D World.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making An Embodied Generalist Agent in 3D World

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:54.957407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:54.957407Z digest=sha256:0fc4558afad65206610c755bd4e88c1548df792c6058ba03936c9aad7c55d288

Observation 2bce66b4-e12c-4561-96fb-633b72938e40 · outbound

This paper cites Rearrangement: A Challenge for Embodied AI.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Rearrangement: A Challenge for Embodied AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.061063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.061063Z digest=sha256:e9255acd0b0ca581c300e6d8a931619ed6f6fa8d627500136321f88ce5d12cb4

Observation 15921749-94dc-40da-880f-362e189737d1 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.137137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.137137Z digest=sha256:23b4f4d1ae3b697b14113e457a2dc73b2ac45abc587024298dfc83b2dd839c0c

Observation bba39573-5795-4d24-8a53-c3c84d3a8ba0 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.233055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.233055Z digest=sha256:f3aa4ce997905951bc331c641a4a61cfc6929e4a11c319f4f25d412153d1a242

Observation 5f10fd22-c294-4a2a-88ce-f1e5cf49199f · outbound

This paper cites Palm-e: An embodied multimodal language model.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Palm-e: An embodied multimodal language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.328013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.328013Z digest=sha256:66cb3c551f0beee048a5a6e7a9f0b47158df7d5f4ad1d551e8d2449097c9ff89

Observation 09681592-73b2-42d2-b461-4ec0bcecad87 · outbound

This paper cites Large language models as gen- eralizable policies for embodied tasks.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Large language models as gen- eralizable policies for embodied tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.428981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.428981Z digest=sha256:786c3a8c76460ebfeed4443c840944bbd385289e54c237aafc9ba4b749ca27b7

Observation 2e61dc64-8296-43f2-8a98-8aa3e3fab4a4 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.547151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.547151Z digest=sha256:2ba441159cc3657e9a17e774baf5801b6b37198ce8b5ae4b4fc9acf6395c6f3d

Observation 84fa2242-9b03-4602-95ac-7cedd8676442 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Gemma: Open Models Based on Gemini Research and Technology

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.659073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.659073Z digest=sha256:28cece8bd73deb781b851f4463ebabf04b8f33a0a2704e6da335312c65ea808d

Observation c8d01f6a-31c6-4a0a-a29e-ca21030d8ab8 · outbound

This paper cites Qwen Technical Report.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Qwen Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.744626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.744626Z digest=sha256:75e5fd4531d51379b52a66561189ba753bc6afd8c92d7d3a4e82cbe7ca594a3b

Observation 03a7cf90-24c6-465b-b207-a1ac23c24956 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Physically grounded vision-language models for robotic manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.818724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.818724Z digest=sha256:9cc150b2210c9e64f3a73ec2133c66f21caf74e4fb9d69e663af70433392008f

Observation d8505a54-c944-4812-b1ab-56058bdd0763 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.891849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.891849Z digest=sha256:5db89ebbb606188bbd6445e29e2025446c307d56bfb1ad85628e2b9a45770ff3

Observation cd42b183-0e0d-4c4f-807a-e76342afb5fd · outbound

This paper cites Multi-skill Mobile Manipulation for Object Rearrangement.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Multi-skill Mobile Manipulation for Object Rearrangement

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:55.986927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:55.986927Z digest=sha256:57dd69ac99f5e88d0178d91467298461526e19085f35d00e44d5c8016d7eaf41

Observation 9c452002-c1d8-4c61-81e2-0d5069078971 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.087735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.087735Z digest=sha256:770e9f8e215909bfda7c94f11cc598ac0ebb04571b9cafd1b575e3d2d9ae18cf

Observation cdf95c37-3141-4763-9c69-7a7712077c8f · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.192260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.192260Z digest=sha256:fa1d2bc83e6b450bb7f22619b9f2778369bc030a36da7b8da990512fcbd04cde

Observation a31e1bfd-41f2-48ea-a237-abce7c0f9499 · outbound

This paper cites M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:54:11.656017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:56.263491Z digest=sha256:0d425adc7966abae8eaf5441029a1e88f417c04853f34a115a5d66c3b96e4e50

Observation fefe6b08-2850-4476-8406-fcf0e1e875a4 · outbound

This paper cites Embodied agent interface: Benchmarking llms for embodied decision making.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Embodied agent interface: Benchmarking llms for embodied decision making

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.353324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.353324Z digest=sha256:3b862838c807048f20bb7f8017dbba0227382b2de038738e9003e9202cc650aa

Observation 9a2f7942-6d4d-4a4a-8a23-45cfbbc8e047 · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.447497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.447497Z digest=sha256:89dc57655af155f27b7c5bfac8f01c7181832ae883ec5c1b5e0f384b47a00671

Observation a852e5a6-9764-4616-a394-990e93212934 · outbound

This paper cites LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.545921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.545921Z digest=sha256:7b431ad5ef5c4bd4de0f6e4fe71420b482aae3147e43f5a3fd1719b2fad797ba

Observation 9c824d0b-2e4d-43c7-b157-8e79aefbaee2 · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.725406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.725406Z digest=sha256:3879920535273cb5d07d42fc8398a938a430b0af44167395311d0545e1b42232

Observation 19189071-c053-470b-8c64-f823e3ed0860 · outbound

This paper cites EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.823209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.823209Z digest=sha256:bf040bacfb4435193c14cb2271c54d9b934bc08af81a5f900f53cdcfbca5a0ef

Observation 6b626a01-d55b-48c7-93c5-59eb075399df · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Habitat 2.0: Training home assistants to rearrange their habitat

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:56.922046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:56.922046Z digest=sha256:21bbd5ce5485ee2a813ec28a97cfc335aeceb9f70be8a9ffa855acfe294e7d44

Observation 15c6aa96-a5b1-4838-a960-763c259df222 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.005818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.005818Z digest=sha256:2943acf04ef31a25295ac5d5d2bb6fc0c42099a68ff42080d8661809207608ea

Observation 80dae249-2d41-4925-af64-4a7145de8b2d · outbound

This paper cites Sun rgb-d: A rgb-d scene understand- ing benchmark suite.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Sun rgb-d: A rgb-d scene understand- ing benchmark suite

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.095946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.095946Z digest=sha256:79161092a8f2c899a2b393e7e162b9684e31005b22af4985c3c56de2735f535d

Observation cde03339-dc71-4810-ba10-3274b15fde30 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.168766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.168766Z digest=sha256:a055c220001cfd1b4f36989753da88f2ea3e4e83768723a429c1494789e5b990

Observation 589b5135-de90-4db7-8927-50b5f19f43a7 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making React: Synergizing reasoning and acting in language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.240312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.240312Z digest=sha256:d23b6b80e56685d2d8373e8a140fa4bb697d7612a835963beca0d5f241fa71a1

Observation 87457872-452d-4681-a83a-e7d9e985fc23 · outbound

This paper cites AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.326324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.326324Z digest=sha256:e30d284c2477da807b9b0e3d614d78a70d9851e0e2d90038f80bd3672ab6d0f7

Observation 26a37aa7-7d09-465c-99a3-a8bb426f167b · outbound

This paper cites Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.380409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.380409Z digest=sha256:b4df7f78dd8f0990d5504ece68fbca0eb35df27b7b154d17a1b56425c27fce27

Observation 20ed65eb-b100-49cb-a67f-ed491cfd3b57 · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Leveraging procedural generation to benchmark reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.458580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.458580Z digest=sha256:48dfe30ec1e0ca03cbfcc7c9ff594e969b82ccef5e6ce6c15c23af53cb0395e3

Observation d8cb284d-c321-4eff-9481-1f696a007576 · outbound

This paper cites RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.551508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.551508Z digest=sha256:cef0556235176d225cb31322029ead845f3f90402cbb94fdf291c05f7cfad07c

Observation eb3a3191-31d9-40eb-979f-18817f69b652 · outbound

This paper cites Adaptive Procedural Task Generation for Hard-Exploration Problems.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Adaptive Procedural Task Generation for Hard-Exploration Problems

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:54:11.407001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:57.649094Z digest=sha256:a1948b9f3a11f4d31e326d74cbe1909feba48ef03e7f948b342c851b5a320bca

Observation 60942058-9243-4d47-908c-449eb87b0c8a · outbound

This paper cites GenSim: Generating Robotic Simulation Tasks via Large Language Models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making GenSim: Generating Robotic Simulation Tasks via Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.740352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.740352Z digest=sha256:10225c3bb5cd8371da4dd754c7f43783929fc910175358c00818d1374f02dc73

Observation 9cbbf7f5-23a6-4329-86b8-4accebf656b3 · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.835282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.835282Z digest=sha256:18451f2032b134c535a8aa599fc0d49ce612d31761e8979f302c5da611fef08f

Observation 491d1c65-0b12-4e12-b7e2-732400a61c2d · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:57.932255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:57.932255Z digest=sha256:06250e1043d2c9d190d79a2ac69c8d9e53a8fb95954955333a186cb28cffcac7

Observation 8e9a6e81-93f2-4cbf-b8d8-6067f5312c51 · outbound

This paper cites Habitat rearrangement challenge 2022.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Habitat rearrangement challenge 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.573015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:58.002577Z digest=sha256:d710b2442eebd9ee5b6db570d4b26521432167b6d87d3443620d6934072ea814

Observation 65561e2a-1be2-4230-8bd2-485387ff7555 · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.124688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.124688Z digest=sha256:47ac6fd4d9b6da01a7e446594a952a439452277f063b0d4a28b6c2086f4bcc32

Observation 860b9867-74b1-4221-abff-77adbc88d483 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.184489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.184489Z digest=sha256:ac499a033cc61c39ad06daeca1036ee42b418f3f4ca627927acf9f9ac0195ac8

Observation a28e3a21-dbaa-4511-bd18-61b1703b31b5 · outbound

This paper cites {\lambda}: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making {\lambda}: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.275583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.275583Z digest=sha256:f23ba6fc9144a17b7cf0ecf27c09fb2d6590c9135593bf8ceccc24ecd1a0cad0

Observation 022070a7-067d-4dba-85cd-84da8c35bcd0 · outbound

This paper cites GPT-4 Technical Report.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making GPT-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.405553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.405553Z digest=sha256:1c3472598afc8c5e95919d3b685786eeab7f0b301ee7770928729c1cb2238d6c

Observation 157df454-c573-47a6-8b78-4a8671dc9d89 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.489673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.489673Z digest=sha256:71be0e79228ffa059fab380296ba0d55cc447919d677ad61e46cb34a7bfe996a

Observation 2bba0d62-520d-49ed-a488-3294bd16f4fa · outbound

This paper cites About claude models.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making About claude models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.408060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:58.621363Z digest=sha256:03c96a7fee9e3fd8d5268a106f534e3cfd7b54dc26ac022bb63ba5e360c7de9d

Observation b1723ad4-33a5-4287-b4c8-dbe1d937df6d · outbound

This paper cites Sapien: A simulated part-based interactive environ- ment.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Sapien: A simulated part-based interactive environ- ment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.722494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.722494Z digest=sha256:1d8a692dab2154443faf1ac893b966fe5fc008a624731f2d038aba655406d35b

Observation 31cfd8f5-151f-467f-84ef-371c1cbc174a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.811914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.811914Z digest=sha256:70c3d3e2f35de27eea912c631f188b0dad94f568a338fe68182dfbd84d1625b2

Observation 98e7bdec-d034-4e89-b96f-203611b907ca · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Finetuned Language Models Are Zero-Shot Learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:58.927954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:58.927954Z digest=sha256:4752ff447dd8814edd0377f8e212870fcacef685b446b8fd620bf47dac0a1694

Observation d0e6ad65-1dd7-4515-8481-bc19822d0031 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making The flan collection: Designing data and methods for effective instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.022673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.022673Z digest=sha256:947f5bbfb23932438ca77fd848bee7a0fa24a403ebfc26ed3193890ebd20cc06

Observation 08a4195a-2f23-4da7-8006-6e5654bf458c · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.239608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:59.119195Z digest=sha256:b72b52c5d49cbfa27415255b3a6bebf8097a0066a073af76c973c2cde6211b9b

Observation 08b3811a-e0fc-4c3e-961b-604bfd8aff1b · outbound

This paper cites Cogagent: A visual language model for gui agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Cogagent: A visual language model for gui agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.194279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.194279Z digest=sha256:14a9701fc8f8517cf1d0ad8c4c0ce3eed01a6200f61048c10c1272fe7c68c065

Observation 69e68284-c532-4ca1-8d46-b976065c6844 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Autonomous Evaluation and Refinement of Digital Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.280411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.280411Z digest=sha256:5c3a328fc2449b4715d1649a1158566dd97200d802dc82ceaddc094d24f208f7

Observation d478c978-2271-4eef-acb5-47450fea440c · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making ReFT: Reasoning with Reinforced Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.379335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.379335Z digest=sha256:4ca15ec525d4f0979f7c9b426ebde130b19da39ba20501f2570a28210e87e452

Observation b0817a52-26e0-459f-93af-865c92ee0d83 · outbound

This paper cites Grounding multimodal llms to embodied agents that ask for help with reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Grounding multimodal llms to embodied agents that ask for help with reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.449785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.449785Z digest=sha256:3baef80213f08a4092149a865c41c04605aad699408a26325a075c54c030bda3

Observation 507ccd29-01c6-4931-9078-8b810bcaa289 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.540056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.540056Z digest=sha256:200b1dc3e634136130057f273b48656605e69693e4c5b4f150e70e3ba38645de

Observation 739da9cf-ded9-4927-93c8-05911f78a8c3 · outbound

This paper cites Grounding multimodal large language models in actions.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Grounding multimodal large language models in actions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:19.060631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:59.614383Z digest=sha256:6a7aad03272a40f28fd811e1bc84d138ac3d86e26af359886d35cfa5102bbd42

Observation e8dd4e21-f13b-468b-b1b5-90ce35bf3098 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.697480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.697480Z digest=sha256:7c854a27c435ecf482b3cc169629cff8dc18dbabb58af8cd8ec16919fddd4655

Observation c3f27528-3245-4652-bff7-d2160513b6c1 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.798027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.798027Z digest=sha256:ce1742d371c350039ff07fca09c92ef549c79345c1d690a6343fda95f504500a

Observation cc9773c0-28d8-4aa8-9f2b-d683421c85af · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Fine-Tuning Language Models from Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.863345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.863345Z digest=sha256:2d15048d971d27691b2bd8140e6898b44cd5acbbe90f80baddbcac5c309199bb

Observation 3cfe34e8-1326-4792-b080-df0193139256 · outbound

This paper cites VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:54:10.983636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:53:59.934887Z digest=sha256:a420c18a03240001f85d443cf9ef9ba0ab824fbb2dc305fe4fe7116b79ffcf97

Observation 9837c052-b1f1-4fb9-a027-bd72f5b4a1be · outbound

This paper cites Training language models to follow instructions with human feedback.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Training language models to follow instructions with human feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:00.020348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:00.020348Z digest=sha256:78ad6a8bda3a5497ad4bb9a5044fdf8738b23ed28b446b00e5971fc9c638736c

Observation aaa36dec-fb2f-4a78-b795-9f6d5efe4387 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Grounding large language models in interactive environments with online reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:18.904679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.341575Z digest=sha256:c9630c061bf32496d15feb303c68e1f9d443dad33cdd7c4dd299926fae465a15

Observation 97f0bd1b-2be7-4665-bbc7-87b794e22a3a · outbound

This paper cites Amazon mechanical turk.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Amazon mechanical turk

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:18.764475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.415231Z digest=sha256:a0176630c31ba1698a46e175c0badff2cd1f4897302fcacfcc25e7df17cb1716

Observation 17e23274-fe3a-44e7-9c82-fd81388d39ad · outbound

This paper cites center" region. For the surrounding eight regions, we designate their directions based on the surface’s heading: the direction aligned with the heading is labeled.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making center" region. For the surrounding eight regions, we designate their directions based on the surface’s heading: the direction aligned with the heading is labeled

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:18.618071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.501803Z digest=sha256:f3477166b1b84223e04d538d6780bd6880ee9656c8cb68916e6f143e84c0c084

Observation fa5d6184-56d2-4c9a-98bb-686c0cf15a7d · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:18.421170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.614772Z digest=sha256:4861d8a70f7defa7457b7a36fd32bcccd2bd61954c503dd4a9a16fbc743cc20d

Observation bfb85de4-ccee-40dd-973b-7b0456d5cd73 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:18.276703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.709463Z digest=sha256:629416f1a1724ca3fa43fbab9bdd6aa0009fc0ad3969282f5414e4442a1b08a3

Observation dd4dc1dd-4ff8-4f29-b379-85e0078a52b9 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:18.127181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.837553Z digest=sha256:dc7fcc6e9dcae06a9befc5eab7ab679d61d2683cdbcf57a08246a5a33dc6d02f

Observation 63d23804-baa1-48c3-a0ee-babf8dd229f7 · outbound

This paper cites Feasible.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Feasible

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.962332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.916794Z digest=sha256:97d8a3e5905005dd614132cab6cfae623ac4eb91c2f325d914a98b214daf1b87

Observation 56623e55-9369-4f8b-accb-fe24a4275571 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:17.791018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:00.997321Z digest=sha256:a4fa8d58373b38354335a073a19ad86b019fc660aff311474466b111d12f64bf

Observation caa15747-9161-4e46-9752-15250857f940 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.651016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:01.084572Z digest=sha256:29ec6474a49915347ddb6d3fca9e14fb401593d65bdcb264db8bd36d90f4f0a2

Observation 76a29f80-09f8-458f-a15f-0d45c2f7eb10 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.486413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:01.745645Z digest=sha256:0cd6f5ad9061ed7e842bd97b6004e5fa8af83d2c8f511f8faf91021aac5a2861

Observation 3dbe886e-a1d0-4a99-ab0f-03b037400551 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.289911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:02.163867Z digest=sha256:3cd77ddd590d322ee55a808434cce59aae2e0a3bbbc7ac2b2719940fbbd5038f

Observation 4ede6929-eb16-406b-a8e7-a0ab6766d809 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.066562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:02.521528Z digest=sha256:0bfbe5a0ca8dac6dba502763335fd80296267f201930c944ec12322098401738

Observation cb745ab7-32d5-4804-8ffc-bec17e621db8 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.819053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:02.955476Z digest=sha256:88116592528a56c72a6f1efb0a9a97e1129bb025389f8d6810bf1f404423bf54

Observation bb20f84f-0ff2-45db-9f6c-3acd49494a80 · outbound

This paper cites Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.624389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:03.173448Z digest=sha256:2387b2c39b83642efc1cacb203c45556cb9b2159363f7d06dd341011a79c92a9

Observation f6e1e9f0-d471-4676-99fa-b272fa9438e3 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.431290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:03.392298Z digest=sha256:1d0ef8f2254e174c518cd9b55f7b3acc33e44df56a41a78392785f61794fba2b

Observation c5c672cc-6e64-4da0-84a7-4cd28d662672 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.184493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:03.786523Z digest=sha256:e37b5e194c92fa1f68814737e01aadf236ba7ccc651ff3082108ec47dcc0b1bb

Observation da61c604-ee1d-4f79-b563-07b1714ac0da · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:15.007876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:04.211528Z digest=sha256:b220ca32a06e6cf5c306541ce7b209ddcf7c2b3e5c3795f73d06081a78d10d21

Observation fba5372f-70b9-4ad6-a192-4c4d43f0e07f · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.764957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:04.629175Z digest=sha256:8926396836a13f9a95a940d709432e9f33cebb22b679ee611fe75179abafb54f

Observation aa90358d-d8ee-4b60-ade0-f37056f8815a · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 106

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.589581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:05.152138Z digest=sha256:0195c1d89c72cc1a0ec2ab39e87a0a29ca5c09db3828c030e4d180c7a056a0df

Observation 0cbccca1-3d35-4a1f-acd1-2a7aecbb5df7 · outbound

This paper cites This is important because the receptacles may not be intuitive.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making This is important because the receptacles may not be intuitive

Reference 109

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.421053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:05.467669Z digest=sha256:2fb1d4cf9268868da1e543ae9faa07334aae2e322238537489cdc196a49eb210

Observation f7fbb0c8-4d20-4986-a7f9-e411d2e87971 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 110

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:14.201698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:05.580311Z digest=sha256:bde1e1fc2b6a9903ee88b5a767f1d0340da02d479e9137c591180b8f654938b4

Observation 12af8a4e-7f72-4ffb-9cc5-6baf120c27a2 · outbound

This paper cites show_receptacle.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptacle

Reference 114

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.919640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:05.985117Z digest=sha256:8c47c7a8b5f3774f04ec6ebfc5af5cd2ad2ecf45f73e7860f12c4f07a124c2bb

Observation ead19c69-2f99-4516-a993-f5b52ddd15b4 · outbound

This paper cites show_receptacle.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptacle

Reference 118

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.697302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:06.393888Z digest=sha256:37fffb2d600b4a8711c70a94cfdbff1c155cc96fe76aed8e297fbcdbf01e5af5

Observation 7a66cf5d-d623-4805-b8fe-4835337b3ccc · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 122

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.517228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:06.638527Z digest=sha256:e9fe5b3644ec4e5175eae51651d7770be545ccf9cd95c131fc199b5fd3a7e794

Observation 9f298546-1cb7-4b93-9abc-c22bebdfecbe · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 126

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.336685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:06.755301Z digest=sha256:3292f981e743ac311433ec45d6b6f26557a88f44227d7be44b0f1163e699fdff

Observation e43a41f9-68d3-4081-a995-b1111a8aabc4 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 127

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:17.472081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:06.851144Z digest=sha256:44df6a4a727c596367073ef9d1021a9a24ee9684836db3067e2409b07be33a16

Observation eec26299-e387-427c-b8b7-1a9b33773c66 · outbound

This paper cites If the platform has no objects, a 3x3 grid will be marked on the platform to help you place objects.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making If the platform has no objects, a 3x3 grid will be marked on the platform to help you place objects

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.316512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:06.945591Z digest=sha256:2a9b74a7f3224a9225361c87acb21dede0961cadf1b206d03e9cc62bd1f3a82c

Observation b476bee6-c400-453d-be23-6de13f74f7eb · outbound

This paper cites show_receptable_of_object_x_of_current_platform.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptable_of_object_x_of_current_platform

Reference 129

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:17.135743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:07.075792Z digest=sha256:0e232aeb55424c94bf2467a9b439fbea55d8004965b53b3c569f0e6d25ffef0e

Observation 8dfe0851-fa84-4ce0-b662-4b3df8ab90e7 · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 133

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.194035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:07.659565Z digest=sha256:274c81f24e205fb49e9ec287ad17113402175fec0a113cc28cc860ba61d0990e

Observation 66eefccd-e5da-4293-8e24-044ce22ff8f5 · outbound

This paper cites This is important because the receptacles may not be intuitive.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making This is important because the receptacles may not be intuitive

Reference 136

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.660557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:08.078388Z digest=sha256:eab5b5b3a19216deeb6a24610669692d0134222030b8706a41bbed599f4693d2

Observation 5da1660c-31f1-466e-9a6d-ece6b36564c0 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 137

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:13.014392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:08.206444Z digest=sha256:22a19a9de3944c43c1309acbb9519203a736ff878c33c99d54505d64788b2609

Observation a54dc73d-f3d1-4427-ac0e-21d48fce2c89 · outbound

This paper cites front," with the remaining regions proceeding counterclockwise as.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making front," with the remaining regions proceeding counterclockwise as

Reference 141

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.632092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:08.812446Z digest=sha256:284996274dd381b1f62630cff12407896430ae52832fd0ceefdaaa48278c64e7

Observation b5b23e3f-e249-41a7-81d3-809e929a45e9 · outbound

This paper cites Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful

Reference 143

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.491374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:09.077039Z digest=sha256:18a99c49ec03d837de02d0075fa62e8e486112e292b513dcb5de3edf8c300c4a

Observation 68d9d04a-736e-43af-bb73-40c35be53013 · outbound

This paper cites show_receptacles.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making show_receptacles

Reference 145

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.355455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:09.396524Z digest=sha256:2791010c1b9b4cd9ccb73a14bda58a0ae4aa9ed7c6a198a44dc06c35e0444f44

Observation 1d4bb4c1-eb42-43b2-895a-96a2038cdf3f · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 149

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.132558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:10.097488Z digest=sha256:ba900a9756655955ab793ec3987988a222a932be034e373a364decce02ece18f

Observation 0b37ff0c-3379-4de1-ab10-80215710c272 · outbound

This paper cites an unresolved cited work.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Unresolved cited work

Reference 150

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:54:16.981330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:10.216972Z digest=sha256:525cc3e31ff10151b26b42c4c9725e30ec814bbcc7d5f5eaf02a42615df72433

Observation 51d0f704-bee8-46a3-97cf-f6242392fab8 · outbound

This paper cites Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making Specifically, for the task asking you put object to empty platforms, try combining adjacent receptacles may be very useful

Reference 151

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:16.837465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:10.327005Z digest=sha256:0f3892d41b573b7a543e17a0caca7b5b40d4834820881e56bac88cb79ea1cec8

Observation 539127ba-b0df-447d-be4f-332f0db22ebe · outbound

This paper cites This is important because the regions may not be intuitive.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making This is important because the regions may not be intuitive

Reference 152

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:12.851308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:10.473014Z digest=sha256:cabce3888680af6e475bad85cf663b912f88684e8d55135b506de3b6d2c882bd

Observation 36847eec-c898-443d-a95f-98b92e2b42ac · outbound

This paper cites You will only receive the same hint informing you your invalid action.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making You will only receive the same hint informing you your invalid action

Reference 153

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:54:11.932375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:54:10.639156Z digest=sha256:cf10f359060e8c1ef7659492b7dc12f420f3df161e63d40be8c8dc9cf1bc58dc

Pith citing papers

Observation 62e2116f-9236-46de-9cc1-b77ce111f607 · inbound

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation cites this paper.

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:12.440926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:12.440926Z digest=sha256:4d2843c07542ae6e6f66411bc781ebd35bcb3aa42fbac322ad6fb8132c197028

Observation 7b68a88f-8926-40c1-8f47-cbfa9a74a8f3 · inbound

Compiling and Benchmarking Task-State Horizons for Embodied Agents cites this paper.

Compiling and Benchmarking Task-State Horizons for Embodied Agents ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:36:21.671674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:36:21.457436Z digest=sha256:d77e0bbcab03f33a8aa0aa67911039f239e96ddec2775dc82347c2b88b30c118