Pith. sign in

Paper Citation Record · LEDGER

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2507.12391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12391 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:09.623638Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:40:21.327031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:40:25.831314Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08cbf44a-4b00-4086-b09c-742698c10219 · outbound

This paper cites A survey on large language model based autonomous agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning A survey on large language model based autonomous agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.672882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.398980Z digest=sha256:e97c38d5b90b6b3fc66ddac4a59873dbbe76fc7b6bb9a3ffc84ccd54f445fb6e

Observation 9986f59e-4f4c-475a-965f-7ee8a08b25c0 · outbound

This paper cites Large language models for robotics: Opportunities, challenges, and perspectives.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Large language models for robotics: Opportunities, challenges, and perspectives

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.514886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.474604Z digest=sha256:43530e1e27b1d073dc71c508895c4b08983fdd340ecabb00fcb1d501650023f5

Observation c65b190b-166d-4fcc-a5f3-a3722221d776 · outbound

This paper cites Human-Robot collaboration in surgery: Advances and challenges towards autonomous surgical assis- tants.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Human-Robot collaboration in surgery: Advances and challenges towards autonomous surgical assis- tants

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.354229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.565616Z digest=sha256:1f0f3718eb061a2cfa8c7b9b56ce9926934a3a088f40bdcc47c21911add49d17

Observation 57623a7e-2b68-4061-abe0-f2ee19db3928 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.147985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.700942Z digest=sha256:1035461585ec5ee4caab74b5a8c92b6d6a0ca1956f4455fd14341862156a69ae

Observation cbc1c722-8cdf-4b41-b90a-9a502e50e7a8 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.952899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.858503Z digest=sha256:b16b8f241d6f89f044b51013409d9180496a160bdd33447a0f047cf292549efe

Observation 3f2e3fba-f0c5-4623-88d5-2bbb2c9d94e3 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Progprompt: Generating situated robot task plans using large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.717772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.985416Z digest=sha256:935b1b50ebbc1c9f0296abc34c053bbcfe6d5e51a7d4ac043c6d8bc253c21200

Observation 19a8e19c-e42c-4bdf-ab88-6750cce49d7a · outbound

This paper cites V oice con- trol interface for surgical robot assistants.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning V oice con- trol interface for surgical robot assistants

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.092121Z digest=sha256:05928082209eeda631da564c87e57618116ed8a5bffa696891c7f21793a2cf76

Observation a94424e7-7caa-4cee-a8cf-f67ecc0d3653 · outbound

This paper cites LLM-based ambiguity detection in natural language instructions for collaborative surgical robots.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-based ambiguity detection in natural language instructions for collaborative surgical robots

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.129833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.202134Z digest=sha256:ce461504247fb798d3fc943313208869471cd3d6e961ec591683355042c3ea41

Observation 1a42df65-51b0-4524-ba32-23ca0dd7a425 · outbound

This paper cites Toward autonomous robotic minimally invasive surgery: A hybrid framework combining task-motion planning and dynamic behavior trees.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Toward autonomous robotic minimally invasive surgery: A hybrid framework combining task-motion planning and dynamic behavior trees

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:10.835604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.392134Z digest=sha256:e197c5b049e4b3f3e5f0955ff0276cc464d8c58e129b4487b7f7f0dc4fea643e

Observation 4a39828d-d7e5-453d-88f5-03b56020e6eb · outbound

This paper cites Exploring embodied mul- timodal large models: Development, datasets, and future directions.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Exploring embodied mul- timodal large models: Development, datasets, and future directions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:10.620752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.518679Z digest=sha256:5e72f439fe193e6373a752b3c8b4bb64a4ffdb2a3cb7ed76c1c04e76dfaadcfe

Observation 47faf2df-fb07-430d-ad0a-1957377fb85c · outbound

This paper cites Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.597731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.597731Z digest=sha256:c1638c438c73f2a608acb144aeaf4b29cc94aa80309b66c341d8fc1250283f7a

Observation 16c1735e-0042-4033-abad-8427410ddd74 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.663316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.663316Z digest=sha256:9cf86ce0f3636c59fb7abb3da1054e884d467ff6dd0101a5e54038fdaac8adbb

Observation 07c40737-2ab5-4210-8b4e-34d859072622 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.750449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.750449Z digest=sha256:6e2fea8f2a2a5085fc6663f3a85f3bc07261459b831cff5bfcfd1b8720a93352

Observation 3c1d1cae-4389-435f-a7d6-0e4ba6a3e83e · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.815246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.815246Z digest=sha256:28aad98cc371e52879f01c5d8dff241bef15492d8754e4ec27c29014fcd2e04f

Observation f1d0268a-a77f-4a04-9f93-2964c0b7155d · outbound

This paper cites LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.926716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.926716Z digest=sha256:baf5ce712199957053506079acd940fbc3b68b81e613c17618b43cb89486fd6f

Observation 3ccd6026-280f-407f-892c-740167a13455 · outbound

This paper cites LLM-Advisor: An LLM Benchmark for Cost-efficient Path Planning across Multiple Terrains.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-Advisor: An LLM Benchmark for Cost-efficient Path Planning across Multiple Terrains

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:50:10.352632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:09.019979Z digest=sha256:0ef2009e04314ff99d027a452a55c7fd843302369ac17565cbec78c74d95f68f

Observation 622f095f-1e64-4971-a277-1bd0feddc52a · outbound

This paper cites LLM-Enhanced Path Planning: Safe and Efficient Autonomous Navigation with Instructional Inputs.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-Enhanced Path Planning: Safe and Efficient Autonomous Navigation with Instructional Inputs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.102714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.102714Z digest=sha256:8891bc25aa17d4360cc6b7baf7459622a4e537f0ee62db32b39277314b14af31

Observation 235df9a0-fdba-4ab2-a74f-909b7a894a4e · outbound

This paper cites Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:09.172796Z digest=sha256:90b232235a269ac21386100205449063f261a8a19b3877e7bc12af25c8a00292

Observation acc50030-12b0-4839-a227-482856a712ca · outbound

This paper cites From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.264012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.264012Z digest=sha256:90615ae873944045af803ed4d232ce8dadf4dc26c546020852f82634786216eb

Observation eeccf834-057c-4c69-be2d-ae4783029d6c · outbound

This paper cites Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.379020Z digest=sha256:35e304c0576d4f390c3c823e079d0323ebaff53b61c204c1624a6063c82b50a8

Observation 6367c078-f26e-4924-bb48-ed3c51739195 · outbound

This paper cites Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance- Guided, Self-Consistent MLLMs for Food Prepara- tion Task Planning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance- Guided, Self-Consistent MLLMs for Food Prepara- tion Task Planning

Reference 21

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:50:09.900096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:09.470153Z digest=sha256:cc6e9ea283cbd8f8252b9fc3bf0cd139e0dad9f77401ee5c89c6902fe901c992

Observation 7f6edbf4-224d-4902-ae90-0d0f32c2d5d7 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.535109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.535109Z digest=sha256:606b8fdd9c8ad85dfbfd700e0bd3c3b1b6f138163114527575f651a00e8b2b80

Observation 20a633f0-d5a2-4c3f-93ba-942cbab5df6d · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.623638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.623638Z digest=sha256:6d131ff5e5dc56e41038d7e7ef43a20ea0909967ace273f70e7d47bfe47878a2

Pith citing papers

Observation ff00f254-9fce-4018-87cd-dad413f06d7e · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:40:25.837858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.327031Z digest=sha256:af262a2e88c257236f09a031b4751a7939b4e5111e026b024f26eeee8a9d2168