Pith. sign in

Paper Citation Record · LEDGER

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2507.12391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12391 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:09.623638Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:40:21.327031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:40:25.831314Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08cbf44a-4b00-4086-b09c-742698c10219 · outbound

This paper cites A survey on large language model based autonomous agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning A survey on large language model based autonomous agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.672882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:07.398980Z digest=sha256:151b14b5ff657a8697a68258db9933e01cd1de3d1bb823ab9f0f62e2f6aed21a

Observation 9986f59e-4f4c-475a-965f-7ee8a08b25c0 · outbound

This paper cites Large language models for robotics: Opportunities, challenges, and perspectives.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Large language models for robotics: Opportunities, challenges, and perspectives

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.514886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:07.474604Z digest=sha256:27dd40f9cf19f98a776742681d2f827d980b566afd597428010794f5d8af5ce9

Observation c65b190b-166d-4fcc-a5f3-a3722221d776 · outbound

This paper cites Human-Robot collaboration in surgery: Advances and challenges towards autonomous surgical assis- tants.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Human-Robot collaboration in surgery: Advances and challenges towards autonomous surgical assis- tants

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.354229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:07.565616Z digest=sha256:e2e586c6c1248db098f527ae64874be97795c9897626dd7e087f932fa3d2ee63

Observation 57623a7e-2b68-4061-abe0-f2ee19db3928 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.147985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:07.700942Z digest=sha256:06d95d5c7c75bf55c6b02b0ca556efdc357788bfe1dbf46d432b0d8ac2a3a556

Observation cbc1c722-8cdf-4b41-b90a-9a502e50e7a8 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.952899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:07.858503Z digest=sha256:6bbceb8acad6747d4459fcc83215e809350d258d563ddbf0547020f1e9c05548

Observation 3f2e3fba-f0c5-4623-88d5-2bbb2c9d94e3 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Progprompt: Generating situated robot task plans using large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.717772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:07.985416Z digest=sha256:e32afc0962138e8686818352aeded760d4877d3c03e2d5201770edf75aa9d597

Observation 19a8e19c-e42c-4bdf-ab88-6750cce49d7a · outbound

This paper cites V oice con- trol interface for surgical robot assistants.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning V oice con- trol interface for surgical robot assistants

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:08.092121Z digest=sha256:1f83bab5ebc0a4f403c529f6130b1b416abd859c110af00f9e97c7846c0badaa

Observation a94424e7-7caa-4cee-a8cf-f67ecc0d3653 · outbound

This paper cites LLM-based ambiguity detection in natural language instructions for collaborative surgical robots.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-based ambiguity detection in natural language instructions for collaborative surgical robots

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.129833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:08.202134Z digest=sha256:1f41d3d782e10e559a449497cf4db634ee75c4b9621946a34c825e252535241b

Observation 1a42df65-51b0-4524-ba32-23ca0dd7a425 · outbound

This paper cites Toward autonomous robotic minimally invasive surgery: A hybrid framework combining task-motion planning and dynamic behavior trees.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Toward autonomous robotic minimally invasive surgery: A hybrid framework combining task-motion planning and dynamic behavior trees

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:10.835604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:08.392134Z digest=sha256:51b2ffb82efd9309a41909b2f932969dee768f0c75f2df597783c355aaff6211

Observation 4a39828d-d7e5-453d-88f5-03b56020e6eb · outbound

This paper cites Exploring embodied mul- timodal large models: Development, datasets, and future directions.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Exploring embodied mul- timodal large models: Development, datasets, and future directions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:10.620752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:08.518679Z digest=sha256:0ab28429991130d10dcb46c4ed007866eec3c27c5c45741aac9406d8fffeb1f4

Observation 47faf2df-fb07-430d-ad0a-1957377fb85c · outbound

This paper cites Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.597731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.597731Z digest=sha256:3682af8f12c47f87dc40faa75ea19c236e577362b33ec9c2f04627fcacfbfbec

Observation 16c1735e-0042-4033-abad-8427410ddd74 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.663316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.663316Z digest=sha256:1beba62d8e91c477b12dd9e487452decd14aa2984c9554c42c38f83991c620dc

Observation 07c40737-2ab5-4210-8b4e-34d859072622 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.750449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.750449Z digest=sha256:0cd19aab7b96c8e632433e2d59ecf8f77cca76a5ae7b5e53eeecd873f9b3ff3d

Observation 3c1d1cae-4389-435f-a7d6-0e4ba6a3e83e · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.815246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.815246Z digest=sha256:b5de10ed27a909569f3307570ac03b655637a7deb5f02996cd62ef59bd1c00e6

Observation f1d0268a-a77f-4a04-9f93-2964c0b7155d · outbound

This paper cites LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.926716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.926716Z digest=sha256:fb023a2e5e66327da76e65259eb39ce03de74ef5de165b71a78390463266ba75

Observation 3ccd6026-280f-407f-892c-740167a13455 · outbound

This paper cites LLM-Advisor: An LLM Advisor for Cost-efficient Path Planning across Multiple Terrains.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-Advisor: An LLM Advisor for Cost-efficient Path Planning across Multiple Terrains

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-17T02:11:25.293435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:09.019979Z digest=sha256:f3c12645b2c31fb9c0747677e9643466745a2f8c148f03b0d4b35bf80239a6f0

Observation 622f095f-1e64-4971-a277-1bd0feddc52a · outbound

This paper cites LLM-Enhanced Path Planning: Safe and Efficient Autonomous Navigation with Instructional Inputs.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-Enhanced Path Planning: Safe and Efficient Autonomous Navigation with Instructional Inputs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.102714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.102714Z digest=sha256:5b3869eec7e1721a4450099061c06269d304d5301d0593aee7db3b7195e241e5

Observation 235df9a0-fdba-4ab2-a74f-909b7a894a4e · outbound

This paper cites Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:09.172796Z digest=sha256:0f925be17b13ee71fa293d9ebe5911c56776c273b5ec6bc44b9c1b33d3696b78

Observation acc50030-12b0-4839-a227-482856a712ca · outbound

This paper cites From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.264012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.264012Z digest=sha256:0dddd336bd3e9cafba005ad7aa4c445905dc1a629194d685d839fd6e0176db1b

Observation eeccf834-057c-4c69-be2d-ae4783029d6c · outbound

This paper cites Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.379020Z digest=sha256:e48dc50258e3c69101ce45b76d2afb7fc7e1379bfb011d205e7a3867309c0dda

Observation 6367c078-f26e-4924-bb48-ed3c51739195 · outbound

This paper cites Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance- Guided, Self-Consistent MLLMs for Food Prepara- tion Task Planning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance- Guided, Self-Consistent MLLMs for Food Prepara- tion Task Planning

Reference 21

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:50:09.900096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:50:09.470153Z digest=sha256:4082d262f822487f08aba4a6f6a56f508a53ddeb432ab5fc359c99d81f396568

Observation 7f6edbf4-224d-4902-ae90-0d0f32c2d5d7 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.535109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.535109Z digest=sha256:d0f91dd768f74324cb737d644f539f6adb8db9708636def5d38d9ded08cd70c2

Observation 20a633f0-d5a2-4c3f-93ba-942cbab5df6d · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.623638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.623638Z digest=sha256:d4880162c60dd3df6198d2807bbeef7ca00467d3f16cb32b7f21647e639b5ae4

Pith citing papers

Observation ff00f254-9fce-4018-87cd-dad413f06d7e · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:40:25.837858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.327031Z digest=sha256:b3542e3c410855ce97b78fd9575e17a419c6e37c573f6cf62710ec8c9f2a8d43