Pith. sign in

Paper Citation Record · LEDGER

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning

As of 10 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2502.08903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08903 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:21:24.869616Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99e568e2-f83f-431d-ace4-4ad52797d845 · outbound

This paper cites Embodied intelligence toward future smart manufacturing in the era of ai foundation model,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Embodied intelligence toward future smart manufacturing in the era of ai foundation model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.759129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.633975Z digest=sha256:05b01b5e90433e3e1689b4ac8d687eb0ceb6b4327cf0e74266d263354875d602

Observation 51c0a5fc-055d-4859-a414-5c75c9b916ca · outbound

This paper cites Navigating industry 5.0: A survey of key enabling technologies, trends, challenges, and opportunities,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Navigating industry 5.0: A survey of key enabling technologies, trends, challenges, and opportunities,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.640204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.640204Z digest=sha256:6ec0387992bb8ab6e65df032f32a5c605e11297a423df04d4913c3e10dcbfc15

Observation 3465b0df-345b-40e3-aa67-d38d2d3b8a89 · outbound

This paper cites Advanced manufacturing in industry 5.0: A survey of key enabling technologies and future trends,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Advanced manufacturing in industry 5.0: A survey of key enabling technologies and future trends,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.733182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.645694Z digest=sha256:e370edded748e2296bfea925f20ca550ca06e63d7a09bb1c9d50800f36e91b85

Observation 3a4efae1-1a48-41ea-977e-ed6aa77142a7 · outbound

This paper cites Large language models for human-robot interaction: A review,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Large language models for human-robot interaction: A review,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.716684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.651712Z digest=sha256:d41bb7b58c7eb3c5d30fff15fb19ab9951403d586ef6990a6ea73dc35788ed23

Observation de8fb097-6b20-4673-a04d-64a6d8db1463 · outbound

This paper cites A survey of optimization-based task and motion planning: From classical to learning approaches,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning A survey of optimization-based task and motion planning: From classical to learning approaches,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.657401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.657401Z digest=sha256:8f4ccab1164d16f4bdd11bbe919a0dc8f2d4f3296c3d53b43293d84c70329726

Observation 53084b09-2c62-47f4-9372-1669815f2edb · outbound

This paper cites Computer vision techniques in manufacturing,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Computer vision techniques in manufacturing,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.674499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.669227Z digest=sha256:b05c098b78963889166bb6f053364a1b762c478731dc8e8eaff219ceb19b1341

Observation 203fdb2d-1d6c-4b6d-9f3c-88b732fd39f0 · outbound

This paper cites A comprehen- sive study of 3-d vision-based robot manipulation,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning A comprehen- sive study of 3-d vision-based robot manipulation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.658425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.675049Z digest=sha256:4a0da721debf2bda71d7b8e72097c134fd41a0c70c10024c764acbe37e1ec84a

Observation e7c61138-a666-4cd9-a763-4f3da896d222 · outbound

This paper cites Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.690542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.680327Z digest=sha256:88ed3259e561edca60aa2d967a945afa0aa9b01d4d3c34d1e2f54dc94c702e33

Observation 05b98e09-6909-461a-9135-c48742e77be7 · outbound

This paper cites Human–robot object handover: Recent progress and future direction,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Human–robot object handover: Recent progress and future direction,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.642300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.685653Z digest=sha256:8f1a53777e64b3aca1a5831c4e4efae723a14e30bb15ea4115484bbd33d97758

Observation c4234f14-6199-41d4-82f7-9f2161be8ee2 · outbound

This paper cites Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.690968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.690968Z digest=sha256:9659858fc6381f81dc0a12a1aa815772c7d9036155cc928da69485fdb53d4827

Observation 7f89dc9b-5cce-4541-9531-f593264bf14a · outbound

This paper cites Multi-modal feature constraint based tightly coupled monocular visual-lidar odometry and mapping,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Multi-modal feature constraint based tightly coupled monocular visual-lidar odometry and mapping,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.615825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.696367Z digest=sha256:e8a254f2736ae33c536d9e8a5a18b6d5f7c79565d5e8aad18eb61ea4184fc6b0

Observation b6a33347-4d35-4141-b73d-4c17fe331785 · outbound

This paper cites Llm-bt: Performing robotic adaptive tasks based on large language models and behavior trees,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Llm-bt: Performing robotic adaptive tasks based on large language models and behavior trees,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.598148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.701794Z digest=sha256:9c3ce02ffb4b6c8879d56f59abb9e16f2e21b29506b8cf49f506379184d365e4

Observation 2506b0b2-1379-4d9c-8689-affbe82a38dc · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.582169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.706864Z digest=sha256:a9ce9f560e47a56de5af0d75e00531b5b59e4d8287b75337b0e25dc4bb0e4499

Observation be13511d-bc3e-45a3-b473-7bd3ed7a2c0f · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.716904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.716904Z digest=sha256:165f3154930155ea743337ac6c9227216b185722e0377bbc6412303958f0eb9d

Observation c23206f2-6fca-46f8-9667-86428a6b30aa · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Progprompt: Generating situated robot task plans using large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.721994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.721994Z digest=sha256:456dc8093fec262e7fba5bbcf87f353399284f6b9795ca29e1b4ed0be11815ad

Observation 8b45d3cb-8d91-4180-a3d5-ed4b7341ddf9 · outbound

This paper cites 3d- llm: injecting the 3d world into large language models,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning 3d- llm: injecting the 3d world into large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.530677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.727041Z digest=sha256:8e70fa770444cf9181ed66cf5b36d624f95d1f87b1661739aeca2a46d07e9b88

Observation 9f79243e-13d3-4c04-bda7-0faf4516e5f4 · outbound

This paper cites LLMI3D: MLLM-based 3D Perception from a Single 2D Image.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LLMI3D: MLLM-based 3D Perception from a Single 2D Image

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.732040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.732040Z digest=sha256:0e054ebfde11f0b5b27c8a37a2c30fd42a0ee5d8b458974b86c4be45366cc1a7

Observation e4303813-9f78-4651-87a1-57c8dd7caa7e · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Rdt-1b: a diffusion foundation model for bimanual manipulation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.513861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.737749Z digest=sha256:4cdcd44d552296268008c35974a5d194bc0226a20a1efca40bd2c58c04951b37

Observation db939c55-5a15-45c2-bea6-2a2e39185b03 · outbound

This paper cites Efficient prompting for llm-based generative internet of things,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Efficient prompting for llm-based generative internet of things,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.742557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.742557Z digest=sha256:0a49dcac218d91a5e0fa1c1ea32339cd1d9bd666d66788bf95b84a80805092a3

Observation 63d58fa0-3574-4969-ac1b-568fd8927efc · outbound

This paper cites Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.747622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.747622Z digest=sha256:0d34ee741ff7995a14ce12b01c8ee10e2ec88c8c179b2e1870c4d61e2c022388

Observation b3dd1daa-2245-406c-9a63-245a6f9dc0e8 · outbound

This paper cites MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.753064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.753064Z digest=sha256:586a8ff0e704c6c6b4156d6776f1bcae70c06dccc059f726d5b7610c57150d12

Observation 2b21e68d-9dcf-41fb-9b49-debcb5b69290 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.758518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.758518Z digest=sha256:c6a8cebddd29898889484ddf35cba11c98df78dff22be057a5865b06ce01d3e0

Observation 2edb81ef-8720-445f-ac1b-a00f0160b2a8 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.487551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.763904Z digest=sha256:32b572dfe1753cd4b255af94b7098fbfcea0d5a8da193fdfdcc9eae1d6fe056f

Observation 5dac971e-7165-498f-bab5-2a201487cda7 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.768834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.768834Z digest=sha256:95bf076289dc757c2ff8d7f0a93cf18f49aee3da474adf0494dffd6ea1712ec8

Observation 1be82fd3-1d8a-455e-a8f3-e46619735dc8 · outbound

This paper cites Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.774141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.774141Z digest=sha256:b169a923c9945e81e2c0202548becf32050a0929c18ae5266ffe7f96ab7e9f28

Observation 5672150c-e946-4f3c-b69c-2e4663618561 · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.566103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.779551Z digest=sha256:e5913292d5022297d6c236ff269b923f9e01195abe9edb1d72d33ac63f005715

Observation d6f29f30-3c78-4bd7-8c59-bf4e3af90faf · outbound

This paper cites Survey on large language model-enhanced reinforce- ment learning: Concept, taxonomy, and methods,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Survey on large language model-enhanced reinforce- ment learning: Concept, taxonomy, and methods,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.471120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.783804Z digest=sha256:404c5088b994559b9eb4bad6a21afe8e87e57bc041afe812e3468787ffeab124

Observation d0ee76e6-55ff-4f75-bc82-6464aa67469a · outbound

This paper cites To boost zero- shot generalization for embodied reasoning with vision-language pre- training,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning To boost zero- shot generalization for embodied reasoning with vision-language pre- training,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.453133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.788247Z digest=sha256:a45e9decd9841f66047bb9c52800d36f71e7d309478b8e8ddec8275c2ec29b6d

Observation 6bdcfdb4-4f9c-43a8-9975-ecc8eeaabb98 · outbound

This paper cites A survey of visual navigation: From geometry to embodied ai,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning A survey of visual navigation: From geometry to embodied ai,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.436777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.792831Z digest=sha256:b8bd6e9390dda7db428747edec96ea122bb81f5e1c99d9c7618e6a5d207f9170

Observation c56990b9-2c38-4115-b99a-27b781f9bc51 · outbound

This paper cites Delving into multi-modal multi-task foundation models for road scene understanding: From learning paradigm perspectives,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Delving into multi-modal multi-task foundation models for road scene understanding: From learning paradigm perspectives,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.421188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.797181Z digest=sha256:ebe9863136af41ce1e21c46658cd83158ec33f3e922015f16ebdb99004165e86

Observation 8e710c2c-f061-4eec-bc6b-dd7432edafe1 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.801646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.801646Z digest=sha256:faa501860dd36b31cd1251477ce22e3f467eedd6607b89370a9c8e931d53e826

Observation ce9e92b8-da09-4c38-9e66-14355fae9f8e · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.806266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.806266Z digest=sha256:f0062ff00f2f8db22f91cf1d493bc8a784d0eaef6c28cb70b651720f875d8e99

Observation 5eef3c2e-3460-4598-bd12-47a07501e139 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.811203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.811203Z digest=sha256:5a07884eeda00c4b36f685748b8918a8374ac0ce3d3366c64983ec3dba629d20

Observation 31967dc6-4202-4999-9f0f-f97c8f0ba81a · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Interactive planning using large language models for partially observable robotic tasks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.405220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.816705Z digest=sha256:9df744fd1fc10e8b339ddf6868ac23795ccb2eea4777c5ec5fcafee54a02e178

Observation a9d24f4f-73ac-4492-aabf-9e8db7fb4ba3 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.821792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.821792Z digest=sha256:8ed5e186195a4146200810818174288478944c23f0e3862f47a1e0591847137c

Observation 5561590d-e2a0-4435-958a-18c289f144d5 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.826991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.826991Z digest=sha256:8ce76ac87c24eedf9a4fdacfcc83b1b3c839c3f7e918950e95fb647834128192

Observation bf7997dd-d115-4180-8a32-710dd69bab63 · outbound

This paper cites Visual language maps for robot navigation,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Visual language maps for robot navigation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.389029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.832463Z digest=sha256:8c6a4bf521b80d94a269ea0cc1b633dc00c3fc3890f881b4838ce7bf2648e0d0

Observation 2aed1b5c-b8c2-48d6-948e-8cd12ab6780b · outbound

This paper cites Ving: Learning open-world navigation with visual goals,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Ving: Learning open-world navigation with visual goals,

Reference 40

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-07T23:21:25.177254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.837326Z digest=sha256:9665dc809d610aa861a21fb685df0bbeba468b13b969c3eb495fc4b477c80935

Observation ce7cb4ab-9fb5-4273-b845-f81d216c5e45 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.842290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.842290Z digest=sha256:0a4dcc167c2a8d2dc8dd68dbc163c2cbbf8349b9d26fe4cd7d1847b41a65e573

Observation 53a05805-5044-4462-ab8f-e913e47f8993 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Lora: Low-rank adaptation of large language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.847764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.847764Z digest=sha256:c8f06afaef7d3ebfd35a25a3d020b43ec28266f84b16fc30608e56daa013ffb0

Observation 87ac9c5b-7f41-4d79-9682-7f5f6d967e21 · outbound

This paper cites Patchwork++: Fast and robust ground segmentation solving partial under-segmentation using 3d point cloud,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Patchwork++: Fast and robust ground segmentation solving partial under-segmentation using 3d point cloud,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.362772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.858979Z digest=sha256:c86df16f13ae5024ee14b8346d19c09d01e0897b58afd4bd2193391d24b032a0

Observation ef31a0f3-933b-443c-8e1d-c271e5f3973a · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning BridgeData V2: A Dataset for Robot Learning at Scale

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.869616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.869616Z digest=sha256:ccb209b5856dc479b6b603069c6bfea7e8c3720f6f77e8ccd3f43811f20fa8af

Observation 5f592ef7-a54d-4ce8-9b5a-dab006be4acf · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.853042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.853042Z digest=sha256:2ecfbc8cc07c26ba23fb09224fb37bc466f80599c5983aa47216bae27c7f3480

Observation bc592cf2-c519-4d80-89cd-abfd98a46259 · outbound

This paper cites Patchwork++: Fast and Robust Ground Segmentation Solving Partial Under-Segmentation Using 3D Point Cloud.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Patchwork++: Fast and Robust Ground Segmentation Solving Partial Under-Segmentation Using 3D Point Cloud

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:21:24.932684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:21:24.864173Z digest=sha256:5b9ddf94b0748d68203746da0162c9da2a1e8238ebaf3a28d7965df4a97a27bc

Pith citing papers

No inbound Pith citation observations are available.