Pith. sign in

Paper Citation Record · LEDGER

Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2305.11176.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11176 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:44:50.810323Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:26:54.507995Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14ef7f5e-d878-4474-bbb2-24627362f663 · inbound

ConfusionPrompt: Practical Private Inference for Online Large Language Models cites this paper.

ConfusionPrompt: Practical Private Inference for Online Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:58:54.849006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:57:55.197897Z digest=sha256:97c83ff58a53a9e8366e8645f5f9f4cc3cc1de3c5d203b43a154d741bcdb6dd5

Observation 28ec5370-4b3d-4212-9945-02add761aeb4 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.588575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:f31eaba90d26119c8694c9b78ebf9d8507198c2623499ec47867e20fcebe4bef

Observation 89f0551f-a3db-4646-a5b3-761676c6817c · inbound

Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation cites this paper.

Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:44:50.810323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:44:50.810323Z digest=sha256:a2350ded2c2e43ba611938b95a74a8c369f7658754a34f09d5d8ba681c170045

Observation 4373253c-8b85-4c23-9d85-3f86929dab2c · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.363277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.363277Z digest=sha256:61866ddd86116579f1115281638b1251ef4e1e4e4fef606ca1a7c1267c14a688

Observation 89abcfe5-4980-4830-a70f-a54d6375978a · inbound

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models cites this paper.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.880512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.880512Z digest=sha256:d63f0a9aa4941e8d40bab699199e66ed6500341d78a373367ae44b4fc48167f7

Observation 8ca5bd2e-b098-457d-9d5d-a490089a89d2 · inbound

Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction cites this paper.

Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:23.742554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:03:23.742554Z digest=sha256:08b33d96d6030822eb5936e4f376b3d9d96a89a54d00fc72512a989ace9e66c3

Observation c0190b68-6726-4d51-925c-eb772a4a5cb2 · inbound

ACTLLM: Action Consistency Tuned Large Language Model cites this paper.

ACTLLM: Action Consistency Tuned Large Language Model Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:06.759535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:06.759535Z digest=sha256:4644d4a2579bf6090d32e3d601c8f63f52fb2f9eb8d7ce084942c10ca2e8e8e8

Observation 41837580-5943-489d-86cf-d0ebd4a8e2da · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.437833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:de9bc5a948106c203a5f067c71c9dbda2991de9c71e39e82af73f3ca8924d9ad

Observation f149ce8b-d424-4047-8791-14a0fd8573fe · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.122220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.122220Z digest=sha256:aa866999149d8a707f1b8b0895ae43df973d77ec211e4d10a069d84ed26df33d

Observation fc477058-e9f4-4f03-aa57-ac6113c7ec82 · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:28:42.012107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:a1923c3e83ef09fac02ebf522c356007abc5f2c19ada9113fb760c4292127c01

Observation ba2340f1-f0a0-4a76-8b2a-df76d4981e15 · inbound

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing cites this paper.

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:30.341166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:30.341166Z digest=sha256:06f563ebeaae2903c7b5dea17069900fbee647c680f313b6ce1d5b7cfe9e584e

Observation 50b1e20c-be23-4d8f-95b1-8a6def55c66b · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.886164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.886164Z digest=sha256:6e2abb5bf05e1818dbbcf6eeeebe5ca28d4568684220681e656d69f3b55322e9

Observation 439c4702-a860-4347-a2b3-4074bbf90a83 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.007179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:9b6b08620c2d7541964017e6f913bf26a2d2a07dc6da6e9e4929dc1bcf20df1a

Observation ed0b6296-7bb2-4a49-8dcb-3f090240dd80 · inbound

MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence cites this paper.

MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:10:34.249058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:05:47.657235Z digest=sha256:a41af057ca2493bb80c208c21b5e5cf235b4d0e10a645ee20502d711ca59c3ae

Observation ff20d6fc-18c7-4d48-bbdb-df59b914bd95 · inbound

ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination cites this paper.

ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:08:32.486743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:07:54.497999Z digest=sha256:e9cf57b73745c8329a2491b047249b0fba428c044f518cc012ef990e064f15b4

Observation 288dafc4-074a-4583-abe1-aaee24236c04 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:20.295381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:20.295381Z digest=sha256:ab38a8f35dd2a88de1ea91aca70bb8d8e9cc0e8583736bfb51453f6283e95b30

Observation 995f3c54-2c1d-45fa-a8e2-970b786bc6ab · inbound

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization cites this paper.

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:24.469130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:04:48.932945Z digest=sha256:134ccfe2116fd7851c0e351fd5d6ddb7e9865913f6b2d1ef17f36b7d40983c1f

Observation be0a79a6-6112-47a0-aa88-ecbeb602750f · inbound

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models cites this paper.

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T09:26:51.544337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T05:29:47.894226Z digest=sha256:a7e752fe3518a2f5a874df78a6944b6bbdc4b3fc123da0c82d313dac0da94461

Observation 9809110e-0ecf-49d2-a9b3-76dfeb8145ba · inbound

Efficient Skill Grounding via Code Refactoring with Small Language Models cites this paper.

Efficient Skill Grounding via Code Refactoring with Small Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:24.153749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T19:55:20.212198Z digest=sha256:110026a100b69411884798444051cc6e78f058c8ce5dbbd28b49d4d59b22491f

Observation 0e5fb998-6181-44f1-8c8e-c8d955e8c208 · inbound

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents cites this paper.

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-27T05:30:35.769413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T05:26:48.431739Z digest=sha256:b03b7b69ec5781ef9df9f98f392af091b54a16b5f799a7b75c7fc99fa1c6db38

Observation 34c3500e-747b-4446-a4f6-566ad15a71d2 · inbound

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents cites this paper.

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T11:44:03.545449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:44:03.545449Z digest=sha256:4c82f513253af3c8cf5b4b1d025df082e7fdf30b9a9de938bce4e5174a6f2ae8

Observation f29e2a24-7b6b-4337-81ef-652bb752d701 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 230

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.761394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:48ab64c187374df61b0e1bdb277e52de86b7f278eb13d9940a17689979e3a3c6

Observation 1645445c-a951-4492-a399-749d9d4b7ff7 · inbound

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents cites this paper.

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-10T07:26:54.509395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T07:18:57.823444Z digest=sha256:de19af74a8d3eb6884d4abf482ecc090840307347aa9bdc6bb860574628a4711

Observation a68f4efd-fefc-40af-aaf3-647509499690 · inbound

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents cites this paper.

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:56:56.620153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:56:56.620153Z digest=sha256:3fe1838dee1aee4fcbbe8c466086cbbe3c5893c5d534b0a54f245ab9857859dd

Observation c237d157-fea3-4b06-8ba0-d60383f9904d · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 237

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:d3cf55314bb178eb77ec00a3b1e95f1bf06a20fc2e24782882117008fa677388

Observation 4c58c8a4-f760-43e9-8d79-5fa90acfec7c · inbound

Self-Evolving Just-In-Time Memory for Proactive Embodied Safety cites this paper.

Self-Evolving Just-In-Time Memory for Proactive Embodied Safety Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:52.432986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:52.432986Z digest=sha256:2ea415e5e27716a2a0c23d6d16c9ac5a679bcf3f18bf2a3c2ecb7a3982894b00

Observation e3d210c8-b04d-4049-83ea-fe0b50d20bb5 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:33.996455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:33.996455Z digest=sha256:6aeb57a3530d0b6b1acee96544a058089062135297fd44897e27e5ebb1740f58

Observation 766e97bd-0a28-4853-bb37-341cae407934 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 220

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.330321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.330321Z digest=sha256:93c795bf851580adda8d6fd3e025e0ee278d0b46dc39db25fe9a91c7879df561

Observation 54530e77-b581-4687-a48e-ad8354d3d278 · inbound

ETA: A New Agentic Paradigm for Embodied Tasks cites this paper.

ETA: A New Agentic Paradigm for Embodied Tasks Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:44:44.230086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:44:44.230086Z digest=sha256:24225e3425a7c41afa374a9fd067aad7654daeb9917c15d849b6a15ad801af39