Pith. sign in

Paper Citation Record · LEDGER

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

As of 18 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 0 inbound Pith citation observations for arXiv:2506.12374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12374 v2

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:29.897079Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 137 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45199b3e-0170-40a6-bbc6-d428e817f69f · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736, 2022.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.081678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.081678Z digest=sha256:1f4eb18f1d6d59edfc1e8334b5adef80bdb2c81619aa8358abd00bb4db7d32bc

Observation e90b2662-f6d1-43e4-b255-88be81a0910b · outbound

This paper cites Learning transferable visual models from natural language supervision.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Learning transferable visual models from natural language supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.150439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.150439Z digest=sha256:ff4a95a1ac0747c1c3685d97a4afd7254882617b75008d0fd9dd2eaf0d308009

Observation a76572f7-59b4-49f5-a18c-98f5025b2ad7 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scaling up visual and vision-language representation learning with noisy text supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.223630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.223630Z digest=sha256:3cc83acbdb9143b78f0ba5a382c0d9016bd4bf887bb0b00299a46f4dd9ec7477

Observation 4ba83d3d-98be-424e-8a83-be2de65f9adf · outbound

This paper cites Zero-shot text-to-image generation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero-shot text-to-image generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.300306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.300306Z digest=sha256:6a57840fcf220527b92162c5101679d82e5034e260b159565db765e17b9ba2d9

Observation 8edbc6a0-d9e7-4f3f-a35a-387f5a3a0de7 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.396676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.396676Z digest=sha256:fa6468fde3efdde1b5d56064200e99ce5ac85ca0241cbf0c41adf69d5f9f15f8

Observation 6518237b-bf54-4ede-81e0-f66736cc7e99 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.479391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.479391Z digest=sha256:3b5cffe0d7c5b6a9b1ed2abe366abe0c48382f9c67eb85383649481f1ca7aaa8

Observation b8a27dc8-e4ae-4833-92ba-aa3a621e239b · outbound

This paper cites GPT-4 Technical Report.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making GPT-4 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.576785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.576785Z digest=sha256:d5d773f5637d5620451eede5c2bb1c9e54ab80dc73ee6b505331b45bdc3677e9

Observation c8333fc0-7cb3-41f0-8bf5-79688f81c4bd · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.641202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.641202Z digest=sha256:cc32c2f13931165665dd7a20a802c977ed7985608e4387d2f62cb083fb843d0e

Observation cbf9ea69-12ff-4097-a862-f5a33c4881ad · outbound

This paper cites Open-vocabulary queryable scene representations for real world planning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Open-vocabulary queryable scene representations for real world planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.750801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.750801Z digest=sha256:3de5ead0cf0ffecb8ab05868999cb0e4c97fb66c1c2a63aa372f4a996e3e6300

Observation b638acb7-e6f4-41bc-bb7b-a8ce4b092d30 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.851964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.851964Z digest=sha256:bfc1ad8b191375ce8266594910acea66e0d6438f2bbc713ddb98e1972799c098

Observation 1fd7d921-6d8c-46a1-8028-a7e487bca854 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.947466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.947466Z digest=sha256:c0ca671d4de4b7bb6bbfd405e1b07c9e96323948e872a2e22f43610895a841a8

Observation 0d1561d3-4d7c-4800-a5ac-e1472a3bf328 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.062840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.062840Z digest=sha256:2a58b2a10febe592a94c692afa98c3f2f8ba4c8cd3c8f0e1ce83593172172407

Observation d401c5fd-abdf-45e9-9f76-541046427ecf · outbound

This paper cites Code as policies: Language model programs for embodied control.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Code as policies: Language model programs for embodied control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.188599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.188599Z digest=sha256:1ed671dcacb43d47fcb875757b2abb2b7ecf85a934e620b4b41dfce5c76fa2af

Observation 3ee45808-8a8d-4e6c-a6c0-1b2f11779223 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.264561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.264561Z digest=sha256:eeb563987bc57cc96fe62aecf44312d939837d7a2d61c53a4bdd5e8220c283fb

Observation f898e174-1805-4323-9a7f-b500b241d761 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.306305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.306305Z digest=sha256:aea0cbd7df22713eb26a4c74a072f6e17a3f91c12d3b3e5cf5477807163aab08

Observation f1494ef9-d559-40ea-9c75-588bc06f24d5 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.372509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.372509Z digest=sha256:fd6f7cadc12169a4b59f2379334707af7de5592c87906934a71ed52bfa3cd2f2

Observation 7699befe-e1bc-4ead-880f-0e1beb5288f5 · outbound

This paper cites RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.482172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.482172Z digest=sha256:5884b4c560d1a0ba4bbf719a54c07ce014f965365b6373c0e22fc982cbbe3507

Observation 7c29c10a-ade1-4c72-91e6-45a7d04c8a88 · outbound

This paper cites RoboGround: Robotic Manipulation with Grounded Vision-Language Priors.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:31.256067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:26.612635Z digest=sha256:b8e28272a31551711471f3bf1bf6b0656229d679aa2210480fffa18719078954

Observation ba1a1b09-4043-4a96-8a6a-f2fe481d4284 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large lan- guage model as an agent.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Llm-grounder: Open-vocabulary 3d visual grounding with large lan- guage model as an agent

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.751072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.751072Z digest=sha256:cc35dee87f7a4242cac715ac64362c82ce027a25bcd72d237eae5e96a5d97aab

Observation 4a81f21e-71f4-46d6-a2cb-b4dd1f8683bb · outbound

This paper cites SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.863014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.863014Z digest=sha256:9740405dd3ad6066903136e06f193edfda9595c56cdc2d01d951eba37bf765f7

Observation 0a19d958-9ac3-4651-8fac-421e7d39b69b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.025972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.025972Z digest=sha256:7aa88113e3f476d175f975d41ca521b3c0f73f583bca2f1ed97fa1c503b63e29

Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.127004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.127004Z digest=sha256:fff4cce7af05adbedad08b78cdb43aaa5c54864816f1b1b32803bfd969bfaafa

Observation 89ed4db9-1ac1-465e-9632-203f668c3226 · outbound

This paper cites CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.167072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.167072Z digest=sha256:5f8247c10427d0485ba2400244ad5057a4751f4eb6c1b0319bee1726863dd911

Observation 6e20543d-3d2d-4381-8c77-b14dd9b70932 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.247850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.247850Z digest=sha256:921f1c49aa2d5c9bbb057cba9b931aebc1b93fa81fb439f38043719a4ccf649b

Observation af814dc3-26b1-45ef-83e8-940e705c8bfa · outbound

This paper cites Zero-shot visual reasoning by vision- language models: Benchmarking and analysis.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero-shot visual reasoning by vision- language models: Benchmarking and analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.372595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.372595Z digest=sha256:1e9b8b57c6fe999ca084ba25b36430152e3ac2094dbf54b95a4538c4bdcec440

Observation 7b350e45-f5a0-40d5-a504-12a5c2a2c2bc · outbound

This paper cites How to enable llm with 3d capacity? a survey of spatial reasoning in llm.arXiv preprint arXiv:2504.05786, 2025.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making How to enable llm with 3d capacity? a survey of spatial reasoning in llm.arXiv preprint arXiv:2504.05786, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.459312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.459312Z digest=sha256:a7ae74a08cfa2e6af77431f5824a9db115d88224f5c5a313b16fea89743f54b4

Observation d6bfe043-b776-4725-ac7b-a1fba9f5b982 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.528961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.528961Z digest=sha256:92254e6ed0e99aba7a3171caa6ffd7e5b84273b75efef322adcd1506a7e60953

Observation c8499945-4b1d-4d4d-89e3-eaac8adcd815 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.599419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.599419Z digest=sha256:03d8fc756d6b7877d71b44da4df7bae96939bc924c95022ae7fe2bbfcf154765

Observation 06b3bc4a-023e-42d1-9676-72c568e0c030 · outbound

This paper cites Agent3d-zero: An agent for zero-shot 3d understanding.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Agent3d-zero: An agent for zero-shot 3d understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.720163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.720163Z digest=sha256:40a372d90de69b933ea08ca951111f34f22d0bf925ec29f81a0955449eae00f6

Observation c990ba0e-9bf6-41c4-9c72-612f6e2a9d5f · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Shapellm: Universal 3d object understanding for embodied interaction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.842050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.842050Z digest=sha256:7cfa26ee12d6083ef575ce3aa8a3b740aedd7ad334efd782b3fe1d77b52a1cb0

Observation d451d68d-d5ea-49c2-ac94-441f12a5004d · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.005755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.005755Z digest=sha256:ffece5d6576383d824ca3249206b12a2bf56fd33a4c67cff3e2eb16f3c3441cb

Observation e5a6f8f5-13d6-4362-af1f-7a0192560783 · outbound

This paper cites INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.102223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.102223Z digest=sha256:9738d4d03ab7e70f4951200600427ec2a40d0679884312afa87ea168c31a82e1

Observation 7764fc4b-8cf7-4717-926f-f3a11a814c8d · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.210537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.210537Z digest=sha256:a84a3825c033ce94f2b64bb709b1b84e8b43da6ac1531a3d7cd49d94718408d1

Observation 3c265a4b-d4ec-410e-bbe8-c6540ef8450c · outbound

This paper cites Manipvqa: Injecting robotic affordance and physically grounded information into multi-modal large language models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Manipvqa: Injecting robotic affordance and physically grounded information into multi-modal large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.331337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.331337Z digest=sha256:efe8418432896cd6373ce97ab61b7f33dee45069b220301cac74a6fbd52137fa

Observation 1ceeb23e-fa10-453e-b8d1-9f5377370db9 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Robovqa: Multimodal long-horizon reasoning for robotics

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.492183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.492183Z digest=sha256:f19eafd645aafc02082eb83565ee5ba0be43688e3d4440ba32a518435c3e12d2

Observation b4f339b0-0fb5-4577-9c16-10da8048e90a · outbound

This paper cites MQA: Answering the Question via Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making MQA: Answering the Question via Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.618914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.618914Z digest=sha256:2fbf6bfc84f15c33a7965513949e23d11e2bbcee8bcfaf91db76da44eddc0ffb

Observation d2251920-b18c-4576-b55b-d1fe88529cf5 · outbound

This paper cites Robotvqa—a scene-graph-and deep-learning-based visual question answering system for robot manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Robotvqa—a scene-graph-and deep-learning-based visual question answering system for robot manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.668841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.668841Z digest=sha256:41269fed6112516486923e191f6bab83fe0c6040a5f6bc06de0f8ee1b8d907e7

Observation f7b5ad7a-a19a-4970-88a9-af407adeb473 · outbound

This paper cites A visual questioning answering approach to enhance robot localization in indoor environments.Frontiers in Neurorobotics, 17:1290584, 2023.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A visual questioning answering approach to enhance robot localization in indoor environments.Frontiers in Neurorobotics, 17:1290584, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.748826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.748826Z digest=sha256:9b340f53f9e6a010677c3d221ce86c4e57f730e82ca87fc4b3ada41fd9d67f5a

Observation 007edbd3-b00c-48af-8651-72a8d2f5f8a5 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.855409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.855409Z digest=sha256:e25a1fba99f34d36dc37c36ea4add0b88124b88e3d49f299d83d8d60d5a88905

Observation 02853fee-6955-4fdc-a7a1-b06257ea453f · outbound

This paper cites Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.983421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.983421Z digest=sha256:354a8509829b8c637a3267d01ac9dd48e4bd4b0dcab7020b2063ae3f59ba4a61

Observation 883cd5eb-1dda-4aad-99f3-c20eb39e9c0c · outbound

This paper cites Open-world task and motion planning via vision-language model inferred constraints.arXiv preprint arXiv:2411.08253, 2024.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Open-world task and motion planning via vision-language model inferred constraints.arXiv preprint arXiv:2411.08253, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.049705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.049705Z digest=sha256:8f8ba0766f121e2f8038d4fd0484eabfac016cd6ce96d393e4a64187d630b3c4

Observation 5343a5cc-08c4-407a-a1ba-15e190d79ea4 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.064263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.064263Z digest=sha256:063eced36d98f89c27737b2a9565cf28558f69c01e947d04f25421c08c82b9b5

Observation 4e079adf-50b6-4d9a-8d6b-0816a3c10146 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RT-1: Robotics Transformer for Real-World Control at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.147844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.147844Z digest=sha256:f14d4a0eeb87498fb5af01c09cddba8529e8d82d5d93d79f67e0c25e8d1bd2d0

Observation 38cca183-d0f7-48ce-8f4d-94170850d9d4 · outbound

This paper cites RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.206479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.206479Z digest=sha256:003141d3017e9d8d62e4f328514cba11016bb0ed11f83e33a6b3c9f15139d533

Observation 6508197f-7505-44fb-ab49-954eae7e8b86 · outbound

This paper cites Scaling proprioceptive-visual learn- ing with heterogeneous pre-trained transformers.Advances in Neural Information Processing Systems, 37:124420–124450, 2024.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scaling proprioceptive-visual learn- ing with heterogeneous pre-trained transformers.Advances in Neural Information Processing Systems, 37:124420–124450, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.283597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.283597Z digest=sha256:7673423cb0df3916b6fea755c5620d9ad6b9e21daddd363431c143ce033f049e

Observation cda072aa-ae8e-446f-91f5-afd6562b6005 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.343756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.343756Z digest=sha256:5164290b5ae8dc9d1925142c4352c61bc06016c119833043e6523ca725534b03

Observation 56a818c4-2a61-47a4-9e76-3aa21b770964 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.426099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.426099Z digest=sha256:44aed1e6b9ccc25a45a8635d252f8e4e9d28ee872da671b1c642c3de8c72a28e

Observation ba9b3811-0d51-4698-b11e-46816b7ede77 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making OpenVLA: An Open-Source Vision-Language-Action Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.508963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.508963Z digest=sha256:56583a585d9d029473af51c3dc107ecc27d8643568a5caced123e4207f893c21

Observation 960a02d4-e7c6-4bec-8684-ef955625343a · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.572917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.572917Z digest=sha256:80607441658935157833094eb9619ea5b4d814c5faf2349c4bf6956a5035895b

Observation 3c9762e3-034f-4bc4-9d83-b3e51d650d2b · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.614110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.614110Z digest=sha256:ecdb2a577cc968838a558c553500ef94bdfb0ba2933999edfe2e20e6c393520e

Observation 0be1e207-5932-42b9-b829-029e1581e81e · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.690756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.690756Z digest=sha256:4870ea881ceddab662e0138fd90d33452727fc96b8ae0a859dd5f0969905a037

Observation 9a508b3a-2b41-4f7d-8f98-022b07172a94 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.709441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.709441Z digest=sha256:0df0cbb2f3518ac725ce1d95eca74f1ad727de9a9e8d2125af8d220530870b5a

Observation 44cfb571-0f28-4941-8826-e53bf5b252d0 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.714057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.714057Z digest=sha256:f20563733138504553e56629c9d6bf4a299a0e5dfd639745befe30fb6ef27fa7

Observation d63ef4ae-4dc7-4a58-9977-ef06d5db41d7 · outbound

This paper cites Skillman—a skill-based robotic manipulation framework based on perception and reasoning.Robotics and Autonomous Systems, 134:103653, 2020.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Skillman—a skill-based robotic manipulation framework based on perception and reasoning.Robotics and Autonomous Systems, 134:103653, 2020

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.718582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.718582Z digest=sha256:4094e5b7313bda6eeb78c0f5c5b457836e129b7556c934ff342c27401e3caa67

Observation 680fdbc5-9dcc-42c6-94f7-1772805f9740 · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.723051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.723051Z digest=sha256:5e1bb8536a8e77f13d97b522e5cb28b9b65d162928e3ac35e01e6919463a1652

Observation 3218ec8b-1e33-4470-9fbc-28ae5e08c84a · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.arXiv preprint arXiv:2502.13143, 2025.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.arXiv preprint arXiv:2502.13143, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.727266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.727266Z digest=sha256:3f44296028cd531b8cd06a96f073ebba3312d43b8b374edace63f1ed50c3e2fe

Observation 0dc6f32f-0d3d-46c8-a402-d3623c0da8b5 · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.731232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.731232Z digest=sha256:837da375b30cc419c8376f014b50f4a307c776b4af81afb7153b3ad6437ae5b9

Observation 3423f687-670e-46b2-9fb0-066bba60553d · outbound

This paper cites KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.734890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.734890Z digest=sha256:25a2774e2a2cb0cd75bc6ee1e336b19b5d7c33cb202fa190d4a68c539e43d696

Observation 2e2bcb92-9f27-4b04-b0d1-e5669a48fce2 · outbound

This paper cites RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.738813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.738813Z digest=sha256:2f92e4011f3872c28bd9ea3b41dc4bffc189a386e5ba7f5f16872c71efacf357

Observation da8c1c21-aa6f-476c-9a6d-610b252029e3 · outbound

This paper cites RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation Learning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.743022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.743022Z digest=sha256:264dfb1a3db175d096bfbc78e414416f3700aa18a06e65984b98b019d5e37536

Observation f3ef6282-fc0d-4948-b435-6dd960cd8532 · outbound

This paper cites Discovery and Deployment of Emergent Robot Swarm Behaviors via Representation Learning and Real2Sim2Real Transfer.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Discovery and Deployment of Emergent Robot Swarm Behaviors via Representation Learning and Real2Sim2Real Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.746981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.746981Z digest=sha256:c2aade0c1d5a833c3edf8556967e7087716579e8e0427be53854a4cd328b667f

Observation ff4c716f-cc6d-4892-9889-7aff6367eb9d · outbound

This paper cites Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.750649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.750649Z digest=sha256:26956de99e3a9698f80e9001379f53feede0ac6bb5c6e356b9d0367467240176

Observation 202612d6-a7d4-44a6-bd43-0bc2841f8da3 · outbound

This paper cites Rl-vigen: A reinforcement learning benchmark for visual generalization.Advances in Neural Information Processing Systems, 36:6720–6747, 2023.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Rl-vigen: A reinforcement learning benchmark for visual generalization.Advances in Neural Information Processing Systems, 36:6720–6747, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.754394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.754394Z digest=sha256:c929c0d0dc70c91c6ccab3c4594bddd703530bc8265afdce9cb0e17773b0fe75

Observation 81530520-73f9-46c8-91f4-6b365c713e3a · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.757961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.757961Z digest=sha256:e6e94abd96e8c32f61bbe6881060d57ea17356484cf8f35851ddeeb754f38a84

Observation 496335a6-57fe-43fc-b268-e855daf8b95b · outbound

This paper cites Efficient real2sim2real of continuum robots using deep reinforcement learning with koopman operator.IEEE Transactions on Industrial Electronics, 2025.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Efficient real2sim2real of continuum robots using deep reinforcement learning with koopman operator.IEEE Transactions on Industrial Electronics, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.761706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.761706Z digest=sha256:f8f329885b7c535d9f38092ede354bcaee0b93989e14137a2a010d22ebe6c033

Observation ffbbd005-5e6d-4c1a-a83c-00dca03b16e1 · outbound

This paper cites Real-time per- ception meets reactive motion generation.IEEE Robotics and Automation Letters, 3(3):1864– 1871, 2018.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Real-time per- ception meets reactive motion generation.IEEE Robotics and Automation Letters, 3(3):1864– 1871, 2018

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.765399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.765399Z digest=sha256:61d846addc2e6d9f849ceef1cb3357efcec14df9268fe3984e469e491290619b

Observation f3f35635-23cf-4e22-9e02-b007767edf48 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.768816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.768816Z digest=sha256:7a345083b15422c6ecb9c1cf90393330dc23ec85901affa293d4c26ad4517262

Observation f574453c-c930-4e62-9c01-d6bda8e08740 · outbound

This paper cites One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.773481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.773481Z digest=sha256:662ef78f5fc948954450d860ddaa4c17d042f5dc4439049b5179bfa3e126f658

Observation adbc6daf-938c-4315-bfd0-5d33f89eb69c · outbound

This paper cites Sparp: Fast 3d object reconstruction and pose estimation from sparse views.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sparp: Fast 3d object reconstruction and pose estimation from sparse views

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.777217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.777217Z digest=sha256:0f85c96c1f695fb9b16037ec93758ade06db4b84b83a8cd34f69d7d2e4908f9f

Observation a35b8401-4e81-4ef2-bf10-f1589bf1bded · outbound

This paper cites Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.780816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.780816Z digest=sha256:900547dbbb198bba54d70e1e377b37d512deeae4c691ee30aa6db98c1add4f84

Observation dc0a08ef-866b-441d-8562-1e5ae190eb2a · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero-1-to-3: Zero-shot one image to 3d object

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.784567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.784567Z digest=sha256:da9c026c10cc70f317ad80d6d4da72fa053ecc406441de24879f47ad8b5f4269

Observation ef4dcf57-55e1-491d-864f-80c272ffa6e7 · outbound

This paper cites Get3d: A generative model of high quality 3d textured shapes learned from images.Advances In Neural Information Processing Systems, 35:31841–31854, 2022.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Get3d: A generative model of high quality 3d textured shapes learned from images.Advances In Neural Information Processing Systems, 35:31841–31854, 2022

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.788055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.788055Z digest=sha256:a6cde052953cb196a965288622901eff1112905145cc4559d2b5ce02d635a3f8

Observation 48bc8d56-a630-4be9-957c-79ecffa15e31 · outbound

This paper cites A-sdf: Learning disentangled signed distance functions for articulated shape represen- tation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A-sdf: Learning disentangled signed distance functions for articulated shape represen- tation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.791944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.791944Z digest=sha256:3bee89883d834ff2ce82239e8ca78a3acb3e257e0ed5fe3e591fcce90e9769d5

Observation 7580dbdc-6cdc-4ed1-9099-4bee1a27c0e8 · outbound

This paper cites Ditto: Building digital twins of articulated objects from interaction.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Ditto: Building digital twins of articulated objects from interaction

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.796122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.796122Z digest=sha256:abbdf37a138046d39c5fa2114f953a84697b5bb32096864bc3b0426ab83f7e63

Observation fa0071f9-0ff6-4537-83a7-1eb67104ae13 · outbound

This paper cites Structure from Action: Learning Interactions for Articulated Object 3D Structure Discovery.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Structure from Action: Learning Interactions for Articulated Object 3D Structure Discovery

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.799748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.799748Z digest=sha256:d9daa6a505f4ab4d33c514f84dbb40914777583540ad3cfcfe0c3df9c082852a

Observation c419f85c-0346-4210-87fb-e8a8ed36e284 · outbound

This paper cites URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.803567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.803567Z digest=sha256:0f0ac1ae1395f57137b66a8a73fa02e7114884f5ae1821c2fd3f887ada66ad52

Observation 20fce2ed-0716-448f-a67f-d57b6c645845 · outbound

This paper cites Paris: Part-level reconstruction and motion analysis for articulated objects.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Paris: Part-level reconstruction and motion analysis for articulated objects

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.807593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.807593Z digest=sha256:e31145f83d034fa4694bd137af80315e973878a26b2ec4b2642ecab3a27929db

Observation 495fe785-5f7c-426a-92ef-5734e6c41c7b · outbound

This paper cites Cage: controllable articulation generation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Cage: controllable articulation generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.811629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.811629Z digest=sha256:8fe1b6db6e9b55361a194953e8977d1de77ed4577ea423a5d9347bde10317dab

Observation 0fcd9752-97bc-425d-ad36-2a43b6db519d · outbound

This paper cites Sam-6d: Segment anything model meets zero-shot 6d object pose estimation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sam-6d: Segment anything model meets zero-shot 6d object pose estimation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.816117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.816117Z digest=sha256:d126b7168001af51651369cd0553ebee4ba605700276cfc02165024779232212

Observation 40a5fc99-2b26-40a1-876d-f1d74922802f · outbound

This paper cites Gigapose: Fast and robust novel object pose estimation via one correspondence.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Gigapose: Fast and robust novel object pose estimation via one correspondence

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.819976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.819976Z digest=sha256:9ebff31b15ff0d9da4635b81f81a02362d60f4a0d77475cc07ed45b15ff720fc

Observation c38baffb-291d-483c-bcbb-d7590ebed77e · outbound

This paper cites Any6D: Model-free 6D Pose Estimation of Novel Objects.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Any6D: Model-free 6D Pose Estimation of Novel Objects

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.823718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.823718Z digest=sha256:062177f30ee4668722c772ccadc4993acb07f391c110bd11f69fbac7a52b2384

Observation 3d8d1f3f-b4e2-4abb-9faa-061c4988e054 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.827851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.827851Z digest=sha256:67683fe992f17751436a48b0cf87c315e765ca4746acd18b5fd574c8e3ae9ec9

Observation ef06b616-a511-492c-a899-22c81d92154f · outbound

This paper cites Foundpose: Unseen object pose estimation with foundation features.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Foundpose: Unseen object pose estimation with foundation features

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.831798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.831798Z digest=sha256:7c10294d39ffce2d3b2d03c0be82fad1ad86bb46a2ddf0a04e221a8e5a6dc82f

Observation 175bfc2c-8433-43e3-80fc-0818e56a3ebf · outbound

This paper cites A comprehensive survey on point cloud registration.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A comprehensive survey on point cloud registration

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.836289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.836289Z digest=sha256:d07ab86f8c8e444c85f77419cd94292341dd75e1092b0dab423f2d5d3555a3f0

Observation 64a7666d-48b9-4c6c-a827-03d561bd4fd3 · outbound

This paper cites Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.840574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.840574Z digest=sha256:225d57f1486379993d60ed67237f8fcd9e3245a5c7c8cfed86514e174847e5f2

Observation 4b60822b-1233-45b1-a150-d639ee99a26c · outbound

This paper cites A tutorial review on point cloud registrations: principle, classification, comparison, and technology challenges.Mathematical Problems in Engineering, 2021(1):9953910, 2021.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A tutorial review on point cloud registrations: principle, classification, comparison, and technology challenges.Mathematical Problems in Engineering, 2021(1):9953910, 2021

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.844471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.844471Z digest=sha256:68ed65a054369599361d34ee5d143314943e17db1d9bd6f3221cdff0e5843876

Observation 6ad01dc7-36e0-4509-a66b-54ff429948b6 · outbound

This paper cites A comprehensive survey of visual slam algorithms.Robotics, 11(1):24, 2022.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A comprehensive survey of visual slam algorithms.Robotics, 11(1):24, 2022

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.848687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.848687Z digest=sha256:c1a1490765984b9bce049f34c9bae4a4236119a82ea19dd23fcb956b9c59c453

Observation 64a42648-e64e-4826-9f97-f59431350a19 · outbound

This paper cites How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.852481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.852481Z digest=sha256:284183ff08ec66db1d29591a4ee4153494b0f24f5c419b152e91e8dc07cd578f

Observation 15c6e1a5-f68c-4fa6-832e-18218550afdc · outbound

This paper cites A survey on active simultaneous localization and mapping: State of the art and new frontiers.IEEE Transactions on Robotics, 39(3):1686–1705, 2023.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A survey on active simultaneous localization and mapping: State of the art and new frontiers.IEEE Transactions on Robotics, 39(3):1686–1705, 2023

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.856398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.856398Z digest=sha256:09c531ff01468bf02864eb5e2d82ed3a9db2e173b33840bdf8f24e82d1dc7e25

Observation 533d08e4-5d2e-4e78-ad27-85348be21404 · outbound

This paper cites Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.859946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.859946Z digest=sha256:dbd39d6aed556023289859e955258fecdacae534f9a2a21bea98faa8c683f92e

Observation 5b6072d7-adf1-4bd5-a78c-54b04a71c4d0 · outbound

This paper cites PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.863672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.863672Z digest=sha256:311dd11b32b12b35519ee62281b95dd720227caf556e1dd8428f34bc2cc73051

Observation daec14fd-34fd-4737-8cad-94abc14cff19 · outbound

This paper cites Sim2real 2: Actively building explicit physics model for precise articulated object manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sim2real 2: Actively building explicit physics model for precise articulated object manipulation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.867354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.867354Z digest=sha256:38412fd66483019f6ba64a4e441945f6787b1969f94ee3e9a1c94ec4d13bc923

Observation de7bba1a-3309-4736-8a1a-7dda6b800c41 · outbound

This paper cites A real2sim2real method for robust object grasping with neural surface reconstruction.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A real2sim2real method for robust object grasping with neural surface reconstruction

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.870998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.870998Z digest=sha256:81dfb8d6c5707bc76d8372e330c2abe9528076fc3317ead646ada8d83e2030d1

Observation b42e82ff-3951-4d4f-ae39-dffb6af3f622 · outbound

This paper cites SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.874470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.874470Z digest=sha256:573f32981c933b325550a66fa1f792e56ef23dec80ba1a953cbfadf0a33a2f28

Observation f01ede5d-9fa8-446c-bdc7-6d8f221cebb1 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Pointllm: Empowering large language models to understand point clouds

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.878448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.878448Z digest=sha256:b1c655e53515d98e34535a49b3324abafceb48d7b2137f0d004bd2cbe5ed6c2f

Observation a181ef1d-f192-48db-ae2a-fa10fe1ae6ef · outbound

This paper cites Point-nerf: Point-based neural radiance fields.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Point-nerf: Point-based neural radiance fields

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.882105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.882105Z digest=sha256:c7f4d6b09f2ec4b4737df852f4e3ddb5139df1908bd68ee9f4943d9f493375e6

Observation df7441fc-37c6-4163-94d6-9fcbeddb9643 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Pointclip: Point cloud understanding by clip

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.885459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.885459Z digest=sha256:d6df1dd7c6564add524a37cdf99f1d6d138481541e5613a1b46795fcc70e2839

Observation 61384a05-d953-4a85-8b61-2d7e14236333 · outbound

This paper cites Text2nerf: Text-driven 3d scene generation with neural radiance fields.IEEE Transactions on Visualization and Computer Graphics, 30(12):7749–7762, 2024.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Text2nerf: Text-driven 3d scene generation with neural radiance fields.IEEE Transactions on Visualization and Computer Graphics, 30(12):7749–7762, 2024

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.889110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.889110Z digest=sha256:74d91e6a6fda867fbfa3f55e7288b6b667362aa94b5af394566f2cec793b5e8f

Observation 6086479a-47b4-4b7c-9399-15e7e233b747 · outbound

This paper cites Pointr: Diverse point cloud completion with geometry-aware transformers.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Pointr: Diverse point cloud completion with geometry-aware transformers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.893362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.893362Z digest=sha256:f01a2a22930b9218a319b4123f3010a0c35259b4409d710f53d18b63e2f9b363

Observation 613b9ae0-f12c-4977-9baa-8f78b495a719 · outbound

This paper cites V oxel set transformer: A set-to-set approach to 3d object detection from point clouds.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making V oxel set transformer: A set-to-set approach to 3d object detection from point clouds

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.897079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.897079Z digest=sha256:3023ee497e8273a9143dfe23e97127e7176dbf0ac9ee861342135e013dd705b7

Pith citing papers

No inbound Pith citation observations are available.