Pith. sign in

Paper Citation Record · LEDGER

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2506.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19498 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:09.155546Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7dcb54d-4c62-4aef-8593-774b4343b1d4 · outbound

This paper cites GPT-4 Technical Report.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.006778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.006778Z digest=sha256:190317ac75fd4d1154fb05a6317648fc667b652675362e58388fc18b60a5b940

Observation 21aac602-01f4-46c6-9983-57593bd09a5e · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.047648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.047648Z digest=sha256:c6caaa5a5a36e31e95e9a7e21847284ae58c13fea4d733866fb7e6e19a2cc9f1

Observation a8fc755f-0a12-4267-98cd-080e70e43da0 · outbound

This paper cites Copa: General robotic manipulation through spatial constraints of parts with foundation models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Copa: General robotic manipulation through spatial constraints of parts with foundation models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.746596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:03.087048Z digest=sha256:400e652c8c30a60c0468723100d547c9fb7da6539cec79fb0566f4aaf58dbe1d

Observation 76538e9f-30b3-4326-9c65-c30a45354c5b · outbound

This paper cites OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.142366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.142366Z digest=sha256:6b00915ce6360c4679a2d15793a6ee179bed2ddcbdebcca58882942e9d44904d

Observation 3cf9391e-3c58-41cf-aaeb-7e7b2b496281 · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.239306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.239306Z digest=sha256:9cb38ad439d2169b827c814fd19129baca0ccae620eedb527d87084e9525908b

Observation b8cd82c8-14e0-4c5a-9b39-57dd03ce7e62 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.288372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.288372Z digest=sha256:1d820ae295fd455a325d9288a90cb4b19d698eb11a4b186dffc61da415ec54e5

Observation 976c396b-5b79-49c7-92ee-e309c1a1840f · outbound

This paper cites Guiding Long-Horizon Task and Motion Planning with Vision Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Guiding Long-Horizon Task and Motion Planning with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.390138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.390138Z digest=sha256:fce695a75e31f18f8b44d8036eb1a41bceaa17b6f7ab208e90a7495da2a48621

Observation 2760d490-e208-4e5d-ab2d-3edd5e53b8ba · outbound

This paper cites Open-world task and mo- tion planning via vision-language model inferred constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Open-world task and mo- tion planning via vision-language model inferred constraints

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.448540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.448540Z digest=sha256:632efceefdcc3b59d5292f3b38866a44e837e3e29e19a031f8d69d42de7dc861

Observation 11c21ccd-35e1-438d-bb10-5eca6d4f2576 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Physically grounded vision-language models for robotic manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.567384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.567384Z digest=sha256:6db5f63574a5e7e2d2a840b1c0b23cdd07e2efa959c7e25a6c398f43d7eed1a1

Observation dfc156ef-59c4-42dc-bcb1-abd204d73f86 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.613132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.613132Z digest=sha256:4c4dc9e24de27fbd6ac9e97e053c203392bc42f8cd619403f2510d2a9eb45494

Observation 729f7f9d-31dd-4201-8883-8523075fde40 · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.665705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.665705Z digest=sha256:7677a0879dda1cbe02e2ad6fb29c0c0bf2f885e64e7501135713a9bbd5f97d67

Observation 853f45df-a4e2-43cf-b81b-ee8f622a8357 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.746754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.746754Z digest=sha256:1580b577d78bb7d7b2a18fae316e68e566edfa6e8ba0cc1925eec4b2bccbaf99

Observation a4dba97b-2dcf-4f91-b96b-a501ea8e86d5 · outbound

This paper cites KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.801070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.801070Z digest=sha256:8744663084d48445c851dec203c674adc2bc4537cea7bd48e5812ba256699cf7

Observation 580133e2-e101-42c7-85c3-fa14c6b5f52e · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.885154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.885154Z digest=sha256:1d526017e302949251d4f0842fb31608cc4ebec882e0d42b79c31b3cffa2a5fd

Observation 995d07c8-f01e-4779-be89-6c9c304d5a5a · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.958565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.958565Z digest=sha256:dc9c5bd81492fe2f78de2a1ab96451ea8cd829ae3b0be5e0678e3775b9a33b49

Observation b1375379-f124-451c-aa10-bf2537dc0f4e · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.998541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.998541Z digest=sha256:27b0488604121438e3c685760316ae6dac7fcacbdd6ec8697daa9ce45d88c000

Observation 9dafe183-77f9-43a7-b1fb-f0083b08dbf6 · outbound

This paper cites Learning to interpret natural language commands through human-robot dialog.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Learning to interpret natural language commands through human-robot dialog

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.730699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:04.066584Z digest=sha256:36d9f832b1fa98313f49fb76fe8329fc53d2d33a0b93f00835c5d679e2807bbe

Observation aab21090-a360-4a35-9eac-4579e90a30e1 · outbound

This paper cites Grounding verbs of motion in natural language commands to robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding verbs of motion in natural language commands to robots

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.719822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:04.142753Z digest=sha256:d5625798b352b070cbbbb6516c2ebfe514dd584ee410553fbaf5e2875189f9eb

Observation 386300ee-6e29-40fe-9074-074470b7b85f · outbound

This paper cites Toward understanding natural language directions.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward understanding natural language directions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.709353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:04.190234Z digest=sha256:076dc557e85ca560a1823d430f1c73e082542d50ff95c0a56769f3aff9487dd5

Observation c3121184-f8a4-4d0c-b472-90cd1def4285 · outbound

This paper cites Understanding natural language commands for robotic navigation and mobile manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Understanding natural language commands for robotic navigation and mobile manipulation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.699711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:04.271113Z digest=sha256:ab5c03aed39552729cfdc8b6550f79ba9de62c138ff6d0c0cad3a3392a340203

Observation ef9b0120-ba87-41fe-8042-ff94900e6787 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.326924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.326924Z digest=sha256:5d8297af394fc9dd71dfcf72b187d0c584288fa80ddd89acb383c0ce62579840

Observation 0f291216-b1a5-457e-a4a0-125447ba092c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.421149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.421149Z digest=sha256:20a785d0a0ec875acba420b3792c1fb6dccbf283438dce5efb4d484bec4a16bf

Observation eba4b8fc-77ba-492f-91d1-011bb97464b5 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.499840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.499840Z digest=sha256:d58ebb693a0567f178177f240779026045912686e8d3cbad76da73139fc8edca

Observation 863da663-d828-452c-a62a-6c80cc8bc02e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.566817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.566817Z digest=sha256:4183173f6fdb7d0109c13c31c4a29ce99807942053ec51b36177f0ddc5b9af6b

Observation 683d88e5-4fa9-4883-b4cf-83956b5d0b60 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.641939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.641939Z digest=sha256:11da365cb7f38335f591b07d05573181e464bca0a55572d9fb2a7c2b5e341f02

Observation 59868033-286d-4dd9-8f01-30a6cca9c25a · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.680233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.680233Z digest=sha256:62c0f50f61a730174659b8de2538d45a838932230604d34936195a905e7bde47

Observation 4fccaf42-5f7c-4145-b561-acea518e2579 · outbound

This paper cites ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.770431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.770431Z digest=sha256:60e6bb64f706c9d1cf4ddebc45bc1d44c43e7dbbbf97612e9c5bdce54b065c86

Observation 173e3d0b-562e-4c94-860b-56bd7c878c87 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.843906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.843906Z digest=sha256:f44d150add08415ca1082fd74ff091c6450b027062bbf91a88af8ea81ad725e0

Observation bac650b5-f260-44c4-a0b8-30472031daa4 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.884890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.884890Z digest=sha256:c367a5742445fb885cf9badd4e2661d9e828b0e3e0f11ae3333465ccda7950ef

Observation 1bb37cf4-a892-44fe-9a4d-6c3904cccbdd · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.988783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.988783Z digest=sha256:0c13810b819c330c4f7bf4d79f95f20da8a0a7a79ea420eecb69b74b809f8ec3

Observation e9f30bba-37a7-4640-9ca3-92b965550e7e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Octo: An Open-Source Generalist Robot Policy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.061262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.061262Z digest=sha256:d5dcf826591c40cc86a481e44d33cbcb7ac370f1749c2ac37c9c7270c42267c3

Observation 74e98e85-7c9f-4150-862e-1c6478143aa8 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.111915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.111915Z digest=sha256:90257feb0dc0a78e78964fc6f52d98091c41855f5a3475362e836bdc7326bf53

Observation 6d3f8af4-8d51-4a1a-a2ee-f9d3d15fa295 · outbound

This paper cites RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.207574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.207574Z digest=sha256:42fcef49642e32acdfb438575589407b0ee03faf318ecea116642e50a09ae09a

Observation 0019b4a4-5b4e-4be1-a351-23cdb4da1fd4 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.276267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.276267Z digest=sha256:04864b61993acb112b1e17ccd2a5d24455921cbb912658e9761b59dfb944d6b1

Observation 431fd637-b19d-4a66-8bf9-2668e5200083 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.316266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.316266Z digest=sha256:42c16ab8996e51bed233d52ba1e88d41b58b0210263b2f225d8df91f57ec4ec9

Observation 42b8a48b-fb42-4115-9dd3-18e67ce0d8f7 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.418519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.418519Z digest=sha256:c7f43d1e026a48ef589e455d9da9ab1ad64042dab1a987433f439442258df192

Observation 5a081d2e-86fc-45f3-801d-ec30236c942d · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.478000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.478000Z digest=sha256:d3e09ef4dd13e22e6bc28e36bebfe2debde02d41d2e2332a28ace6e5b8410520

Observation 5053b701-0997-4d04-bd89-610c97133537 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.527284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.527284Z digest=sha256:334bb18e5387e9358f3e7bd24ae7a0cca33a451c83aefb971c52dbc371dd645c

Observation 792276d5-a035-444e-b92e-4bacef746649 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.638223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.638223Z digest=sha256:43c5b59e16fa7cb53dfb26284701d05e48cb6404e2035fb97bbfbb4abdeb4904

Observation 9e993ad7-d19a-4958-a9e1-6f1068b6278c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.690842Z digest=sha256:35f21ae66096fd666a9f5aa3ce76bafaa06ad500b6b106edb163a8a4455d7ad8

Observation 5ebc3e1a-ec6a-40e4-8276-dbf108045ac7 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.781091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.781091Z digest=sha256:6b2461bd6f61c56f23e8edcf62e435a85f695e5cd6080e4ec0bbe14171df46b1

Observation c8e34603-f4c2-4037-8cca-a920b2473925 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.850037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.850037Z digest=sha256:167eb13b1ae8bf4b5e24e77c95c75ba598ba4cfc3073ffa7cee7260b3ee0e9d1

Observation 948a6dba-6d7e-4e96-a0c6-8a10ee1607c8 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.893446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.893446Z digest=sha256:d1d81166f95f2b4bd2f4511b9b53e2c7282b0792bc7ab1360ca6da4828f9d772

Observation e3447b32-c4e8-4a2d-b955-4ed8aa480917 · outbound

This paper cites Code as policies: Language model programs for embodied control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Code as policies: Language model programs for embodied control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.953173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.953173Z digest=sha256:54d5cf179d830fe74eb4a97efaebf8497e32e70243c35ea2fa41cdbc888237da

Observation e5024913-6ebf-4446-a471-4ceca1687d6b · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Progprompt: Generating situated robot task plans using large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.677317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:06.080619Z digest=sha256:e76128d5f3c0f1118bb6e214a49e632d97749920a35029527dc8abcee9c63b26

Observation 782b5a8f-d56f-4d43-85d6-09706f9e0715 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Chatgpt for robotics: Design principles and model abilities

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.667052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:06.222417Z digest=sha256:f2f84c3523092e85aa84c3bc7070c8420df451f7174a303cbd75d32d2197ce34

Observation 549d0afc-d638-467f-9554-76eafc693ad5 · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.341930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.341930Z digest=sha256:bfc3411d3c170f3ed2c1c67a43ed0d977e8cc9e3b925037d87827ca328aa1a79

Observation 33ee52c1-a2c9-4f89-9348-2d41e098b838 · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundation models defining a new era in vision: a survey and outlook

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.449159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.449159Z digest=sha256:f4cb8ae3fe11937e1c6d9df78209356e92b13936cf41f22a7cb87816b3082dfc

Observation 4f3a9ea5-f4d4-4d98-b2e7-3ce1934b8957 · outbound

This paper cites Yolov10: Real-time end-to-end object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yolov10: Real-time end-to-end object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.513413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:06.591235Z digest=sha256:0ba88bb5603f9187d10149fe66d057f7318b7a2e719e68b0d8217af3242675ae

Observation 762187bf-5d7a-4126-8efa-83b428afc9f3 · outbound

This paper cites YOLOv12: Attention-Centric Real-Time Object Detectors.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models YOLOv12: Attention-Centric Real-Time Object Detectors

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.730254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.730254Z digest=sha256:0cf803806db039ec4857cdd5ecfde0dede416f03cfbd28ac128dcd28f9f46888

Observation 10cdafa9-2d9b-4743-913f-e11df4e5d3cf · outbound

This paper cites Yoloe: Real-time seeing anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yoloe: Real-time seeing anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.841797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.841797Z digest=sha256:093ed0b7a0d0d0aa14066ac751c4b745f67bc8451c9b361463291b0b37d68e03

Observation 06330ba7-25e0-43ef-9875-729e32e584f9 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.986608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.986608Z digest=sha256:bae467a14112d980a82d9c21071e0ceeb10b0e8e25d80a5a5fd93abcf1ea43e7

Observation cb635328-fa44-4d6b-81b1-06ef8efcb96f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SAM 2: Segment Anything in Images and Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.109157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.109157Z digest=sha256:932cd0d6dced3578d985327cbeb6a8571deb8484e46bc0afeaaa57dee9a0d1a3

Observation 8b365f20-6304-4397-ab46-0b0f8a3e1c87 · outbound

This paper cites Fast Segment Anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Fast Segment Anything

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.220820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.220820Z digest=sha256:b3b77fdfd98ca108bbc9bcc0d8978cc1447764b13acaabdd580b02558f9b9463

Observation e99c5c4b-e05b-42e4-aebb-d10330b70b74 · outbound

This paper cites Segment everything everywhere all at once.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Segment everything everywhere all at once

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.344271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.344271Z digest=sha256:670d96513d6c5c4b569b00e2522ab17b8f07f455ea5fc2d033dfe8f1021ffff4

Observation 96c81ed4-59b5-4a9c-a761-ad09e2dac5ba · outbound

This paper cites kpam: Keypoint affordances for category-level robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models kpam: Keypoint affordances for category-level robotic manipulation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.351801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:07.470139Z digest=sha256:db324c19b98aa18e0f64f05f9e9d622531476187340b6abb19ad348c1e8ba807

Observation c2077bc1-821e-4c72-8444-ed19d164cdc9 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Any-point Trajectory Modeling for Policy Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.568520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.568520Z digest=sha256:3763271157a1ef24a286a39b137d80b4b5153704c0006a813d69f28d08ac7549

Observation 8da321ad-6616-4dbe-a41e-20a59c43a8cc · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.661869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.661869Z digest=sha256:e6570fe1b0b4b8e52f6996d4b0be3f0e016809447e6e01dd541949e66f2ce964

Observation 8b477562-60e4-4212-88fb-6c53e5443b26 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.113415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:07.861016Z digest=sha256:e94889f45ad82955544a52387ad3429547ebd6a113ffd428ff768f2f96d47ddb

Observation ee1f49c4-f2d0-423e-a5db-1332efd13cba · outbound

This paper cites Sam-6d: Segment anything model meets zero-shot 6d object pose estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sam-6d: Segment anything model meets zero-shot 6d object pose estimation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.842969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:07.911112Z digest=sha256:9052550d3a2b0b9bd128fa72e9b231ab6d4f72c5d7d9bd21dfb53f6a9b7b1d7d

Observation a807a592-c2bf-4450-a059-72ad9a73c895 · outbound

This paper cites Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.687810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:07.937920Z digest=sha256:aa17f17dd62688eb83bb6f3e3b984ef7f12ccc3b4f3299f16fd40add44758db8

Observation b6cd504a-d784-4ee8-b767-f44c04bce151 · outbound

This paper cites Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.485692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.021674Z digest=sha256:51bb9a09a65e869199d09f90bd0a751ef13a4651199294554ea4beafb99f2120

Observation b9fca9de-91cb-4823-9fcd-ca78b24839be · outbound

This paper cites Onepose: One-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose: One-shot object pose estimation without cad models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.329092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.096588Z digest=sha256:fc7adac8038c1988b254350577526eab6f535b9c0fe3d204934e82159c9314bf

Observation fa9de27d-3787-4f29-a4f2-df1ca297a8f3 · outbound

This paper cites Onepose++: Keypoint-free one-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose++: Keypoint-free one-shot object pose estimation without cad models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.119342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.158463Z digest=sha256:e28b14668ee4270dc6cc7d294a9d1dcf887e46aa262735bcf594cf47bf488ef7

Observation 4b407ffb-5d6c-4cca-973e-98c757567aaa · outbound

This paper cites GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.497856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.250575Z digest=sha256:e341f15da53718c3bea80dd79eaa62a046a4b3f01d54a63b3828c485bd56baf7

Observation 746ef3ae-d489-4e7a-ba7a-ef9107518ad2 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.280921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.280921Z digest=sha256:7bfd46940ecde497cee7d66dce462fa087c900fea8adaa6409ed839e6bb884eb

Observation f33928b8-72ce-43a3-a0f8-3f30a67c1475 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.358863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.358863Z digest=sha256:8e9ae2f1962f1b4a3ced017c31238d74d56d1b91e998641f9a02a0e3ed0aa25f

Observation 6b33626f-3e6d-49ac-a012-b6b4d3b5ab12 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.430220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.430220Z digest=sha256:e9f1f99294a9ca26c2d5455afecd5d1b7ea0ed00c0af12edc9a824421fe526cc

Observation c570fe5d-f911-48a9-acee-dc952887e705 · outbound

This paper cites TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.483436Z digest=sha256:17f33133ac52ac5f5a28fa93ea8322fa66ebbd0c0cbadacb44b21a05cca41b89

Observation c5c84f78-1d05-4bdc-b997-dee6732cd87f · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.982349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.570568Z digest=sha256:9e5398867b7a16ef6cbcb135c82c1941229cda61fe9c744ad0718d6d2e33f102

Observation 52372e6f-62d9-47c3-ba65-9551ac4a6178 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.608592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.608592Z digest=sha256:81c09a099884d78f6291e0c8830e879b99f5a43daebcf4caa17eba98b3cc581f

Observation 3ce2b293-2739-4dba-8e62-d862d532f07d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.686757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.686757Z digest=sha256:4dec44b83a3f5e842c4603c9a813b1a51fa4affe199cfb3f837d45afdec66791

Observation 0a6daf92-271e-4862-b982-7d1c9c3cc231 · outbound

This paper cites SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.750059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.750059Z digest=sha256:47fda439b769f47a503343b4e6ed165d8f9abd6a0b0c269839452ba5fb836d44

Observation 53ca0041-7873-444e-ac7f-7068092a9c50 · outbound

This paper cites 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.785173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.827515Z digest=sha256:210e112ad48caddf0852f87418d4d06c13e78fdc5e0efe21f1b453a27b4411c1

Observation 76c4b5a8-b5b9-4cfe-99ab-e89056c77a1d · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.914602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.914602Z digest=sha256:9f6dc50fb1cde2c6b37d2b0fcf42e70b706ddab538644dd4ecad071a2ed83963

Observation daadb4f5-6f63-488a-acb5-81076bc53dda · outbound

This paper cites Pointodyssey: A large-scale synthetic dataset for long-term point tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Pointodyssey: A large-scale synthetic dataset for long-term point tracking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.666620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:08.982750Z digest=sha256:fdc9f326c20c51b644701acb046ca2594548a8dc738a65b0c42a397ad3b38867

Observation 191d4ab0-115c-4010-960f-05535a61a749 · outbound

This paper cites Cotracker: It is better to track together.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Cotracker: It is better to track together

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:09.050283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:09.050283Z digest=sha256:25a1f3dde7ea3fc9f93c93be3150781b16e0e03e5cfa484c54f818d74fb0105e

Observation ae492222-1f9a-42d4-81cf-495c70edff99 · outbound

This paper cites Robotap: Tracking arbitrary points for few-shot visual imitation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Robotap: Tracking arbitrary points for few-shot visual imitation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.496351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:09.075532Z digest=sha256:9f65a43c7d71f5d39d1406e16920440c6399a19bc96215d2bb693a6ed1c4cc08

Observation 83a67dae-83d4-4528-88d5-e146ac0b2397 · outbound

This paper cites Recyclable.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Recyclable

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.345022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:12:09.155546Z digest=sha256:75c56db724c9c97c754aa7de82fc1fadfb1fe748b1c55d93b8f100a6e5f58aec

Pith citing papers

No inbound Pith citation observations are available.