Pith. sign in

Paper Citation Record · LEDGER

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2506.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19498 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:09.155546Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7dcb54d-4c62-4aef-8593-774b4343b1d4 · outbound

This paper cites GPT-4 Technical Report.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.006778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.006778Z digest=sha256:27e9a4393c8d725c26669726c6fec0fb61b9f81c5b4326ba7035b784d136a8b8

Observation 21aac602-01f4-46c6-9983-57593bd09a5e · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.047648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.047648Z digest=sha256:5c56c59744c2ff81104f6a9e16a7a373d41d250c25c9374ab359488f5f644d40

Observation a8fc755f-0a12-4267-98cd-080e70e43da0 · outbound

This paper cites Copa: General robotic manipulation through spatial constraints of parts with foundation models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Copa: General robotic manipulation through spatial constraints of parts with foundation models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.746596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:03.087048Z digest=sha256:cb09c45cd9b764f37faf077b449263b0d8cc640bcd32c5048b07fde27bcb8559

Observation 76538e9f-30b3-4326-9c65-c30a45354c5b · outbound

This paper cites OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.142366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.142366Z digest=sha256:288ee9037149442fbfb3b8be789755e389eb19def2aa53853e176a5f563ed348

Observation 3cf9391e-3c58-41cf-aaeb-7e7b2b496281 · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.239306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.239306Z digest=sha256:e156ff302dcbdf4862edb75cd840bc44241f6297e8ef830312d19a07a9374eee

Observation b8cd82c8-14e0-4c5a-9b39-57dd03ce7e62 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.288372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.288372Z digest=sha256:fdc46bf90f3b42383379d93f44558bc0b126a1ba8fdb84abcfa0fd019640de66

Observation 976c396b-5b79-49c7-92ee-e309c1a1840f · outbound

This paper cites Guiding Long-Horizon Task and Motion Planning with Vision Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Guiding Long-Horizon Task and Motion Planning with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.390138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.390138Z digest=sha256:5230b0e0438b50c5f9033491eab1d15a862fd9b54d8aec4c30482b0b822e5c89

Observation 2760d490-e208-4e5d-ab2d-3edd5e53b8ba · outbound

This paper cites Open-world task and mo- tion planning via vision-language model inferred constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Open-world task and mo- tion planning via vision-language model inferred constraints

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.448540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.448540Z digest=sha256:204eb970363376ce9a9ce59c73bd92536b837ce13f3d0b41bbf2619b6b301b7b

Observation 11c21ccd-35e1-438d-bb10-5eca6d4f2576 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Physically grounded vision-language models for robotic manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.567384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.567384Z digest=sha256:a5a5deb360d8e4d0badd6adc7890852c35585bcacb25bdef3917b0195a9a9f2c

Observation dfc156ef-59c4-42dc-bcb1-abd204d73f86 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.613132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.613132Z digest=sha256:0d08fdfe8b2ed5a48ecea2147364d1658f69b92c74c8278b374177c6afcf7326

Observation 729f7f9d-31dd-4201-8883-8523075fde40 · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.665705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.665705Z digest=sha256:778f5f408fbadc9d53b51a172d204a88b683ae10f354f4e143d691cff26b5a30

Observation 853f45df-a4e2-43cf-b81b-ee8f622a8357 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.746754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.746754Z digest=sha256:28c935b137afaa0698e91f550b308abf6b0d2ed9f4609145f5f4a45398218c8e

Observation a4dba97b-2dcf-4f91-b96b-a501ea8e86d5 · outbound

This paper cites KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.801070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.801070Z digest=sha256:5d71f02e2f915cce8f12319640ef1c3b18d32c256ff6eb0c717106f871b1d544

Observation 580133e2-e101-42c7-85c3-fa14c6b5f52e · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.885154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.885154Z digest=sha256:90eb4689dc86dccdc398a3a7de230dc6ef25849e69eb62603a72b41f1afa7076

Observation 995d07c8-f01e-4779-be89-6c9c304d5a5a · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.958565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.958565Z digest=sha256:cef16b160b61e298fb0ecfd52cf9bce0552446057222323374bab5eab2c9ebd7

Observation b1375379-f124-451c-aa10-bf2537dc0f4e · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.998541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.998541Z digest=sha256:da5b98507631c5778b4bdd144dc00a4b6e76ca9173ae7cde5db6911abba8ae4f

Observation 9dafe183-77f9-43a7-b1fb-f0083b08dbf6 · outbound

This paper cites Learning to interpret natural language commands through human-robot dialog.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Learning to interpret natural language commands through human-robot dialog

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.730699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:04.066584Z digest=sha256:a7298dda10c2898dcc8b13b747ba3889fbc01024c6c35fedb6a6bfeae0c142c2

Observation aab21090-a360-4a35-9eac-4579e90a30e1 · outbound

This paper cites Grounding verbs of motion in natural language commands to robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding verbs of motion in natural language commands to robots

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.719822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:04.142753Z digest=sha256:4428ea304fd6cc30bd3ae7a2e9964c7c5da2bd1675bca503a051d4b0e3a791ec

Observation 386300ee-6e29-40fe-9074-074470b7b85f · outbound

This paper cites Toward understanding natural language directions.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward understanding natural language directions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.709353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:04.190234Z digest=sha256:42195ff8ed50ea782ccab10207915972471a751ef73e3d5337ad24e1eb4dcdf2

Observation c3121184-f8a4-4d0c-b472-90cd1def4285 · outbound

This paper cites Understanding natural language commands for robotic navigation and mobile manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Understanding natural language commands for robotic navigation and mobile manipulation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.699711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:04.271113Z digest=sha256:58d661714496c6f1b421465dd4ec316a30f943c50eae355d3310a34b34ab6884

Observation ef9b0120-ba87-41fe-8042-ff94900e6787 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.326924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.326924Z digest=sha256:601f1a8b497b190b2e3ac6a739e696311f452255a5ab21f760558954669ada2e

Observation 0f291216-b1a5-457e-a4a0-125447ba092c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.421149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.421149Z digest=sha256:ef189599d17da6d10f66e98dcbad41c9933252d20531b6a37f3a5b60f5702324

Observation eba4b8fc-77ba-492f-91d1-011bb97464b5 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.499840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.499840Z digest=sha256:af828ae0d8f3cbf1b85422aa3c53f203ab4b24416eb4e3b1655ee995f1a723d6

Observation 863da663-d828-452c-a62a-6c80cc8bc02e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.566817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.566817Z digest=sha256:04ae950c4b80858be0c97d384816c608de917ca5fe8a1c9ed256ad6981405d21

Observation 683d88e5-4fa9-4883-b4cf-83956b5d0b60 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.641939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.641939Z digest=sha256:ca406b78c281567ef5ad61fc588de27694022fe0e26318c9dce4a1a667226909

Observation 59868033-286d-4dd9-8f01-30a6cca9c25a · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.680233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.680233Z digest=sha256:158313e5dca34954de543b3e7dab3696e1923b847a93631160515649edc5d436

Observation 4fccaf42-5f7c-4145-b561-acea518e2579 · outbound

This paper cites ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.770431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.770431Z digest=sha256:8252a0b0a1dddbf757f653e522f4defc5504ebdc3d53aa0a049b65cb622f7636

Observation 173e3d0b-562e-4c94-860b-56bd7c878c87 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.843906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.843906Z digest=sha256:fd0f7a67561b421ecd2b07612ba92b95a4a04e14e0dfa8eb40111501e2dec052

Observation bac650b5-f260-44c4-a0b8-30472031daa4 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.884890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.884890Z digest=sha256:f9ae4b47e1c989aabe812941028688f04ebe6305d6c1f9696a43a534ac905007

Observation 1bb37cf4-a892-44fe-9a4d-6c3904cccbdd · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.988783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.988783Z digest=sha256:60876ed1785e347544d3b19f6f227450291a140e509a6b0bda122fd4ad7b3e92

Observation e9f30bba-37a7-4640-9ca3-92b965550e7e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Octo: An Open-Source Generalist Robot Policy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.061262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.061262Z digest=sha256:f7968644ce731ea66f4090dbf72f7e58f21c5b844c03b11490313c489c559211

Observation 74e98e85-7c9f-4150-862e-1c6478143aa8 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.111915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.111915Z digest=sha256:a69f3852dfbdd6aba4a0b7b081a6962963d39266a2cba442eed5799034176302

Observation 6d3f8af4-8d51-4a1a-a2ee-f9d3d15fa295 · outbound

This paper cites RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.207574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.207574Z digest=sha256:20e2656e994086183fd120ebb7b25d4f8474b373e39807c25d859891111569f0

Observation 0019b4a4-5b4e-4be1-a351-23cdb4da1fd4 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.276267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.276267Z digest=sha256:91169afd4d9b02146d3b0d45e22e3a9988ba09ebdb6f97f6a59940eb6527c9fa

Observation 431fd637-b19d-4a66-8bf9-2668e5200083 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.316266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.316266Z digest=sha256:a4a8b3db2574d925f418bc0699c2035d37aac4856cc01f016403ea937a125f57

Observation 42b8a48b-fb42-4115-9dd3-18e67ce0d8f7 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.418519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.418519Z digest=sha256:476e390e7105169266bde30d61376eba4862059486d5dbb88f44357e6aa28c66

Observation 5a081d2e-86fc-45f3-801d-ec30236c942d · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.478000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.478000Z digest=sha256:52e3a542e5acc94c6db215abc31a05deba9c733437da78fcdde1ba19699377d3

Observation 5053b701-0997-4d04-bd89-610c97133537 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.527284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.527284Z digest=sha256:35f1569bc93f9bb847c0f1caa16f30d4d6fe72251a7021eaf5a56022536e1653

Observation 792276d5-a035-444e-b92e-4bacef746649 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.638223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.638223Z digest=sha256:dc08cb85383e47528318a95e89360b887dcf57bd0425294e214ca508691a91c8

Observation 9e993ad7-d19a-4958-a9e1-6f1068b6278c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.690842Z digest=sha256:d55638c641d603c034b9cc3a82a657e2fd49f9a4dee68c5d386d8313799e324c

Observation 5ebc3e1a-ec6a-40e4-8276-dbf108045ac7 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.781091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.781091Z digest=sha256:205a00e6daafa35c0492457021b8ec0782382f8583699c55da3ab610f5301d4a

Observation c8e34603-f4c2-4037-8cca-a920b2473925 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.850037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.850037Z digest=sha256:ea0686c9ec2e8f63d693b803f28311fd47f5781294ab9e1dd6d0624b5f61bea9

Observation 948a6dba-6d7e-4e96-a0c6-8a10ee1607c8 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.893446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.893446Z digest=sha256:be94e7d98ce6b95208f85242aef77e4a81263dcf7954e0884000074dec61f620

Observation e3447b32-c4e8-4a2d-b955-4ed8aa480917 · outbound

This paper cites Code as policies: Language model programs for embodied control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Code as policies: Language model programs for embodied control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.953173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.953173Z digest=sha256:6ca6febb55651557bc4cf2d593429d5a7c7a022af31b3fad60231d7c875d84b6

Observation e5024913-6ebf-4446-a471-4ceca1687d6b · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Progprompt: Generating situated robot task plans using large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.677317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:06.080619Z digest=sha256:9aac92b7e8059e981b4c901b829b0067f40ec56b07fb5cdb128e3d8512b5c566

Observation 782b5a8f-d56f-4d43-85d6-09706f9e0715 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Chatgpt for robotics: Design principles and model abilities

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.667052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:06.222417Z digest=sha256:36a43d48bf8eeed8c641984ae5be739a015d096b4d39584c8a83824c82c6f781

Observation 549d0afc-d638-467f-9554-76eafc693ad5 · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.341930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.341930Z digest=sha256:f6b151fea6506241fd48448c8fd34e46c261b85ead740545ba5c14771c50da6a

Observation 33ee52c1-a2c9-4f89-9348-2d41e098b838 · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundation models defining a new era in vision: a survey and outlook

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.449159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.449159Z digest=sha256:53d18da451864cdf627d3aebf3ea9b3cb369a61c2e80028199def9e1b6c7ef03

Observation 4f3a9ea5-f4d4-4d98-b2e7-3ce1934b8957 · outbound

This paper cites Yolov10: Real-time end-to-end object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yolov10: Real-time end-to-end object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.513413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:06.591235Z digest=sha256:b71b60e179c03a4a2ffd63edfefd69f5b281dbd59f7bafa6da5eb8e4238711cc

Observation 762187bf-5d7a-4126-8efa-83b428afc9f3 · outbound

This paper cites YOLOv12: Attention-Centric Real-Time Object Detectors.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models YOLOv12: Attention-Centric Real-Time Object Detectors

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.730254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.730254Z digest=sha256:4b0401b091fed71b1e84e51f0e0e9f017c867f01db3f48315b89685cf38d02bb

Observation 10cdafa9-2d9b-4743-913f-e11df4e5d3cf · outbound

This paper cites Yoloe: Real-time seeing anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yoloe: Real-time seeing anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.841797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.841797Z digest=sha256:4ca6d83e40db8a66beb474e986483e03339174c6570fb1466a3c07632dcf8733

Observation 06330ba7-25e0-43ef-9875-729e32e584f9 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.986608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.986608Z digest=sha256:c6d5135f5ca64e325dc341e7ea24c4e5be35a19cad7aa179b859c1ae5c501f2f

Observation cb635328-fa44-4d6b-81b1-06ef8efcb96f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SAM 2: Segment Anything in Images and Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.109157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.109157Z digest=sha256:20ae98568e49f6125e53c1a7d0233d6ca95b16afd8571ea802e39e9a8d97248f

Observation 8b365f20-6304-4397-ab46-0b0f8a3e1c87 · outbound

This paper cites Fast Segment Anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Fast Segment Anything

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.220820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.220820Z digest=sha256:758c6b194617e29090c4415f183998602e05fd57fe8fbb6610d1456995c184fc

Observation e99c5c4b-e05b-42e4-aebb-d10330b70b74 · outbound

This paper cites Segment everything everywhere all at once.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Segment everything everywhere all at once

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.344271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.344271Z digest=sha256:68d86cf98a158bfab67dce619acf5ef972c56d2ec984f64c0a34b4e1c080e530

Observation 96c81ed4-59b5-4a9c-a761-ad09e2dac5ba · outbound

This paper cites kpam: Keypoint affordances for category-level robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models kpam: Keypoint affordances for category-level robotic manipulation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.351801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:07.470139Z digest=sha256:ac85b026d5ee880ddbbdf4a35c033d85ba3d847b6f6d4e156ab21ffc2e21868f

Observation c2077bc1-821e-4c72-8444-ed19d164cdc9 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Any-point Trajectory Modeling for Policy Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.568520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.568520Z digest=sha256:e1902473acd0bf041eade2896906edb90d60e66df4bdc46e45c1351fae098813

Observation 8da321ad-6616-4dbe-a41e-20a59c43a8cc · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.661869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.661869Z digest=sha256:e770f5ab836f5207d3c0d6bef36a0988516e223b847febd9e9234abbc98e0eaa

Observation 8b477562-60e4-4212-88fb-6c53e5443b26 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.113415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:07.861016Z digest=sha256:ac455585fa918f00ba529de268fb49145e4c214763ea8a7514fb5ecdeff8a6fa

Observation ee1f49c4-f2d0-423e-a5db-1332efd13cba · outbound

This paper cites Sam-6d: Segment anything model meets zero-shot 6d object pose estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sam-6d: Segment anything model meets zero-shot 6d object pose estimation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.842969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:07.911112Z digest=sha256:d02cc12a654d8498e2db7048b8ee3554ed2b9e1d806e72d1d59b0e642050da1a

Observation a807a592-c2bf-4450-a059-72ad9a73c895 · outbound

This paper cites Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.687810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:07.937920Z digest=sha256:67de0cd32d5ef3741cad9259846e99d23ea68d40f437e7379740894193c90443

Observation b6cd504a-d784-4ee8-b767-f44c04bce151 · outbound

This paper cites Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.485692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.021674Z digest=sha256:1e3cc4dac4cec290b15bc3ed110131dc5907b45055e3b0da174b9b5787d00a75

Observation b9fca9de-91cb-4823-9fcd-ca78b24839be · outbound

This paper cites Onepose: One-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose: One-shot object pose estimation without cad models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.329092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.096588Z digest=sha256:d14745d88b580c7351ae857b8d8328fa771356a5c334074b1dc8682c1341c29f

Observation fa9de27d-3787-4f29-a4f2-df1ca297a8f3 · outbound

This paper cites Onepose++: Keypoint-free one-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose++: Keypoint-free one-shot object pose estimation without cad models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.119342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.158463Z digest=sha256:318398d5a107a790f94f3295293fa919aa33082069f9ee61c797ed45c038ac6e

Observation 4b407ffb-5d6c-4cca-973e-98c757567aaa · outbound

This paper cites GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.497856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.250575Z digest=sha256:3033b813cfd0a72d91e76f1844dee467a250ca5fed14ae63931ce7f3da30a47e

Observation 746ef3ae-d489-4e7a-ba7a-ef9107518ad2 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.280921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.280921Z digest=sha256:e97b51adc16732bba1d4e3f62aadcf852c975fda5970a94ba2b4a656aac1c1bc

Observation f33928b8-72ce-43a3-a0f8-3f30a67c1475 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.358863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.358863Z digest=sha256:1a1ec08559e65b4ce35bdd00599b2b5ac1198cd2941570f87fbeab5017592a0f

Observation 6b33626f-3e6d-49ac-a012-b6b4d3b5ab12 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.430220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.430220Z digest=sha256:acb544ac81b4182489b30222a49a91884e273e835d67d7cf297428b81edd1aa8

Observation c570fe5d-f911-48a9-acee-dc952887e705 · outbound

This paper cites TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.483436Z digest=sha256:e9fae93559fd12e4ec069231e75a540b459a984d3243af72ac26ea3c16d9aee3

Observation c5c84f78-1d05-4bdc-b997-dee6732cd87f · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.982349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.570568Z digest=sha256:7fc4fe1cff88a21ff505a7d6472eac92667c26e982c09f8a1fc31d836694da9e

Observation 52372e6f-62d9-47c3-ba65-9551ac4a6178 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.608592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.608592Z digest=sha256:4ef7996806774f7f93879704f16b2da75409e8c195f7a91504fbb2b5170b22c9

Observation 3ce2b293-2739-4dba-8e62-d862d532f07d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.686757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.686757Z digest=sha256:0de833f66ad6f6fee1df5947148cd0af761c455dd2d751cfca3729db2f8fa88f

Observation 0a6daf92-271e-4862-b982-7d1c9c3cc231 · outbound

This paper cites SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.750059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.750059Z digest=sha256:7932c9cc3f63eb9367cda601400a060de58f0d950a9d46d8b92227d1b5e45c6a

Observation 53ca0041-7873-444e-ac7f-7068092a9c50 · outbound

This paper cites 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.785173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.827515Z digest=sha256:54a93ef7130d39f32b9bf1c6fc4a8dbf35977b269ea6cf50516c9160382aa238

Observation 76c4b5a8-b5b9-4cfe-99ab-e89056c77a1d · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.914602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.914602Z digest=sha256:5dcae90c79c70812298ab9512c2d7a901fd60380c2558e11af93dea6a5918488

Observation daadb4f5-6f63-488a-acb5-81076bc53dda · outbound

This paper cites Pointodyssey: A large-scale synthetic dataset for long-term point tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Pointodyssey: A large-scale synthetic dataset for long-term point tracking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.666620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:08.982750Z digest=sha256:6d811027e4057f7266e56694c1cf7db514a9ae5930b3938a56b54ac49b41e566

Observation 191d4ab0-115c-4010-960f-05535a61a749 · outbound

This paper cites Cotracker: It is better to track together.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Cotracker: It is better to track together

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:09.050283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:09.050283Z digest=sha256:045eae22090c9d7a251a7e8c6cfd15376512ff8f2bb0177155f728c0dffa69c8

Observation ae492222-1f9a-42d4-81cf-495c70edff99 · outbound

This paper cites Robotap: Tracking arbitrary points for few-shot visual imitation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Robotap: Tracking arbitrary points for few-shot visual imitation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.496351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:09.075532Z digest=sha256:d8a61e8070c3aa330ca632b4b19e4e6555d1d2abc8edbfdc4473a650eba89b10

Observation 83a67dae-83d4-4528-88d5-e146ac0b2397 · outbound

This paper cites Recyclable.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Recyclable

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.345022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T23:12:09.155546Z digest=sha256:274ac6b7a7b7f494ebd9dbd4b2d602c3407dd3631c504935e04ab79a3653f771

Pith citing papers

No inbound Pith citation observations are available.