Pith. sign in

Paper Citation Record · LEDGER

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2508.09032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09032 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:34:25.833722Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T00:49:13.291897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:08:58.805621Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c79802e-87d9-49c9-918e-f28944f26181 · outbound

This paper cites Inner monologue: Embodied reasoning through planning with language models,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Inner monologue: Embodied reasoning through planning with language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.542080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.542080Z digest=sha256:bd0112807083d4518998b3d824e2714546f66f78dd2d4b7b767d33b505d4c735

Observation e50cfe5e-3f45-4182-9d51-6b138e65db61 · outbound

This paper cites Application of pretrained large language models in embodied artificial intelligence,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Application of pretrained large language models in embodied artificial intelligence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.547596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.547596Z digest=sha256:d927fe46d1701a09a1c4c94acb4800e8791722d74c0b177bd0fb02668b13beeb

Observation 7df9b5f0-a8af-4a57-a956-13f95bc731e5 · outbound

This paper cites Evaluation of pretrained large language models in embodied planning tasks,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Evaluation of pretrained large language models in embodied planning tasks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.552715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.552715Z digest=sha256:52f65c00b9b0e588d93b47f09bffb4860a5dc95e91b2f1bc9982128b42d38701

Observation 7a88b2c3-33d7-4794-8f4b-6eaac602a5fe · outbound

This paper cites Palm- e: An embodied multimodal language model,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Palm- e: An embodied multimodal language model,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.449912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.557331Z digest=sha256:a28404e3ac14fccfbfb6cd31454d7ceeac42a9b381736313198b785058385981

Observation ba33e3a2-62e4-43f0-a00f-beda6566fed4 · outbound

This paper cites Common sense plan verification with large language models,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Common sense plan verification with large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.434694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.675049Z digest=sha256:2a2129e79f3dd4f0e38ab0a31fa116b1cbe5fd1835fc070f94f1e0c543b20783

Observation 362e8a88-eb42-440c-bc1b-c59613465990 · outbound

This paper cites VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.680454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.680454Z digest=sha256:037056734eb651b601e7fa28f13ba132645de545adf45123e773568da92619e8

Observation 27cd7e8c-c175-4aff-9c68-87813bca387d · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Llm-planner: Few-shot grounded planning for embodied agents with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.419539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.685823Z digest=sha256:032a04fa5c8ab584c5e0446ac489f60e09b57984cae51126573eb9f1ab5212c2

Observation f385eb77-eff2-455d-a1e6-ca83ad8f847a · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.690219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.690219Z digest=sha256:4f5363c4ffe5bd0b0fa8d1c7670ce5707bfc0e3143af47e1cbf3e0d2b418f5ce

Observation f10d1671-62ba-4553-9245-7642b2582cb3 · outbound

This paper cites Lookplangraph: Embodied instruction following method with VLM graph augmentation,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Lookplangraph: Embodied instruction following method with VLM graph augmentation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.694982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.694982Z digest=sha256:4ce6fea3ba64c34bd3dfb0735dc1744384fb96b9a10a048f45c4b509bd924e41

Observation c5f25143-2221-4230-a4cb-d9ca78cdc3c4 · outbound

This paper cites Lera: Replanning with visual feedback in instruction following,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Lera: Replanning with visual feedback in instruction following,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.393805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.699553Z digest=sha256:c4e3ddc936e1918c21d6eb65e68e900040e2166bd3f811ea93084cc7657b5248

Observation 661dc391-4b9b-491e-9677-72f612a0bf8f · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding VIMA: General Robot Manipulation with Multimodal Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.703990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.703990Z digest=sha256:f3457e18dfe48f278cba995e17be6e187385ede62a8e0308a4861d4f9908a0f0

Observation 519e80d1-1c3c-4ed9-b04b-b5a2f2d0f326 · outbound

This paper cites Fine-tuning multi- modal transformer models for generating actions in virtual and real environments,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Fine-tuning multi- modal transformer models for generating actions in virtual and real environments,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.378478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.708901Z digest=sha256:ad4abcff999ab1aa46c56e0042ca8c5357c3e50230d9fe0d3e7955499e10d636

Observation 445cdcaa-2446-4380-a9a8-32ba03dbba72 · outbound

This paper cites Octo: An open-source generalist robot policy,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Octo: An open-source generalist robot policy,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.713763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.713763Z digest=sha256:a6d004a481944244450738f7005b1830230ea444198108e23c08a903224be0c8

Observation b5dfc1e8-a6c5-46f0-be24-8517420a92e7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.718240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.718240Z digest=sha256:ffc686c674a185e4af7fe4efc3985cf9d0df7822ce17f11686d95b528ce4e14b

Observation 1b91a08a-166a-43c7-9bbc-cb7f3ae0251e · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.722997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.722997Z digest=sha256:c9c0d67381f60644f31d94a7597f008f08dcf82b2d1600692326de626e58ff99

Observation 7ff86df5-5a33-4c63-a3df-c7dddc97c685 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.727844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.727844Z digest=sha256:e20e34079e26fc970a356a05938eda9ce999d50f20d2afe53e05d0a7b02db568

Observation d1e8a458-9d5f-432a-bcd1-7596f83da34d · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Robovqa: Multimodal long-horizon reasoning for robotics,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.352672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.732838Z digest=sha256:33098614299faa2a29a1412bab3e8836517a282699ab3009d9554359eb69bf6a

Observation 575497c8-4f22-48ca-94b1-1982e83188bc · outbound

This paper cites Mastering long-context multi-task reasoning with transformers and recurrent memory,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Mastering long-context multi-task reasoning with transformers and recurrent memory,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.336917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.737497Z digest=sha256:2dc64a0ae4f61bca1a6f9d56a4580c353c4b82611dbc3d12d5044879becea06a

Observation afab4d89-87c2-48d7-941c-16fd576c3400 · outbound

This paper cites Babilong: Testing the limits of llms with long con- text reasoning-in-a-haystack,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Babilong: Testing the limits of llms with long con- text reasoning-in-a-haystack,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.320255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.742016Z digest=sha256:4133338d9a7285c4f8144c00cf5ec8b7429bddee01aa38d4fa20aab2e1576e65

Observation 862d4eed-8104-40b9-bfd3-6ee28530e885 · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.746610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.746610Z digest=sha256:fdf79f51ee2f7057b7bb902c5b71016f31fecd9a6a7073a5df13e9c707493708

Observation f01f22ae-6a1d-4d0d-bdd6-6350199ca179 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.751730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.751730Z digest=sha256:91f704c2991e872f9637677487d34dc47a76a3c54c4289df734c9b23ab7f6322

Observation 8f8fadbd-8a5a-42f5-833a-911185d0387d · outbound

This paper cites RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.756414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.756414Z digest=sha256:d32d39ec9e99c65ee3993b59ba508f6cd7582841d715cd00d11afbc53169d860

Observation 7343542e-ba85-4d74-80f0-4352ebac6e93 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.761568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.761568Z digest=sha256:36626732a8241c6ebe94d3b01929415ad4d50cc8f6990d0603803f36ad99f3f5

Observation 709475e8-3d5e-4c86-989c-a309ea588af4 · outbound

This paper cites Sem: Enhancing spatial understanding for robust robot manipulation,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Sem: Enhancing spatial understanding for robust robot manipulation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.767070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.767070Z digest=sha256:f76bbb5099b2587877699d89d046c0f7c2a6eed0fe882fab04ada76eb67b580f

Observation 6dc7c7ce-30d6-4f84-8998-520fd08e5270 · outbound

This paper cites RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.771689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.771689Z digest=sha256:67ad342ae6b40c4a20f0a21944b84bd9683783009983f481fb26c3d70f824a0c

Observation f4c68434-6c73-4a5f-90b3-5b5f38204027 · outbound

This paper cites World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:34:25.974845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.777343Z digest=sha256:5593576dae6b45353d1234f92cf5fd7aa3cb613ed27a681325db092954ffdbec

Observation 446949bb-978c-49ef-b901-6fdf9ff1c1ab · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.781988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.781988Z digest=sha256:a37c554ce6067cd37e71870705d5930cfa28032b4539b69c16bc252936239a10

Observation 5073f28c-aad3-4f17-9c3a-72165fc0c8cc · outbound

This paper cites Magma: A Foundation Model for Multimodal AI Agents.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Magma: A Foundation Model for Multimodal AI Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.786667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.786667Z digest=sha256:c2a18dc1adb6e8bbd32e82c465b547b30a0e9e719e69bf2536cd0f8ae1f1e92c

Observation 259d631a-a390-4e41-84e1-78ce45374c78 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Sigmoid loss for language image pre-training,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.304612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.791653Z digest=sha256:20e2f0608ac68975d6121debed0a0ecb5ec1c4d8c3c0e31b61e91aaf11627492

Observation fb5df039-beb0-426a-8126-77188708c701 · outbound

This paper cites Cotracker: It is better to track together,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Cotracker: It is better to track together,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.288849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.796117Z digest=sha256:4388057215711305201757fd4f2f9e3e6c715800a83d63965d3cc180b5e0767d

Observation 732633ba-e588-490f-b434-6b081c09498d · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.800692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.800692Z digest=sha256:4c8a0b0cac8d62f8f72b7a74778fd905284da89f85777b60571bfce097346027

Observation ca905dde-d301-4d60-8301-9c66a6a2328d · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.805861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.805861Z digest=sha256:1e89ea73c18255414cca0f2f8087c0331720d6da2ac0ec307794c767abb74fcf

Observation 3470ffc9-d1b7-4435-93d4-d206091ef020 · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Evaluating real-world robot manipulation policies in simulation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.273758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.810457Z digest=sha256:aab1094999d03a6e9bac91f2ba62f334bbf6a8da15c892b39f24cee90da20cf3

Observation 20d6ac6e-e5b6-4aea-98cc-074334fc4432 · outbound

This paper cites robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding robosuite: A Modular Simulation Framework and Benchmark for Robot Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.814757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.814757Z digest=sha256:def52305385b9e5146e880ae19c3de5b6a9b99648ac7fee6d7d10143006d3393

Observation adcc1023-9d2a-4357-bc9a-79654c1e0a0b · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.819442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.819442Z digest=sha256:53e58339cf21599d5525b172fcd6e534b0506bdc458f9fc2bac7d995d291c7ef

Observation fce96eb1-110f-4e64-93c8-2faaad3f6794 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Bridgedata v2: A dataset for robot learning at scale,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.258080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:34:25.824184Z digest=sha256:b287b9a0cb06fb9f3bc879c78203b93fad8918efb3b18e4738812d60ace325dd

Observation 8c452362-9e51-4aa5-ae8e-a55710f7966b · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Lora: Low-rank adaptation of large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.829203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.829203Z digest=sha256:6bfa86b9bcf8b90e5ffbb2a92c832d4749b5148360dc1773dc0e8ee91eda6590

Observation d31ee40e-ac83-4eb2-9b8f-51a197c954c8 · outbound

This paper cites Decoupled weight decay regularization,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Decoupled weight decay regularization,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.833722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.833722Z digest=sha256:5a621c0b6c0de62e0f312f6144471791455edf2b45b683fdbb68267788145079

Pith citing papers

Observation 2a29d3a0-d9c2-4eea-a9e5-33468fe17bd4 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.167232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:e136b778087640a56ad9919799ea1948a5d805baddaad5d83bd98fa2301bdd97

Observation b5664072-da6e-481d-a15f-a72a0b2156b5 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.611287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:d1fda27590df8cc7d3fe2582d502353e2ffb6d1046b78bf686d47b027de261a2

Observation 6dd17db8-49e1-424c-aa14-ecd4a8a41852 · inbound

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation cites this paper.

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.807388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T00:49:13.291897Z digest=sha256:388aeb77e80704418d079276c29dd551d1056ac0656e0a4007d32603d9bbbb74