Pith. sign in

Paper Citation Record · LEDGER

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2508.09032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09032 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:34:25.833722Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T00:49:13.291897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:08:58.805621Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c79802e-87d9-49c9-918e-f28944f26181 · outbound

This paper cites Inner monologue: Embodied reasoning through planning with language models,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Inner monologue: Embodied reasoning through planning with language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.542080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.542080Z digest=sha256:3f2b37367add9c972fdcbe391e68443e6a05919c0ec639aa1efa096c04b67c4f

Observation e50cfe5e-3f45-4182-9d51-6b138e65db61 · outbound

This paper cites Application of pretrained large language models in embodied artificial intelligence,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Application of pretrained large language models in embodied artificial intelligence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.547596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.547596Z digest=sha256:f213d36aff03830d0800c6be362bb95d6a4baa5e0e6ca4e1481f9d0c4c3f0554

Observation 7df9b5f0-a8af-4a57-a956-13f95bc731e5 · outbound

This paper cites Evaluation of pretrained large language models in embodied planning tasks,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Evaluation of pretrained large language models in embodied planning tasks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.552715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.552715Z digest=sha256:1d188096c16bdaf6fb6190494fa1a51ec5d01dd382a4a3a65b53410d93a873e8

Observation 7a88b2c3-33d7-4794-8f4b-6eaac602a5fe · outbound

This paper cites Palm- e: An embodied multimodal language model,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Palm- e: An embodied multimodal language model,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.449912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.557331Z digest=sha256:2dcea243dd40568b7307710cc94d1c1a77f45e194571588164d304bd4fae0541

Observation ba33e3a2-62e4-43f0-a00f-beda6566fed4 · outbound

This paper cites Common sense plan verification with large language models,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Common sense plan verification with large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.434694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.675049Z digest=sha256:796b61981ac584c047b414a771b72287a40098411718883699ff359e8ff2ab18

Observation 362e8a88-eb42-440c-bc1b-c59613465990 · outbound

This paper cites VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.680454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.680454Z digest=sha256:391c65d2cfe0146ac8439701e066b410eeaf5af040824ced2b362a7cdec2c1ed

Observation 27cd7e8c-c175-4aff-9c68-87813bca387d · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Llm-planner: Few-shot grounded planning for embodied agents with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.419539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.685823Z digest=sha256:a899c144609c6920d9ba2fc753f7066879112709181c427978ebb24f301d979b

Observation f385eb77-eff2-455d-a1e6-ca83ad8f847a · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.690219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.690219Z digest=sha256:4eb1e019af9f9a495d3ffb0ebffd491457f04eca1f6b09c32225bb600f1747d4

Observation f10d1671-62ba-4553-9245-7642b2582cb3 · outbound

This paper cites Lookplangraph: Embodied instruction following method with VLM graph augmentation,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Lookplangraph: Embodied instruction following method with VLM graph augmentation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.694982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.694982Z digest=sha256:afd2b9645ccf72ba5e246d61ba1b3b464ec8b26058a30bbfb447d75d7d458663

Observation c5f25143-2221-4230-a4cb-d9ca78cdc3c4 · outbound

This paper cites Lera: Replanning with visual feedback in instruction following,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Lera: Replanning with visual feedback in instruction following,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.393805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.699553Z digest=sha256:2916d77747660db26999c99598ef461b69e715dc5a219445c2034fb1349dd8d1

Observation 661dc391-4b9b-491e-9677-72f612a0bf8f · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding VIMA: General Robot Manipulation with Multimodal Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.703990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.703990Z digest=sha256:e8fce37430eb4f6badeaa859fa53f2dab1ec88ef8582da362c62bb3ad68efd1f

Observation 519e80d1-1c3c-4ed9-b04b-b5a2f2d0f326 · outbound

This paper cites Fine-tuning multi- modal transformer models for generating actions in virtual and real environments,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Fine-tuning multi- modal transformer models for generating actions in virtual and real environments,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.378478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.708901Z digest=sha256:19e053299d5d38e6ffff88c30e696238bdf77eda17d24d39fe293b1d7dce66c9

Observation 445cdcaa-2446-4380-a9a8-32ba03dbba72 · outbound

This paper cites Octo: An open-source generalist robot policy,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Octo: An open-source generalist robot policy,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.713763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.713763Z digest=sha256:a1e3cc4e9583251fc6f093e74de2b3f3a7344132d8812c1229cd5c7661c2d4d0

Observation b5dfc1e8-a6c5-46f0-be24-8517420a92e7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.718240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.718240Z digest=sha256:56d57179575384d764c642ff087b1ce84fa8373fd2aa5f294ab0db54c03cfd9a

Observation 1b91a08a-166a-43c7-9bbc-cb7f3ae0251e · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.722997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.722997Z digest=sha256:d81688883becb0569dcdf25b72415e75aada40766b1ee3b3d61b0ab866459d7a

Observation 7ff86df5-5a33-4c63-a3df-c7dddc97c685 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.727844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.727844Z digest=sha256:e603430bf3ea04e079f292c9c040a37f602bf593d7dfcc08cba53cb573caf59a

Observation d1e8a458-9d5f-432a-bcd1-7596f83da34d · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Robovqa: Multimodal long-horizon reasoning for robotics,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.352672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.732838Z digest=sha256:a4e17f19194bf508f7e4182cb8ef5de10e1a6c6b9c8b799caba50f7b3155802e

Observation 575497c8-4f22-48ca-94b1-1982e83188bc · outbound

This paper cites Mastering long-context multi-task reasoning with transformers and recurrent memory,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Mastering long-context multi-task reasoning with transformers and recurrent memory,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.336917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.737497Z digest=sha256:cd31db652ce109bd10b99394f1193d0cb1d96532221ab505160e2cd65920814c

Observation afab4d89-87c2-48d7-941c-16fd576c3400 · outbound

This paper cites Babilong: Testing the limits of llms with long con- text reasoning-in-a-haystack,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Babilong: Testing the limits of llms with long con- text reasoning-in-a-haystack,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.320255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.742016Z digest=sha256:7ab712f94a4a8265651287f51299048c5cecd5e4910141687f46c5d7bc70bc98

Observation 862d4eed-8104-40b9-bfd3-6ee28530e885 · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.746610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.746610Z digest=sha256:039e8a96f44ce4154b2fe7c35c5d716cd3fef26aec46eb430cff9c6977b7641a

Observation f01f22ae-6a1d-4d0d-bdd6-6350199ca179 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.751730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.751730Z digest=sha256:b2c6786b2e05b1064acd93e1b16582c22ba35f4f1f1ffe3ec04e7c16ebb74211

Observation 8f8fadbd-8a5a-42f5-833a-911185d0387d · outbound

This paper cites RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.756414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.756414Z digest=sha256:12ba440e13332f3509e1ea4073ff79fda121409f490e51aac6ab6736510db873

Observation 7343542e-ba85-4d74-80f0-4352ebac6e93 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.761568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.761568Z digest=sha256:4a6207c0adf5b955ea84de60478eb5282424d89238c479c33c4a941a37599a93

Observation 709475e8-3d5e-4c86-989c-a309ea588af4 · outbound

This paper cites Sem: Enhancing spatial understanding for robust robot manipulation,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Sem: Enhancing spatial understanding for robust robot manipulation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.767070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.767070Z digest=sha256:6706bb8fc17f36b93c96f3000a7936781c96b065a71a6a176a7a2d5e52edcb39

Observation 6dc7c7ce-30d6-4f84-8998-520fd08e5270 · outbound

This paper cites RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.771689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.771689Z digest=sha256:61cbbff547f644f07f5eca5b621213947651b8594fd63ac54773377e25e58d86

Observation f4c68434-6c73-4a5f-90b3-5b5f38204027 · outbound

This paper cites World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:34:25.974845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.777343Z digest=sha256:bfabb59cb9910a1ea5683d80851979edc9ea6c10ccaf7b48490e84b766df33b0

Observation 446949bb-978c-49ef-b901-6fdf9ff1c1ab · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.781988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.781988Z digest=sha256:fca46d08e5e9f27ff6b0ee1e3782e26dfcdf324ba9e104b9e77f10afbed90980

Observation 5073f28c-aad3-4f17-9c3a-72165fc0c8cc · outbound

This paper cites Magma: A Foundation Model for Multimodal AI Agents.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Magma: A Foundation Model for Multimodal AI Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.786667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.786667Z digest=sha256:e88674e9cf053a026dbb4e1f42ba2ce65f8f8ead1845c2b5ac36fd546e7ba460

Observation 259d631a-a390-4e41-84e1-78ce45374c78 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Sigmoid loss for language image pre-training,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.304612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.791653Z digest=sha256:d79f712888e44dd34bbbc8cc4ae4414d86920d9ca29d14ae59a860cbce184137

Observation fb5df039-beb0-426a-8126-77188708c701 · outbound

This paper cites Cotracker: It is better to track together,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Cotracker: It is better to track together,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.288849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.796117Z digest=sha256:e51b0bd4dbb2ab560401ed3caeca21f0cd7b528938a0d4093eb90ab7fe4bc3ee

Observation 732633ba-e588-490f-b434-6b081c09498d · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.800692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.800692Z digest=sha256:774741c2e014bcb2c6d83c785d5c2fba76c92d73710964a88bfebd75cb1834da

Observation ca905dde-d301-4d60-8301-9c66a6a2328d · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.805861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.805861Z digest=sha256:c7b4e217b23290c06a75c7f412bd6722fb278640ac4b06c140cb7c52473289c0

Observation 3470ffc9-d1b7-4435-93d4-d206091ef020 · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Evaluating real-world robot manipulation policies in simulation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.273758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.810457Z digest=sha256:ca8215e1fc4591fb056dedd00407df7ff1bfb184d71676792a316afe663b21bd

Observation 20d6ac6e-e5b6-4aea-98cc-074334fc4432 · outbound

This paper cites robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding robosuite: A Modular Simulation Framework and Benchmark for Robot Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.814757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.814757Z digest=sha256:f651c22cd29e5398f605f94ef1d72e323c244cdcb3c3159d9c89ace90e1e18a1

Observation adcc1023-9d2a-4357-bc9a-79654c1e0a0b · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.819442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.819442Z digest=sha256:895dfa3772e4dd31d90c4d72a6e7a99b9605177b02fe30390372e97c5d555e89

Observation fce96eb1-110f-4e64-93c8-2faaad3f6794 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Bridgedata v2: A dataset for robot learning at scale,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:34:26.258080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:25.824184Z digest=sha256:ca82161700c8d7397a53321bfa0aa821c9f02fc0824e7390ad19cb36b35145d0

Observation 8c452362-9e51-4aa5-ae8e-a55710f7966b · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Lora: Low-rank adaptation of large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.829203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.829203Z digest=sha256:f2b645b7bc48bf8c17cca88a9506cad5a1f56ad354fa244c3e5dbdd0ed7e4604

Observation d31ee40e-ac83-4eb2-9b8f-51a197c954c8 · outbound

This paper cites Decoupled weight decay regularization,.

Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding Decoupled weight decay regularization,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:34:25.833722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:34:25.833722Z digest=sha256:1f5d93da81e1bb8e0c0287317319523fef04d7427dfebead9f04f62654dbdfb6

Pith citing papers

Observation 2a29d3a0-d9c2-4eea-a9e5-33468fe17bd4 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.167232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:18d46ab8d345888509a4894f92eb97c129cda730a857e9d3bc44ba958b0fabd6

Observation b5664072-da6e-481d-a15f-a72a0b2156b5 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.611287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:368a0cbb7a45891b55c5314d2bd9b37d006931bceab24993dccce9ca1262486d

Observation 6dd17db8-49e1-424c-aa14-ecd4a8a41852 · inbound

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation cites this paper.

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.807388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T00:49:13.291897Z digest=sha256:e397e054f6d2bf19e5a84b53984cb775a3a05824446266407bf3a40873552e3b