Pith. sign in

Paper Citation Record · LEDGER

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.09448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09448 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:20:10.182935Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8af9adce-2dc7-4e0f-a2e0-de686139b078 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.061498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.061498Z digest=sha256:ebe50cad2da6fbbf63255908cc41ac727c96839d41953489827ee79d7afd780b

Observation d540df9c-ad99-4d5c-96e0-59a3c2b72b4c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.066611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.066611Z digest=sha256:bb95b214ef6207979364966dbe68bd3c3985ac0d5e95b0bda0efb85bd8992e94

Observation 26e91fc9-82ca-4b9b-9fff-3067fc20a1fc · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.070360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.070360Z digest=sha256:50eeec3956ee803944cf847504b1dd4f94d18344cea874951d1ce1aa987eb353

Observation 2ca6b3fb-0633-4398-a319-1160854f3757 · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.075185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.075185Z digest=sha256:16ebaf984c0bb4b5fc884ad4009ba589889fa39204a707dab5c0d39e93735a92

Observation 142c2801-48ca-4b6f-89bd-e6c2793c91f7 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-time training with self-supervision for generalization under distribution shifts,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.079556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.079556Z digest=sha256:400ba000e6f9b39c35e338b0b803aa3623613c6af1efafc41d8ffe59e184bf57

Observation c96c18bf-6874-429e-9c42-f13a36a879de · outbound

This paper cites TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.083589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.083589Z digest=sha256:d66ceb9bb2a1d0e22b685b9c47fbd9237f05faff81763a51e15f08341b036a34

Observation 2386901c-2bbe-4a1a-a360-a391b682802d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.087933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.087933Z digest=sha256:cbd634441cbfe362deb492ce226e46cf8d3a658e161a0b28aa4e39c44379b683

Observation 530d3316-935d-443c-987c-d563c8395ad2 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.091425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.091425Z digest=sha256:f0902893bde49d0625f98666a649ceba00a84d9d26f56fd72b822f025d91e3f9

Observation 30546ba1-37a7-4189-8d63-03e6ee442367 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Octo: An Open-Source Generalist Robot Policy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.095016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.095016Z digest=sha256:fdf869035e15c7b0b6d8047861749d5ab4125cd5a5e433d3b94432964cac825e

Observation c55b179d-bf12-4861-8f3e-2fd4017756b1 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.099143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.099143Z digest=sha256:ffef51206397aa5994d1aae2dbe29d833c1d78451b0fcddf6ef5adac0ef062ed

Observation 6da57c20-49e1-477d-978a-5d5fa303f2b6 · outbound

This paper cites FutureVLA: Joint visuomotor prediction for vision-language-action model,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction FutureVLA: Joint visuomotor prediction for vision-language-action model,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.102975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.102975Z digest=sha256:1ec2056ff14c6ed7c84a71774bd15c5f932d6f9bba64d4103eba6e7a8b8f5ec0

Observation ce1a0f9c-d55e-4cf2-b557-b0fc1de32300 · outbound

This paper cites Causal World Modeling for Robot Control.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Causal World Modeling for Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.106868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.106868Z digest=sha256:8b681b98baa1e28d96156aa015e8c66bf2b1099ebf6a685109f725712fe9542a

Observation 3864ed74-e1cd-412c-9e55-137c536323b3 · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.111381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.111381Z digest=sha256:5e1a09389d3afe698280a9de0fe001a649b3476f7083d12a08e40aa41a0d262f

Observation 0474f3cf-5908-4785-bf6f-568a8b5bc9c1 · outbound

This paper cites Continual Test-Time Domain Adaptation.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Continual Test-Time Domain Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.116348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.116348Z digest=sha256:f28783d5233ffc9a208b62823b79cbaaac219610ec708ec9d284aac602454524

Observation b7301b84-b894-4346-aabe-264d0dd3b191 · outbound

This paper cites Efficient Test-Time Model Adaptation without Forgetting.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Efficient Test-Time Model Adaptation without Forgetting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.120811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.120811Z digest=sha256:723b75339f5b8413b2dfcc57874d785322bf3acb0b02000e79c41b6b1b8449b0

Observation 3a9b7b0c-04b0-4e5b-ba74-3074f4b90492 · outbound

This paper cites Test-time training on nearest neighbors for large language models,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-time training on nearest neighbors for large language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.124848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.124848Z digest=sha256:60dfa0839c90ad705e6b2d406e0d4b95092a836082390c8506679bcb47d8dd5c

Observation c8ff7b68-ae3a-4b23-85b7-2c44caf3f6e7 · outbound

This paper cites Test-time prompt tuning for zero-shot generalization in vision-language models,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-time prompt tuning for zero-shot generalization in vision-language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.128394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.128394Z digest=sha256:ac2e5e4c89e0b836fd8cd35f2d1ff6a435163dd2610e9a21a0aa2fb4cf58485a

Observation 74e0c7a2-7a77-41cc-9f93-ccf8bb5c8277 · outbound

This paper cites Test-Time Training for Visual Foresight Vision-Language-Action Models.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Test-Time Training for Visual Foresight Vision-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.131763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.131763Z digest=sha256:e24f6bb5137d64708ce48e55e1f54fca0aaebd09e7bebf75f88b69338f180963

Observation 510231cf-0c98-45f7-aa44-d460ec712d80 · outbound

This paper cites On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.135778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.135778Z digest=sha256:097a5e49a3f671ba7e17cc99f02773eb8f32859a14a916f4fb31fd066d28f2fd

Observation c8e325b8-d7f1-4bc1-84d0-fd1450d1f2ac · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction BridgeData V2: A Dataset for Robot Learning at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.139974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.139974Z digest=sha256:3c366e027974a1988f5c7ed6cbab6fe35c14f1df522a5557a534593226cdf4a7

Observation 66a2923f-3800-4a3a-8592-9cac3a8b9aa3 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.143939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.143939Z digest=sha256:c6beed2d5c78b5a9d252221314c0169a0baa2d951ac1e28dccc8ac71cc572e15

Observation c1169a87-6de4-474d-9f74-4b48ae8ff20a · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.148239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.148239Z digest=sha256:2cb5644df5e43b27a0ef8641d41b7dd807ba32effcc8bb9ce68769180a60fec6

Observation 8f862a66-1709-4dc3-9993-1b13d72ec0b7 · outbound

This paper cites Magma: A foundation model for multimodal AI agents,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Magma: A foundation model for multimodal AI agents,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.151970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.151970Z digest=sha256:c4c46748bd2f804299e29a683c47cf47372ab0342c46f0854817cf0d388c7678

Observation 6c4d3083-627a-4d0e-8965-24c2f5a6c559 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.156291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.156291Z digest=sha256:173f438d7bb0e00add3eda96a0986eedc67fabfe9371e002b30b0df6ef68debf

Observation fb691054-b0e9-497f-bad4-6b78624c3ea9 · outbound

This paper cites InstructVLA: Vision-language-action instruction tuning from understanding to manipulation,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction InstructVLA: Vision-language-action instruction tuning from understanding to manipulation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.160193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.160193Z digest=sha256:25aa7aa249a5cacf4f7b85ab8510310ffd88c6f11327a0afc2f6a526035b9de7

Observation 23ea9715-6dcc-478d-915c-f39e9ed88c6c · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.163707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.163707Z digest=sha256:fac1cdbf62d1a9ceb0c306c7e644277b2108541aa175f0ba1280a974a0a4d1b6

Observation 22fd791a-005a-4cb2-868b-be8ac04b3595 · outbound

This paper cites ThinkAct: Vision-language-action reasoning via reinforced visual latent planning,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction ThinkAct: Vision-language-action reasoning via reinforced visual latent planning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.167723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.167723Z digest=sha256:357d986bbf174a51457849d3d996300c299ee6e21c78f97235ebc139dc36cf1a

Observation 5d16377d-b58c-4042-83c1-c9121d8308ff · outbound

This paper cites TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.171516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.171516Z digest=sha256:27a0a7ce177b5e7030048fa6f2f8736fdf8be1ec28b67a31af6b5b419e04b590

Observation 4578b24d-9fc7-49a8-816b-69b7f5d34e25 · outbound

This paper cites Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.174839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.174839Z digest=sha256:7324a12e8d73fa3934e04457605dfefbfdb7823e7868fd83bbe40ca9450c608a

Observation 6f28aa92-404e-4f33-81b5-df9c888393a7 · outbound

This paper cites FAST: Efficient action tokenization for vision-language-action models,.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction FAST: Efficient action tokenization for vision-language-action models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.179001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.179001Z digest=sha256:3ba1aff9bb9a528599e8bbc31cd4f8f1da2e6e2601c15835cbbe7d42ff34f388

Observation d29cd0b6-3f15-494c-adcd-f4730ffa8fd7 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.182935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.182935Z digest=sha256:8e7c4c6393cf594cb5f6cad9409f1a121bcc70a66318378da1e21177931314c3

Pith citing papers

No inbound Pith citation observations are available.