Pith. sign in

Paper Citation Record · LEDGER

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2608.11671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11671 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:39:40.605712Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact4
  • verified fuzzy15
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f5c8d6a-811d-4c38-ab78-5befcc6692e4 · outbound

This paper cites Qwen3-VL Technical Report.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.159696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.159696Z digest=sha256:73f36157fe9a8825392079f1d8d4a21022717bffedc3fbdbfefed2d6346c8d55

Observation 67d75d5f-0648-4b58-b217-e545c44d0264 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.166058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.166058Z digest=sha256:6dc21834d04063eac8886833de30a94235c2ade5c4c29b80120017dc03f577e1

Observation 6dac024d-934a-4e63-9c28-17108811c050 · outbound

This paper cites Motus: A Unified Latent Action World Model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Motus: A Unified Latent Action World Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.171426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.171426Z digest=sha256:1c4466b975193dd4740cc44d399e194f305e4887222be7cb334f87fb7bd7d095

Observation 3d188d7a-3dcb-41a8-9f41-95057c8b2db6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.177152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.177152Z digest=sha256:fd362fe9c967d22564960310df45db0bb78d203dd31cf9cb270ca974a0269abd

Observation bc553b3e-ccae-4b9a-9985-c34c2ecf9c15 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.187843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.187843Z digest=sha256:f765202e94fabb179cbb5ec4ad5f6552e21f6d0fbe62a8b3c6373d8f8a0cec7e

Observation 0d1af0ab-83cc-4257-9d54-d35208203a99 · outbound

This paper cites See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations.arXiv preprint arXiv:2512.07582, 2025.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations.arXiv preprint arXiv:2512.07582, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.193129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.193129Z digest=sha256:a70ff02bde881ebf618b963e0b96e02431fd7d9a4122bd12337148471bc0ffe8

Observation 35e919bf-d023-467d-b3ed-a0256c2a742f · outbound

This paper cites Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.198497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.198497Z digest=sha256:ebb5265c3c6c1464525eeae6116641fddda96fba51ecf049a664a332868a9011

Observation 78dbd86f-6df9-4060-87f0-f9c36e3b9296 · outbound

This paper cites From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.204306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.204306Z digest=sha256:3ffb9492eeac366766d166ded654b38db284b7d9c04fbcd73b3875db6ab4fdca

Observation 92387e10-3135-4277-ae10-736ecc3fef1a · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.209987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.209987Z digest=sha256:3f48afe31f1a20024106e40142a69b3b8652862a09a8a7a3cb4a96249e8a4702

Observation 59b30ab7-0fa0-4397-a48c-9c111f559777 · outbound

This paper cites See what matters: Differentiable grid sample pruning for generalizable vision-language-action model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models See what matters: Differentiable grid sample pruning for generalizable vision-language-action model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.138077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.215675Z digest=sha256:8e072fb3466737ae30d0033c91b7dd2b0f21d273128488e08e96ef828cc61283

Observation e1295b98-35fa-4107-820e-4a881ae8b493 · outbound

This paper cites WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:41.579937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.220903Z digest=sha256:47a2b9007b2cd5df0db1323a55d8068ab9687a5058d570e6df9e2a4baefedf4f

Observation f09b947f-1824-459f-b2ee-31d585b4c5ce · outbound

This paper cites Icrt: In-context imitation learning via next-token prediction.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Icrt: In-context imitation learning via next-token prediction

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.118206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.226605Z digest=sha256:abed3a9a49356b4cfb4e14909701359ef625c6680c9192e40afd44d6aa196564

Observation 5ca88b4a-50bf-46f0-afc4-e195fcc9caa2 · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.231644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.231644Z digest=sha256:78e3557233f62d1c034c53c66c04f82b876e50141a2469f8a7bc23720c0226a6

Observation b74b918f-6ff3-4a17-bc9a-f7d4e5c201ab · outbound

This paper cites Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, and Anirudha Majumdar.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, and Anirudha Majumdar

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.237022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.237022Z digest=sha256:1ac2f3c6d34806c453c8f13837b75599bae1627fee87a9ad025b00cec151a71f

Observation bc03fbe7-39c9-4a07-a4cb-08641716db5b · outbound

This paper cites Motion dynamics learning for few-shot embodied adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Motion dynamics learning for few-shot embodied adaptation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.098520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.241659Z digest=sha256:b298357798b4353bdae7b6821e4f7c2dd5c1e4f8a24265efd0e015f9b0e1d78b

Observation dc9ad96b-87e0-4072-bc9b-3499977baaba · outbound

This paper cites Thinkact: Vision-language- action reasoning via reinforced visual latent planning.Advances in Neural Information Processing Systems, 38: 82782–82802, 2026.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Thinkact: Vision-language- action reasoning via reinforced visual latent planning.Advances in Neural Information Processing Systems, 38: 82782–82802, 2026

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.079782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.246629Z digest=sha256:edc1843709004bf91a18ba19dfe53bb46315743a7445901698b3fcf59456d485

Observation 21ad72bd-a67f-4dfd-a627-4ab833b294fc · outbound

This paper cites Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.372922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.372922Z digest=sha256:0cbc9c667068f2b4f7983080f3a67feda95d5f7503062cf3d054c8588b4abed3

Observation 289bd257-aba7-4d12-a147-a9f0cf32edeb · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.378753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.378753Z digest=sha256:c620c8f16780bcc7c26daaf3ec6c3cf51f057764afae94912354b1f75a9462b6

Observation dd1e365e-fbf5-4751-a3d7-ed85140f453d · outbound

This paper cites Ra-vla: Retrieval- augmented vla for test-time adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Ra-vla: Retrieval- augmented vla for test-time adaptation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.061087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.383828Z digest=sha256:ac227aa35f4178ed9794191b9b3fa7dbe9a9e4deee5b127c2aa1c3ef20038f07

Observation 767bc788-6f2e-467b-b80d-9b6644aeec2e · outbound

This paper cites Ra-vla: Retrieval- augmented vla for test-time adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Ra-vla: Retrieval- augmented vla for test-time adaptation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.040916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.389054Z digest=sha256:8e88ef1e9c860163e99fd7b7fb8ee44d3d037ba6b87a9a16ab5b353c458457ea

Observation 66fbaafc-a41f-4c2b-b8aa-67c6d9b5983c · outbound

This paper cites RoboTTT: Context Scaling for Robot Policies.arXiv preprint, 2026.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models RoboTTT: Context Scaling for Robot Policies.arXiv preprint, 2026

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.019250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.394348Z digest=sha256:4b24f3b511d52cc3ed5751af1e0df48210c4774cb3dc6ccc32c4cdcd3498905d

Observation ebb4701a-c999-4648-a9d6-2ed844b4d514 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.399303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.399303Z digest=sha256:040c449233caa4e4f4f5639e82db92baa2e87bcacde80f14e472b6b85d2f0a6c

Observation b433c465-28c0-4c40-9039-47a72fc38fa3 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.405179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.405179Z digest=sha256:5b9f657e1f2fa6be44182fdc477b36be9525ce0ae701c191d1a9454a7af79a40

Observation 0faafd41-c780-4e71-95a5-70a99bf761fd · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models MolmoAct: Action Reasoning Models that can Reason in Space

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.410246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.410246Z digest=sha256:e34fc4e2b0e3b1d1a56a96b71952c0f02eae64fd02108f1c3346c0d25531e2d4

Observation af116d28-9cb8-4d60-be99-bb45c4d5b95f · outbound

This paper cites CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.416285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.416285Z digest=sha256:e20a4ea373ae87ffaf1e3d73d82592eab970063b6c19d0eab26282f4246ff40f

Observation 81313a33-46db-4747-aafb-d1ca6837f4af · outbound

This paper cites Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.421768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.421768Z digest=sha256:6b544c2429f05a00e4f2d6966ec573940a5df7c111c079fe5ca527293d4b14b5

Observation 6a9848be-5e94-4b18-97ae-ed3eaedd743d · outbound

This paper cites LA4VLA: Learning to Act without Seeing via Language-Action Pretraining.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.432225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.432225Z digest=sha256:2cfedbbd45ff631e3a8555e2c42d827c4cbd4a358e9a2b20204e4e34d6f4e873

Observation 4b40be1e-1cc9-498f-b0fc-8bb09341e4e2 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.437052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.437052Z digest=sha256:92c4c0499eca063a18762818b4e743022be2660bc2fe1c6201a9880b37d3ab9f

Observation cea6bdc0-b60c-4b96-8748-1bd43bed470a · outbound

This paper cites LocoFormer: Generalist Locomotion via Long-Context Adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LocoFormer: Generalist Locomotion via Long-Context Adaptation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:42.000853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.442473Z digest=sha256:123f422e9850a0baa9f27dbb0385fde3b8d0d45a82c11cbc68ea5ffdcd08a29b

Observation cb272567-7abf-4b62-a6ba-ad96b84fc08e · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.447818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.447818Z digest=sha256:4f3157f416fa43a63f15e8c7a0118c54ffaf3a90ae70598537c4e9008d69ef08

Observation 984c62fa-cf3d-49b1-b3d5-6a27c331b914 · outbound

This paper cites Behavior Prompting Policy: Demonstrations as Prompts for Manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Behavior Prompting Policy: Demonstrations as Prompts for Manipulation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:41.258504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.452698Z digest=sha256:6f4962477994bbab9f3b78b4411a58334d16e3c354fc430d2fb6ad548f02375b

Observation c8b74a6f-d11d-4d31-bba9-b692ebb2c443 · outbound

This paper cites Action-aware dynamic pruning for efficient vision-language-action manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Action-aware dynamic pruning for efficient vision-language-action manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.982044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.458321Z digest=sha256:c0f3554e7d3e31e46e1528ff07efe5bb014016ce0a2c3346934b7775204b60ee

Observation 85f6ee72-9c7b-4f32-9dcb-9b5b11063e0c · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.464253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.464253Z digest=sha256:ed08e357e26776b675c5bd8027473a629fa1f6ae2e322477734704cb5f215aed

Observation 93bb0d3b-252d-4ab8-bff1-f5be796a5a06 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.470070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.470070Z digest=sha256:721af42f946193f7750d8f369b542fbe2ab7e4eec151e5730a92a78146c83a27

Observation 0b5fe059-9135-45c7-92f3-d01d315ab801 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.475421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.475421Z digest=sha256:50dfadee0a4076f596b6b7d6fa6569ff862706c6df4b3331de23ad7994d0e860

Observation 70f45c2d-7d60-4706-93b6-f443ff7cef90 · outbound

This paper cites Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.480220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.480220Z digest=sha256:6542ccf60f26fbdd541ad0638581a4d16d1eb01584995f0fe926e8c0c2f37f01

Observation cc00cbe3-e0d0-45ba-b9f1-25cf18ce645f · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.485551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.485551Z digest=sha256:edaa369b58239824fa8c5ef234d73fc560bb22ff312f4e1fc42deb506bec2082

Observation a0b64ab9-d276-4788-96aa-e8b824aea016 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.491234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.491234Z digest=sha256:b8c835dedbe37d748caba6c9c096628c071fbcadeda7dd8f51d55bb88d3c424a

Observation 7b68db90-0d21-4b5d-8502-9ea65dd1f2be · outbound

This paper cites RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.496275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.496275Z digest=sha256:bdc423a1f67ce38ec34dd961b84a163e6582c59be175d5c8deb27cb1acbfa6c9

Observation f1865d81-e0a2-4a80-8d75-fcf68bc8ea9b · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.501945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.501945Z digest=sha256:9fb29a80a45aa90aa46009b348a1d0b4786cd46f81a6ec2e3a090fd57dd09eaf

Observation 19f7e96f-fdaa-4a2a-988f-5868b448f60d · outbound

This paper cites Test-Time Training with Self-Supervision for Generalization under Distribution Shifts.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Test-Time Training with Self-Supervision for Generalization under Distribution Shifts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.508228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.508228Z digest=sha256:94a0f39338904d98549a4b4e5a0d30ff7375e2a86d87a89995aa8d3ac6224ef4

Observation 41f76f0c-886c-4559-8c30-c280a22a1732 · outbound

This paper cites X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:41.067131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.514334Z digest=sha256:3585ea32ee7a5da3509df2310faac8ec90d8de4de69e792a4ff1309a5e41a75e

Observation 479b66dc-ecf4-4bec-b458-164d4580e070 · outbound

This paper cites From Foundation to Application: Improving VLA Models in Practice.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Foundation to Application: Improving VLA Models in Practice

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.520150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.520150Z digest=sha256:8af2f945a96ad5f929cf5f5ac34ad6364d476c5822d594fdc820b6be14757b71

Observation 70ca2af7-44b9-4716-b5d6-e4d845a1c7b2 · outbound

This paper cites A V A-VLA: Improving Vision-Language-Action Models with Active Visual Attention.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models A V A-VLA: Improving Vision-Language-Action Models with Active Visual Attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.964349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.525612Z digest=sha256:9fb0fd9aa3073ae4d7c9c6e44f50d3c7d6a45ce5309a9808b2f65f928f118d95

Observation afe0bddc-4a58-4ea8-9541-71736e7cf7f9 · outbound

This paper cites Towards efficient embodied reasoning: Mixture-of-depth compute allocation for vision- 17 language-action model.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Towards efficient embodied reasoning: Mixture-of-depth compute allocation for vision- 17 language-action model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.947017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.530558Z digest=sha256:38ddbd5e565031b055897666ce5ecbc0adc09385f463e52f1484701c4b638d99

Observation 878b0789-b711-4571-8299-f330b4f86d39 · outbound

This paper cites Vla-cache: Efficient vision- language-action manipulation via adaptive token caching.Advances in Neural Information Processing Systems, 38:164448–164473, 2026.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Vla-cache: Efficient vision- language-action manipulation via adaptive token caching.Advances in Neural Information Processing Systems, 38:164448–164473, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.535540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.535540Z digest=sha256:452da50f9c7574b380eba2ea4b7c5812da88abafaa4974b3664fac7dbd314b1d

Observation 45094728-647b-4228-89c2-1813c683421f · outbound

This paper cites Affordance field intervention: Enabling vlas to escape memory traps in robotic manipulation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Affordance field intervention: Enabling vlas to escape memory traps in robotic manipulation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.918310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.541031Z digest=sha256:6c8150d14cc401144d7e4032e4cdec467a922cfd785b0b8b77b715386e7c9d66

Observation ce9f1612-a1dc-4988-9b10-dde7eb44ebcf · outbound

This paper cites Latent Action Pretraining from Videos.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Latent Action Pretraining from Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.545967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.545967Z digest=sha256:b544e8961734bdf729cbd67c4cff7229973ed24daf607480855fa32a756182da

Observation 60d41e58-ce4f-472d-984a-7d2b30795eae · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.551955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.551955Z digest=sha256:027b2e060e722a0f998877410391b203b095123a10d410722512fe95c8b1b931

Observation bc9939bf-886c-425e-a7e4-3c06e6301528 · outbound

This paper cites Hancock, Mingtong Zhang, Tenny Yin, Yixuan Huang, Dhruv Shah, Allen Z.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Hancock, Mingtong Zhang, Tenny Yin, Yixuan Huang, Dhruv Shah, Allen Z

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.557935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.557935Z digest=sha256:562f5e07fe27c1b06c03570528440d3936c8c6569216371e25961d9565d06987

Observation 4b74d118-0022-4d72-b5a7-8e64f05164a0 · outbound

This paper cites VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.562799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.562799Z digest=sha256:904c9276d2af9a14cbf9953e749c1883ad193f5ed7fe3e59e954f9ec3ae28568

Observation 17528a3f-09df-4915-b63c-4472aee2cf7d · outbound

This paper cites Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:39:40.751326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.568412Z digest=sha256:980cf9d3879d2d4c90f0c9bf2e6fb5e9f274c1d7a5d06a6a0a737c79a4067863

Observation c9c63866-d224-4bf7-840f-dddfeabfb829 · outbound

This paper cites TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.574154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.574154Z digest=sha256:b8a76f97345b1a28258ae55bbaac32b453694006ec2352293054d2b4cb4c217c

Observation 106b0c91-6b38-467a-8957-6146547487af · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.579241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.579241Z digest=sha256:58bf8ae319ca11365c4a6791c55ef663cce2bf4d2f3e258ec5795eed0d142f89

Observation d009941a-08dc-4f66-b5a6-f5cbe72df6a1 · outbound

This paper cites Retrieval-VLA: Training-Free In-Context Adaptation for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Retrieval-VLA: Training-Free In-Context Adaptation for Vision-Language-Action Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.901566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.584585Z digest=sha256:6dce368ade5fdada9cda3a606f90f78790d9e45fa33ed6ce286c1b25816b97fd

Observation af0c0057-cb24-4949-a60b-f467895629d2 · outbound

This paper cites Retrieval-vla: Training-free in-context adaptation for vision-language-action models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models Retrieval-vla: Training-free in-context adaptation for vision-language-action models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.882434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.590092Z digest=sha256:7d8584a270d9a3645750a7dc9b999ec48c689faf1391dfee5cf56b879ee90daa

Observation 0f32b0c2-d391-4553-90e7-afdd546692c2 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.595363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.595363Z digest=sha256:4f8588913802ab65b7c08cabf00a3638ec6e8b3ea49fe6398624e6a4b1e5148d

Observation d1be27be-a1ef-4583-917c-99e785ed12d8 · outbound

This paper cites ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:39:41.864856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:39:40.600800Z digest=sha256:7f3afa42a190f0622d44606bc19885f5f1df68ee0782cf97c1ae8db4d969fc18

Observation 0f1c1164-5e90-4490-af96-a4da7be0e23c · outbound

This paper cites LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.605712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.605712Z digest=sha256:4a635a33a0083548c0f0a525d857f9a8d306ab3dfc44c7b2c093365794ca4528

Pith citing papers

No inbound Pith citation observations are available.