Pith. sign in

Paper Citation Record · LEDGER

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2507.01016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01016 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:07:26.130586Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:59:53.042021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.485742Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bb710e7-e1dd-4a4a-9df3-36d5f420dc2e · outbound

This paper cites GPT-4 Technical Report.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.122914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.122914Z digest=sha256:fb53a6f9d7a26299ace7c2cfc0258532996c1c2740304f6a9a2f451062b8cc29

Observation d2b560d4-2189-4451-aa56-58d7370ba347 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.189706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.189706Z digest=sha256:bbcc676ab1b3a67c18ef5f49bb35a072087cca0bf43cf298d70a560d2dc59ea9

Observation 8cb96dad-3e34-4e3b-be2a-3df86d9bd178 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Flamingo: A visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.458950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:22.275615Z digest=sha256:0ce1f1fbefd38588470cf63f92ff65dcd753d8ca551d74cb501a4c02089c121e

Observation bab2aed8-d8df-48d5-b6d8-f8f3ca6e8650 · outbound

This paper cites Minivla: A better vla with a smaller footprint.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Minivla: A better vla with a smaller footprint

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.333788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:22.363654Z digest=sha256:345591caa6e588ceb0f2ce4fa515b68f09b515324a94fff878d38f6b5347ca1e

Observation 9b752740-a23b-42b6-a8c4-1e9ce8e0b41b · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.451956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.451956Z digest=sha256:7870b5affa10fb640d1e53a2a8bdadd8e5fb29ca99148b49bd9bcfd977e37bab

Observation 6df64edd-df2d-43d0-b50b-7a9991989922 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.506991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.506991Z digest=sha256:05b88c908b85cad3706cc7fe70cf3d939fa47a1960c14b26a6a7e056c5ab3a20

Observation d43c688d-cb8a-4c24-815a-9141f9a3bea3 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.569402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.569402Z digest=sha256:0e50f5f66a0eda8d2d2a21ba6f47345e8a6de37653bfa55523cab18d32e16129

Observation 3a2658ac-b097-4a57-99f2-2d45dbf88426 · outbound

This paper cites Anyvlm: Unified vision-language model for any robot morphology.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Anyvlm: Unified vision-language model for any robot morphology

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.147454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:22.651487Z digest=sha256:21030e6470a211ef280581dec350d6533a8615e3eb0966705b554ad7ef09500d

Observation d4f6ece7-02f1-4998-ac13-45b3652590c6 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.706473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.706473Z digest=sha256:ed19529b83ed48b6853f8699c4c50a218fdbec3ae036e8d73bc0f5e1914e9430

Observation 3827fc8f-d211-4da5-b32d-9e48b3d2b97c · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.759775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.759775Z digest=sha256:af7d2ec0fc6f97b7f5ad33712e808b2f4f94081f387c8d2d62d68e6505b58b51

Observation 7457aed0-8bc8-4c67-8070-4a723c9fbbd1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffu- sion.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Diffusion policy: Visuomotor policy learning via action diffu- sion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.002060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:22.864369Z digest=sha256:293deeef471d5e3b14153d2f011cb3f3e966159a23c38e7545c340f91ca688cf

Observation 49ca3dd1-337e-4860-8823-47335f20c2c8 · outbound

This paper cites Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.903469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.903469Z digest=sha256:658080db7d08cc1e357b1832c4ba1deee84e1af4e286866ed5079f75f256842c

Observation f6a7de39-98df-4cdc-836c-6243442d69a2 · outbound

This paper cites Palm-e: an embod- ied multimodal language model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Palm-e: an embod- ied multimodal language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.918616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:22.959736Z digest=sha256:a979e6f0644cd8e43e849afe4d9c568fca99e234f8d9c66b658d0c0dac8cf5ad

Observation 97837d45-4a19-4c9f-a7d8-8072d999b00b · outbound

This paper cites Exploiting llm quantization.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Exploiting llm quantization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.833382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:23.014734Z digest=sha256:4201ea56f36c0afb5f6b2d7238292c249f77429f79b60113af2a10d6b4a8ff5d

Observation f2dbaa8c-661a-4e26-ba7c-87affb1f7875 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Taming transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.731198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:23.128367Z digest=sha256:442e75e5741997d8fedb97107d5f60014ced6517c7293663cbcb11a00d3880eb

Observation c5c4ba28-880a-40a5-b658-21c9594fb4d5 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.240619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.240619Z digest=sha256:ae593efd16895e8471f8c4620f142e13be5309c89467068a96d251af65fd075a

Observation 96be1bc3-b35f-431b-b1a5-c6304326c03f · outbound

This paper cites A new algorithm for data compression.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers A new algorithm for data compression

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.620501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:23.348202Z digest=sha256:04be3b10c2b47c46a01dc292e4d34f7a7186c81b35c11f08ee80b6094c15595e

Observation 743aa2af-c4ba-4d72-96f2-2b3cd9565335 · outbound

This paper cites Act3d: 3d feature field transformers for multi-task robotic manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Act3d: 3d feature field transformers for multi-task robotic manipulation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.531754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:23.457029Z digest=sha256:9f2a31d8364c48cf3c452cf0940e1175f788585430074d6ca41ba37ed26f5b57

Observation 85db7f8c-2df1-49c7-836d-3b681e7bcffb · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.545915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.545915Z digest=sha256:f39bf6ebc65ee582d946492705599b3633d5089c15ecbffd2f088f3c091b71fe

Observation b1115345-e57e-4ec5-a187-8b9c1ed1f47e · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Deep Reinforcement Learning in Parameterized Action Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.686753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.686753Z digest=sha256:265695e3d06c8dad3b230132ea6c79a29eccf796a323a18e46ccaed785c4cc24

Observation 8e9902a1-93fd-498e-b383-c9e4935b72b7 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.792753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.792753Z digest=sha256:f1f343399357087f9679982c22a9bc4c60ffb2b41d9d47ca8ae02ceb83e48e80

Observation 13dca03f-3454-44ce-8b2e-858468f84538 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Rlbench: The robot learning benchmark & learning environment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.442533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:23.874297Z digest=sha256:922a6f8c2a03999334a9cab03fb9401a338839c47e7658a20e3737794b7ea272

Observation 398ef956-5728-4acf-aeae-20f2aa4fcf33 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Pyramidal flow matching for efficient video generative modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.982056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.982056Z digest=sha256:3ccc99471bfd491923a3f30a3c2ec4d9abcc58c26763afd86554f2ba0f6540f9

Observation 60582574-bb70-4fc5-abd2-93782b851e18 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers OpenVLA: An Open-Source Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.072732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.072732Z digest=sha256:3f06302217635c6044625075dff449a3d34f84e31c82b6a3419fe649c7490ea7

Observation aa408865-f62b-4e67-8fa0-68500444dcf2 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.169056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.169056Z digest=sha256:06537872c6a196b63054f35fbf00f7aa11dc11c064caefb79cf0cda25ca8d70a

Observation ab9bdf39-151d-42fc-93a5-b7188949b3f1 · outbound

This paper cites Behavior Generation with Latent Actions.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Behavior Generation with Latent Actions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.226708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.226708Z digest=sha256:34b610f42317bb4c6524ad17ce2da22b85ee33ab9ea6e4552b390c98f33ff495

Observation 12ce03b6-eb68-4fdd-b832-d35b9d18ff2f · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.295662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.295662Z digest=sha256:1376a225533905c111b7b67c5572f36f69f463d6cfb5f4f95053f9ec294ba0a1

Observation ff2f85e6-844b-4225-8780-158474d501a5 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.365470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.365470Z digest=sha256:6cba6d3b529e7542e7869bde90290cb01ab18502a725de8b98a09aa28c30cd43

Observation c61f6e63-fbab-4d01-a09e-f871f4665d8d · outbound

This paper cites Language Models are Few-Shot Learners.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Language Models are Few-Shot Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.445625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.445625Z digest=sha256:17a6cf09dc0b3d0fb68eecdfb5fce076abbf88ccb9ee160f906fca53ee22e7d9

Observation 2dc1c629-12aa-4ffc-b9c5-24dd9705b76a · outbound

This paper cites Quest: Self-supervised skill abstractions for learning continuous control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Quest: Self-supervised skill abstractions for learning continuous control

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.322251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:24.514228Z digest=sha256:75ee3d035174115b6642c8165de12364819193821a65a50c34ec13b4882aba70

Observation 5beecf7a-5133-43f5-a2de-2873d75a10cf · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.596296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.596296Z digest=sha256:163a17af9fcb2f839f6e16aefeae87efb68159159d161d252b0b922b1b680646

Observation 2b13a5cb-123d-45f7-a071-a5d63ffe1527 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.226869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:24.653961Z digest=sha256:df6e4b031a09994959042d4990ef5cbb1dc477660b4f1bc0b1d9cb78b6ed7551

Observation 0c87a25f-e31b-4a1c-bcc9-b8cbf4155947 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.690665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.690665Z digest=sha256:88807610f22b15507836db9f0f2b1fc2c2c7fb5004f99ecc5858883f8d38dd11

Observation 49f8508b-3bab-41b6-9b32-6d4ba7b90ecb · outbound

This paper cites Language models are unsuper- vised multitask learners.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Language models are unsuper- vised multitask learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.133600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:24.774942Z digest=sha256:c60cc509cf15a8d5f017a079c7442377b992327f073fce40a5d19eac390760b0

Observation 3a536382-8929-4bb4-a7a0-e36e558a26ef · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Scalable Image Tokenization with Index Backpropagation Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.825964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.825964Z digest=sha256:7a6c2c8147526c1ddb29bd082b801ece0782fc7bb0a6bb3c5c5f8c93d2cf4b62

Observation 7366d8dd-f706-4cfe-b1e1-b8382b733708 · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.915904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.915904Z digest=sha256:b86d8539b1bf914ac8fca3ee375e475559fe5858c68a2389026c87bf2cc95c4c

Observation 4db8a6af-59f3-4159-a494-9b86e285c1fe · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.973912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.973912Z digest=sha256:ca883ec8ee897e40d945ec31b27b2ae439e015f42697cde26bbd574a1e17327b

Observation 8554926e-dcd0-4b2a-93ab-00b3c145a3de · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Octo: An Open-Source Generalist Robot Policy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.030945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.030945Z digest=sha256:bcdc020ea265fd602c855654b94fec03f7f77c2f686995050930bab18d8be5be

Observation d9d6a6d4-90fa-462c-9991-fda81792a0fb · outbound

This paper cites Neural discrete representation learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Neural discrete representation learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.105749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.105749Z digest=sha256:0483091658e53f4686ae3c610d96bda6a779325d4e3a04a6a41ae4e8804fd618

Observation 9c69885d-71c4-4ad6-b823-a8a55f78fdc6 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Chatgpt for robotics: Design principles and model abilities

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.044323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:25.151700Z digest=sha256:3400c3877c38d80b753e7b01eb2806bec960b83731c1aa9fb27b399943faaf16

Observation bd82c91d-c9d7-45fc-a6a3-c4788a484fdb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Any-point Trajectory Modeling for Policy Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.244285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.244285Z digest=sha256:02c7afe296f810b95edf6d9d6fa1560809a0876061b3673b8fae66ba320832d8

Observation fd872960-f77d-4bab-9005-d9e24eec506c · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.323052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.323052Z digest=sha256:863b379ce8b434a84934328c0d6a98482504235ce93e45691af8de8931d5ac5a

Observation 748b3e90-2048-49ad-a580-f499ac889134 · outbound

This paper cites Transferring Foundation Models for Generalizable Robotic Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Transferring Foundation Models for Generalizable Robotic Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.362504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.362504Z digest=sha256:115023de09ef8969bac1fefa94e37a5cf56bf58765f28213b1f3db48e08b7581

Observation da1bc365-02d9-4d08-a4b0-9b76864969d4 · outbound

This paper cites Spatiotemporal Predictive Pre-training for Robotic Motor Control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Spatiotemporal Predictive Pre-training for Robotic Motor Control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.446793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.446793Z digest=sha256:bf352ee3aceba5955f5e548f62baeedae04b925db71605f3ba707312ceb4ef95

Observation e1d8b0a8-5381-4bc9-bd51-22bf5b3618a4 · outbound

This paper cites Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.504405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.504405Z digest=sha256:23b98263143ccfec5e4a821a7c2f6c5616c5febd83b790a4b708149d6461677f

Observation 5a69ef97-ae9a-40d9-9334-e9c04c16722c · outbound

This paper cites Latent Action Pretraining from Videos.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Latent Action Pretraining from Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.536516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.536516Z digest=sha256:128dc87e54c415cb2aa7a1af111c293a9368845c7bf76f2c9d50c9e6c3903de6

Observation 1b81bb15-fd8f-409e-9e5f-108a76b3b3ba · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.595946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.595946Z digest=sha256:d420e3b7b9e51acb6f17e1cf80a2d199089988ebe59c7c33f79afb5b44dedff1

Observation a4a5ee28-9536-481e-b8cc-fca2f56b8fbf · outbound

This paper cites Soundstream: An end-to- end neural audio codec.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Soundstream: An end-to- end neural audio codec

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:26.923590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:25.664555Z digest=sha256:ea1c253c3805ddd43ed00accb7b8980fbd395217547e5db8f530587cb75be65f

Observation 296c81bb-282e-4713-b691-debf1cd58ac1 · outbound

This paper cites Moviedreamer: Hierarchical generation for coherent long visual sequence.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Moviedreamer: Hierarchical generation for coherent long visual sequence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.702842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.702842Z digest=sha256:921ace39302a969055d0dc55aadbce012fb3d1f60b5478a0360af15d2c45b28e

Observation 850984de-93ef-4ff3-ab41-3585a9bf2cea · outbound

This paper cites Diception: A generalist diffusion model for visual perceptual tasks.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Diception: A generalist diffusion model for visual perceptual tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.791210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.791210Z digest=sha256:1f2e50db552db7fa6d7900c7d427b07f33b9ba4ce08f838c7c67485ea2a7dda1

Observation 9b04ccf3-2d15-4542-855f-c4cece5049b7 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.850810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.850810Z digest=sha256:494a19d957b0d0854de2cfde004772ea180dbbd80bdcec677813a6c1d5628fde

Observation ef8e333e-3baf-4787-8619-ad432f15f447 · outbound

This paper cites VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.936596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.936596Z digest=sha256:8ae67d96fc943d9c459387597deb52f7e8c23e22c81a06b7401d9633c4190b0a

Observation 50648549-db24-47ed-9cd1-4360a95c363b · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.998032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.998032Z digest=sha256:bfde9bd763bba88efb1ca10ad09c9f8306ca6d36e25ee959325f763b6050d80e

Observation cb1aa5e6-43e7-4995-b644-82f7e86dbef5 · outbound

This paper cites Point cloud matters: Rethinking the impact of different observation spaces on robot learn- ing.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Point cloud matters: Rethinking the impact of different observation spaces on robot learn- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:26.828145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:07:26.049381Z digest=sha256:f193ee4fba8403aae773733c4cdac598cfc2cd77e5020278bd8206b8ccc86d5c

Observation d873b553-398d-4f19-8b8c-2d544b0d7e47 · outbound

This paper cites SPA: 3D Spatial-Awareness Enables Effective Embodied Representation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:26.130586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:26.130586Z digest=sha256:e6a7516f8a6f3ed916164545f07637ae1dbac913ce5891fd86d276974ca338d5

Pith citing papers

Observation 3a81d664-dbb7-49fc-8697-3ce1f15be229 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:53.042021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:53.042021Z digest=sha256:7e6ab62e293d1a457ab624dd587cf96f044bce8206dea49fe7b3c777643d6729

Observation 4f2fbe16-6b13-4c90-9544-c2ded5458320 · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.248500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:00139d18f64fab766782f84e52df142bd1fb41b2bcb1105c7b57f29d5ab52aaa

Observation 5a469e98-446a-4ff9-8737-5725a9922559 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:48.923259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:48.923259Z digest=sha256:44309b873d3ab1fe2bf9ee468e00b331fb6999a6ca49510532d5f28811660fc6

Observation d3391853-c917-4460-8253-b556712ed4f2 · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:40:08.704943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:53d8e73d32fa209bbea0f62b8c0c526e92cd4e0ba7459b61416a536451de2afa

Observation 2b267576-52e1-4a81-b10e-890416b67a99 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:05:10.051106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:b1e43b498b568d33cb8b8bbfa6fece20a02270984646503946693b96af8b3e2a

Observation 408a266b-fe64-4d6d-a97c-0ced98d15a1f · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.978308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.978308Z digest=sha256:ff52545cdd1adb531fe019876af1441b71a8122fd37a66105254785185a2bc55

Observation 97714748-41f3-43a9-b146-f3ad49458b1d · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:57:12.827282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:750378debed7995bba4b9f6cd6dae9c1c1abfbf3e3ef1f0fbca43572c480ab68

Observation 6a8e3e34-85de-437e-998b-533c131989ae · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.671785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:e3590eceee7e07410210742991ffb36b8427f0aef5fe96d19f4da73b40748ba8

Observation eb409383-3e7e-4ba8-9645-4bef12279702 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.666110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:26618480f7edbea8ecdb8be8e60f0bb3ea06c62d6271ff1c85e3ac89280cd59a

Observation d7a5fd1e-bd22-4e94-bdcd-8398898ae269 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.265667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:d9ebe218a1cf81d0f43f93e54148105fe957076f67ff7b1c216268b470f35a56

Observation 3df18ac4-53bf-401e-94a0-bec8be992c4c · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.911642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:5331a47efac5e7864f7cfc87066e78d9281b48215d33b603b0315e0108c9f4fb

Observation 29b1bc9b-80ac-436d-a26e-b6ec5eceff9d · inbound

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation cites this paper.

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.487289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:14:42.111428Z digest=sha256:49799152afc299d4fe657131929b9ee346c4593af0a931ec57708845cbe933f1

Observation 64cce292-40e0-4212-8839-749231c2e2d0 · inbound

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation cites this paper.

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T05:40:47.306935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:40:47.306935Z digest=sha256:776a0873c9528be5a1b60cab2cf847d36da4e98fc68595ea46e9bef59042dd8c

Observation f734bcda-218e-4a07-8473-f5a4d39067db · inbound

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots cites this paper.

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:01:26.608329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:01:26.608329Z digest=sha256:50955fa45a49ff2522ab0a8b74351f667b2c2717e2866499eedfce1bc79f44ee