Pith. sign in

Paper Citation Record · LEDGER

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2507.01016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01016 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:07:26.130586Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:59:53.042021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.485742Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bb710e7-e1dd-4a4a-9df3-36d5f420dc2e · outbound

This paper cites GPT-4 Technical Report.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.122914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.122914Z digest=sha256:a4609e2efb2d0e99ff62c26574e90c61e34be889a448f98ea7703f9dbb049a76

Observation d2b560d4-2189-4451-aa56-58d7370ba347 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.189706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.189706Z digest=sha256:db4b495d4127b3f2d6339abaafb9aedf50f43e6d68a23812c2659ecd62c17d2a

Observation 8cb96dad-3e34-4e3b-be2a-3df86d9bd178 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Flamingo: A visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.458950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:22.275615Z digest=sha256:9af96f05bef9f774883d0f6eb3a5ed58d20ad0e03be783e381db2d440d739f03

Observation bab2aed8-d8df-48d5-b6d8-f8f3ca6e8650 · outbound

This paper cites Minivla: A better vla with a smaller footprint.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Minivla: A better vla with a smaller footprint

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.333788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:22.363654Z digest=sha256:34788ca3acf77f666346f660bfb9b9fae915244d0b9854dadccc98fe591bb98e

Observation 9b752740-a23b-42b6-a8c4-1e9ce8e0b41b · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.451956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.451956Z digest=sha256:204810ae543598d19750be7e2bf03b96f044fa71abbf314d7869e590abb80d7c

Observation 6df64edd-df2d-43d0-b50b-7a9991989922 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.506991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.506991Z digest=sha256:84ae759adaf05352868f9ce2e9b91fea5725268c4fd1e67820015387b8350877

Observation d43c688d-cb8a-4c24-815a-9141f9a3bea3 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.569402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.569402Z digest=sha256:6a0d7cfcdaec29541fab2ed268c5c2da416d139ebdc78414def9fdc93ff83d80

Observation 3a2658ac-b097-4a57-99f2-2d45dbf88426 · outbound

This paper cites Anyvlm: Unified vision-language model for any robot morphology.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Anyvlm: Unified vision-language model for any robot morphology

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.147454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:22.651487Z digest=sha256:8d29bfdbc1b21113e37712a3aa8ef9a61a9fd409e40c2ff6789bd76b69c2b87d

Observation d4f6ece7-02f1-4998-ac13-45b3652590c6 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.706473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.706473Z digest=sha256:c11ee727637b5205926a0b9b32407434c91083e4732c20af1f1240e80ac58a7a

Observation 3827fc8f-d211-4da5-b32d-9e48b3d2b97c · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.759775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.759775Z digest=sha256:a6444e0970f6d2baf71f4ec1c2ec7b30f69b6ea7a1116f59b1e70288670b32fe

Observation 7457aed0-8bc8-4c67-8070-4a723c9fbbd1 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffu- sion.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Diffusion policy: Visuomotor policy learning via action diffu- sion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:28.002060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:22.864369Z digest=sha256:5446aab1b81a78a821e06afaf74672420e3cd0879ddd72fa604eda6a49ad8dfa

Observation 49ca3dd1-337e-4860-8823-47335f20c2c8 · outbound

This paper cites Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.903469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.903469Z digest=sha256:4cd6be36ec1ed8a6e3b51af770d71a9e43b2f6822553fa32f7e6b41eb2d9c5e5

Observation f6a7de39-98df-4cdc-836c-6243442d69a2 · outbound

This paper cites Palm-e: an embod- ied multimodal language model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Palm-e: an embod- ied multimodal language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.918616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:22.959736Z digest=sha256:48fcae10bc0035f86793af727d4ab95eddc3d78177255aaaa24a8a50f75209d7

Observation 97837d45-4a19-4c9f-a7d8-8072d999b00b · outbound

This paper cites Exploiting llm quantization.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Exploiting llm quantization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.833382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:23.014734Z digest=sha256:1315246508ad1c2e36d3a542b366c9aaa37cdc03ba830b72b9ec7e41663dfc00

Observation f2dbaa8c-661a-4e26-ba7c-87affb1f7875 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Taming transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.731198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:23.128367Z digest=sha256:d3a841d4dedd1116c923f168386f7a3feea77d592acebd6468762a1e7eb24107

Observation c5c4ba28-880a-40a5-b658-21c9594fb4d5 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.240619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.240619Z digest=sha256:1613ae079c01d0de375d2ec25b95213e78e725efd2f4ecc825c004188734867d

Observation 96be1bc3-b35f-431b-b1a5-c6304326c03f · outbound

This paper cites A new algorithm for data compression.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers A new algorithm for data compression

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.620501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:23.348202Z digest=sha256:0663c3d2660e331121d02977a51712bb4c65a62cfa1e2eb4131b9292a5934270

Observation 743aa2af-c4ba-4d72-96f2-2b3cd9565335 · outbound

This paper cites Act3d: 3d feature field transformers for multi-task robotic manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Act3d: 3d feature field transformers for multi-task robotic manipulation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.531754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:23.457029Z digest=sha256:74aa71c0452751f50d928b025df8ec6c7b0a710d27c353a497f1df4a2b90cbf8

Observation 85db7f8c-2df1-49c7-836d-3b681e7bcffb · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.545915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.545915Z digest=sha256:4d3f1db3f99994bae70b2f302ef0e5e820d0faa87bca3d2fc1c992c9e431845f

Observation b1115345-e57e-4ec5-a187-8b9c1ed1f47e · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Deep Reinforcement Learning in Parameterized Action Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.686753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.686753Z digest=sha256:075e948ed1f440cae354af9990da3ee589fea47433c0da8b66ae0e9fb7bd9551

Observation 8e9902a1-93fd-498e-b383-c9e4935b72b7 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.792753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.792753Z digest=sha256:71e7c565bd53772972884dbe7ea56879dbc1f5f63bec805d805256720766c1a8

Observation 13dca03f-3454-44ce-8b2e-858468f84538 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Rlbench: The robot learning benchmark & learning environment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.442533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:23.874297Z digest=sha256:168de742821a1297ca3e7a3a2c20ddc6d7fd6c03ebb8cdf7de1d4e9b0a95771a

Observation 398ef956-5728-4acf-aeae-20f2aa4fcf33 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Pyramidal flow matching for efficient video generative modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.982056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.982056Z digest=sha256:944cf9d2bb053cba5eeac28ba52662f03e3c19af745a85902cc046280cc28863

Observation 60582574-bb70-4fc5-abd2-93782b851e18 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers OpenVLA: An Open-Source Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.072732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.072732Z digest=sha256:2fc0a2f53c7a729363d8dc68072dd64a674d914d4c9daf20aa86a178beee90bf

Observation aa408865-f62b-4e67-8fa0-68500444dcf2 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.169056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.169056Z digest=sha256:7e9269aa31a38431303f4a5045c4c011fb829bad0ad7a9aa4760421150f7a82b

Observation ab9bdf39-151d-42fc-93a5-b7188949b3f1 · outbound

This paper cites Behavior Generation with Latent Actions.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Behavior Generation with Latent Actions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.226708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.226708Z digest=sha256:42bb23f3088795c9578d2c91f2408cb8b593bf13d0407e194a804b2ca699affa

Observation 12ce03b6-eb68-4fdd-b832-d35b9d18ff2f · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.295662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.295662Z digest=sha256:2c9f7422c5d9697a0404c5c293f9d240f6203ac6e1e481b771f9de49fc740b5d

Observation ff2f85e6-844b-4225-8780-158474d501a5 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.365470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.365470Z digest=sha256:37df900ba8a7a9983b696dc720453469e2b91e2e30f4ad13c1d51b9455c14cee

Observation c61f6e63-fbab-4d01-a09e-f871f4665d8d · outbound

This paper cites Language Models are Few-Shot Learners.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Language Models are Few-Shot Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.445625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.445625Z digest=sha256:1502ea9938a435e56bc7fb8bf18f65ccbaa8ae47686e927c5bed0647f7b34271

Observation 2dc1c629-12aa-4ffc-b9c5-24dd9705b76a · outbound

This paper cites Quest: Self-supervised skill abstractions for learning continuous control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Quest: Self-supervised skill abstractions for learning continuous control

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.322251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:24.514228Z digest=sha256:531f59ebe1fa155a54ac194c354dc38ff61742ade9fed4ab9f419a0ff9f4f47d

Observation 5beecf7a-5133-43f5-a2de-2873d75a10cf · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.596296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.596296Z digest=sha256:225b0d229642f070001596cf4156b169b3b0ba51a355a48a34f39f5d1c650e5f

Observation 2b13a5cb-123d-45f7-a071-a5d63ffe1527 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.226869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:24.653961Z digest=sha256:d780d6d45140db147c8c799cc3dbb5e854eaac690b41ff5e251bbd38d4d2c115

Observation 0c87a25f-e31b-4a1c-bcc9-b8cbf4155947 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.690665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.690665Z digest=sha256:a4796cfa0464eb4af3cbd1c25344182c528a84c3f72db15cddb3bd93a3962404

Observation 49f8508b-3bab-41b6-9b32-6d4ba7b90ecb · outbound

This paper cites Language models are unsuper- vised multitask learners.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Language models are unsuper- vised multitask learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.133600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:24.774942Z digest=sha256:31506af166e86fe01cb54bdec89e74fbf1026ffffd4d42f7e90b8df03a4978f4

Observation 3a536382-8929-4bb4-a7a0-e36e558a26ef · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Scalable Image Tokenization with Index Backpropagation Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.825964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.825964Z digest=sha256:e2bc2f2a35cdd841737742302b85a8467075499efa88ec60609ba06cc3a10e2c

Observation 7366d8dd-f706-4cfe-b1e1-b8382b733708 · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.915904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.915904Z digest=sha256:d5328913ba5581932acc32bc6444a95eae49c2b34b0611486b183502cc8ad17e

Observation 4db8a6af-59f3-4159-a494-9b86e285c1fe · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:24.973912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:24.973912Z digest=sha256:fedb3e53ae565999cf2a6d5f6cc4295d5a20b7c62303dfe875fc9f15b5b0763a

Observation 8554926e-dcd0-4b2a-93ab-00b3c145a3de · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Octo: An Open-Source Generalist Robot Policy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.030945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.030945Z digest=sha256:03678261396bd72893fcf39b5ffa2c73c8e308057efa1160c058dd273d9d9be1

Observation d9d6a6d4-90fa-462c-9991-fda81792a0fb · outbound

This paper cites Neural discrete representation learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Neural discrete representation learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.105749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.105749Z digest=sha256:9317386febdfa1774e0ca28c664243e2e466777225e720c238746f3f203c331d

Observation 9c69885d-71c4-4ad6-b823-a8a55f78fdc6 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Chatgpt for robotics: Design principles and model abilities

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:27.044323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:25.151700Z digest=sha256:3fbcb533c79610e4fa0f577f83ba7031279245e253adbb27574b258b8a17c27a

Observation bd82c91d-c9d7-45fc-a6a3-c4788a484fdb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Any-point Trajectory Modeling for Policy Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.244285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.244285Z digest=sha256:9885cf3306b7535a7d23ad2261b9f025c48bfc0f484c7211132622e203573c65

Observation fd872960-f77d-4bab-9005-d9e24eec506c · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.323052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.323052Z digest=sha256:300508a982869f61222dd522fdb9e22bc5871d2d21f5b79748d13634b72a7663

Observation 748b3e90-2048-49ad-a580-f499ac889134 · outbound

This paper cites Transferring Foundation Models for Generalizable Robotic Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Transferring Foundation Models for Generalizable Robotic Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.362504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.362504Z digest=sha256:c95baf07f5ef46c963f03c5b60b346173a2fecc8c5f67ba8b2b239c89df1b6aa

Observation da1bc365-02d9-4d08-a4b0-9b76864969d4 · outbound

This paper cites Spatiotemporal Predictive Pre-training for Robotic Motor Control.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Spatiotemporal Predictive Pre-training for Robotic Motor Control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.446793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.446793Z digest=sha256:252c01c0a7f2345686901f609fb4d61578befe5b015849b79b837776c136b0ac

Observation e1d8b0a8-5381-4bc9-bd51-22bf5b3618a4 · outbound

This paper cites Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.504405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.504405Z digest=sha256:112f5e58a2d28e2eb3ad42052da8c064f533757756fd75bd4ec51b0f533290fb

Observation 5a69ef97-ae9a-40d9-9334-e9c04c16722c · outbound

This paper cites Latent Action Pretraining from Videos.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Latent Action Pretraining from Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.536516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.536516Z digest=sha256:915d7562de36a2b6ac6073750e66e0a9a7f9e7769a0922d59fe2937da755dca6

Observation 1b81bb15-fd8f-409e-9e5f-108a76b3b3ba · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.595946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.595946Z digest=sha256:0ce2bc2d69d911022711ad4d698670e128f8729559b518e3ed5a7d1e4314e3b1

Observation a4a5ee28-9536-481e-b8cc-fca2f56b8fbf · outbound

This paper cites Soundstream: An end-to- end neural audio codec.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Soundstream: An end-to- end neural audio codec

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:26.923590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:25.664555Z digest=sha256:75a3fb8c8efbff7c7246f60a40ad8581a0c83dbddc41c58a4373a5200f22f533

Observation 296c81bb-282e-4713-b691-debf1cd58ac1 · outbound

This paper cites Moviedreamer: Hierarchical generation for coherent long visual sequence.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Moviedreamer: Hierarchical generation for coherent long visual sequence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.702842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.702842Z digest=sha256:08c5af909a9595e52ff46fada51f9c1719c123f93724e862742887a0876ad300

Observation 850984de-93ef-4ff3-ab41-3585a9bf2cea · outbound

This paper cites Diception: A generalist diffusion model for visual perceptual tasks.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Diception: A generalist diffusion model for visual perceptual tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.791210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.791210Z digest=sha256:04311d9966feeda1aab77bac222247162634c442a9885d365518c90a57c8b5ef

Observation 9b04ccf3-2d15-4542-855f-c4cece5049b7 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.850810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.850810Z digest=sha256:6b0792410b6749e45cfdaba8ce4ab48797e683497aa2a1f8490bc39aec7ae1ab

Observation ef8e333e-3baf-4787-8619-ad432f15f447 · outbound

This paper cites VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.936596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.936596Z digest=sha256:d37ccd7ff351bf55121975b49dd7dfc5800f1b549bf486448040d426f14d7e40

Observation 50648549-db24-47ed-9cd1-4360a95c363b · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.998032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.998032Z digest=sha256:d24d823239ac5462d9e594b2d28b202555cc720eeeaa1af27290db46fe8c5b6a

Observation cb1aa5e6-43e7-4995-b644-82f7e86dbef5 · outbound

This paper cites Point cloud matters: Rethinking the impact of different observation spaces on robot learn- ing.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Point cloud matters: Rethinking the impact of different observation spaces on robot learn- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:07:26.828145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:07:26.049381Z digest=sha256:7d67a7eea729564b05e1c0becc8fe302fd31d313e65132c9b98aaf255b0a3f0c

Observation d873b553-398d-4f19-8b8c-2d544b0d7e47 · outbound

This paper cites SPA: 3D Spatial-Awareness Enables Effective Embodied Representation.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:26.130586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:26.130586Z digest=sha256:98017233436ea8923bf06b740f1fdef0b64b063df24f4c3e46543ccaac0b1728

Pith citing papers

Observation 3a81d664-dbb7-49fc-8697-3ce1f15be229 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:53.042021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:53.042021Z digest=sha256:7ea958736159441892643e815a73b9ba5c26d92d5fc189cc30c91087391fb8c1

Observation 4f2fbe16-6b13-4c90-9544-c2ded5458320 · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.248500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:e1994ce6da9da7104e1a3463c56045974f1503629d4e0104230d5651300c9e19

Observation 5a469e98-446a-4ff9-8737-5725a9922559 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:48.923259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:48.923259Z digest=sha256:21f535a9f6ae731159263288a5b98fabe2db34d5f39854d76457defaff18d303

Observation d3391853-c917-4460-8253-b556712ed4f2 · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:40:08.704943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:dfd395110a8e9e385bb467f28a9b229c441c30c5d34311e01f17263f798fee6f

Observation 2b267576-52e1-4a81-b10e-890416b67a99 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:05:10.051106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:8c6c705a01a1c890625bc4d32b7ae9af0d3de21f62ee6c576c7d3e6d043dbb75

Observation 408a266b-fe64-4d6d-a97c-0ced98d15a1f · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:14.978308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:14.978308Z digest=sha256:a5c00770e2ff126fc8d7a5199bc2823181df5cbd8bfbe7a53a666da0ed192cb6

Observation 97714748-41f3-43a9-b146-f3ad49458b1d · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:57:12.827282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:9561d714800d7dd2a4a7333bdda0a4d3d9b7c9943df6beae8ef653ac27636264

Observation 6a8e3e34-85de-437e-998b-533c131989ae · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.671785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:343a9dbb31e79d5fb0683f7fc5e60f13b0e1fc74806bca41f70e322a693f69e2

Observation eb409383-3e7e-4ba8-9645-4bef12279702 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.666110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:fa39c19c9046619cacf5ed30111622c1ac5dadb154493e7ac0c7fd040ba81c65

Observation d7a5fd1e-bd22-4e94-bdcd-8398898ae269 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.265667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:571827915c3e0d45d9a8eaf95e35fd136577b130e0eab76aaeb932089ddc3390

Observation 3df18ac4-53bf-401e-94a0-bec8be992c4c · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.911642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:9119f825541dcb7900290494b2d33eabeaea2cb3d00981c069cc15366106a29d

Observation 29b1bc9b-80ac-436d-a26e-b6ec5eceff9d · inbound

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation cites this paper.

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.487289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T10:14:42.111428Z digest=sha256:ef14ab7a24cd3c6f1a2662cb19c19a59ac6572f151b1e45f6d0800715f9c4257

Observation 64cce292-40e0-4212-8839-749231c2e2d0 · inbound

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation cites this paper.

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T05:40:47.306935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:40:47.306935Z digest=sha256:2fa156b67d74dd840b610c61bebc4ab88a4132dceabe958e2255dbf3b882bc80

Observation f734bcda-218e-4a07-8473-f5a4d39067db · inbound

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots cites this paper.

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:01:26.608329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:01:26.608329Z digest=sha256:e67460b83ed8eeecee04c91bccb2e07ceeb3c3fe1c0b2d42024d4aa647982738