Pith. sign in

Paper Citation Record · LEDGER

VLANeXt: Recipes for Building Strong VLA Models

As of 3 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 6 inbound Pith citation observations for arXiv:2602.18532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.18532 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T12:58:30.777235Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T22:39:15.211339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T14:09:53.718432Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact34
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2529d4c-c5b3-4386-af60-45f20360ac71 · outbound

This paper cites Qwen3-VL Technical Report.

VLANeXt: Recipes for Building Strong VLA Models Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.843841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:cfb0b7c86e5f445e51b0849c89e1583f4f120430689399783f73c36d048d3213

Observation 7a76bb13-8fbc-4c1f-a856-83b19701ba6b · outbound

This paper cites 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks.

VLANeXt: Recipes for Building Strong VLA Models 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.938258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:068fa6e45043b9f99a43d6912c6ba5b37a9442ce4ecf7320945176ca0ce7a8e1

Observation ea65ca17-e5fe-4943-9cd2-aad2b1be43dc · outbound

This paper cites Motus: A Unified Latent Action World Model.

VLANeXt: Recipes for Building Strong VLA Models Motus: A Unified Latent Action World Model

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.853621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:7e782198c7f0a40090f801e3d9a37393fb75afbaecdfd0ff71ad71827265f31f

Observation 91a89579-96fe-4a66-bd99-578a2e942d9b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VLANeXt: Recipes for Building Strong VLA Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.829471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:5d7a60f8a715fe830ae5dc82742adb86daee682b3d6bd33e07b75e5aae4743d3

Observation 6a830836-c062-455f-9806-e218bc442433 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

VLANeXt: Recipes for Building Strong VLA Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.800641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:8b3193a7d55ad3e02d11c09b3223be1a06b6c9c7fa5561a34e874e4a852fb7b8

Observation 3403115c-1040-4070-a9d6-984910f7010d · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

VLANeXt: Recipes for Building Strong VLA Models RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:36.254972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:37ca0ededf4427fab57e70b9626a1c4f7629567677ab316d11e98e5c221b40b5

Observation 07a9d8cc-c5b5-4d83-a626-80ea2d0ad0fa · outbound

This paper cites Combatvla: An efficient vision-language-action model for combat tasks in 3d action role-playing games.arXiv preprint arXiv:2503.09527, 2025a.

VLANeXt: Recipes for Building Strong VLA Models Combatvla: An efficient vision-language-action model for combat tasks in 3d action role-playing games.arXiv preprint arXiv:2503.09527, 2025a

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.917410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:657263bc8d04c9c2929c21940dbd007bf01cd7bdadc888ad0bcddad9b3baebdf

Observation 3a721f1f-b528-43cd-9b49-b910113cc8c1 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

VLANeXt: Recipes for Building Strong VLA Models Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.819893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:48c13c72920ca17ecd4199b97909ea7e730beafa9ce39fd7d1b1b1c68e3c1dd5

Observation 2843b0aa-9727-4034-a106-5d3bff1f88ab · outbound

This paper cites Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration.

VLANeXt: Recipes for Building Strong VLA Models Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.825098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:5db6c0c7aa83a13d484139e2f0e0296918726f8669ab8ec95571154d164c4dd1

Observation 8b3cf350-377f-4801-b242-20b4adf402eb · outbound

This paper cites Srpo: Self-referential policy optimization for vision-language-action models.

VLANeXt: Recipes for Building Strong VLA Models Srpo: Self-referential policy optimization for vision-language-action models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.943727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:7fd26b7e97a9e493fc54b0538d016f2f5d988421151b93546aa320312db3b16e

Observation f82c2aad-f36e-4573-be43-5453fa722962 · outbound

This paper cites Vla-0: Building state-of-the-art vlas with zero modification.

VLANeXt: Recipes for Building Strong VLA Models Vla-0: Building state-of-the-art vlas with zero modification

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.948808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:9377540c3b9b521f8f1aaa9b6dea84e14ddc5b08e806657b64f773e61e690092

Observation b207c2ce-82e7-45ea-878d-caf95d688861 · outbound

This paper cites The Llama 3 Herd of Models.

VLANeXt: Recipes for Building Strong VLA Models The Llama 3 Herd of Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.877521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:468ab6e64af86466a0498c6b009e8a3f372fdbba6b84e365592a9df942f8cbf7

Observation eaa56b9f-aad6-4800-98a0-4a53e04b47e6 · outbound

This paper cites Vla-reasoner: Empowering vision-language-action models with reasoning via online monte carlo tree search.

VLANeXt: Recipes for Building Strong VLA Models Vla-reasoner: Empowering vision-language-action models with reasoning via online monte carlo tree search

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.902948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:7075c478b7ea6638cee13cba38d21d486e59ea1efebd2efcef4ec4620b07c72d

Observation f0cc7abe-4ecd-4a84-8899-020d74d74d93 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

VLANeXt: Recipes for Building Strong VLA Models Training Large Language Models to Reason in a Continuous Latent Space

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.872549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:0822157ef6d2876298ce2c19d12541819983284520ae789caa11f1b4af48e2dd

Observation 14e156da-bc7c-4d62-92c7-27c5c93c74ea · outbound

This paper cites ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning.

VLANeXt: Recipes for Building Strong VLA Models ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.922795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:17b54f787a9d084b3405fd255e681a835d8fce95b5bfd64cfede97b3194396fe

Observation 562aa024-58df-464a-823f-066859c496f7 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

VLANeXt: Recipes for Building Strong VLA Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.958097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:bc5235e393d1344599e624b449d4f3466d87fd696aa810bdac18f070b3ccade3

Observation fb9fa478-13d0-4ec9-b9ca-ca2f7c7e2e85 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

VLANeXt: Recipes for Building Strong VLA Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.863649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:51006d22e74887100184d13d13b9a77ad78144666ddfd5eece740be6e32aa0c4

Observation 7feaa7fa-d48d-45e3-9943-bdf0ef59ee9e · outbound

This paper cites Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414.

VLANeXt: Recipes for Building Strong VLA Models Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.953953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:1030f9ea54b96d529c247fac536d99764d444ae894eadb70d2638bea49f317b3

Observation a2995e19-b286-431b-8389-ed660c291555 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VLANeXt: Recipes for Building Strong VLA Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.983615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:be159ffcbee53fcf3e831c0f55bab2d5e48345da1b38d549be70e48ea98ee7eb

Observation 553266e8-9b48-40fd-9460-5071d629a452 · outbound

This paper cites Adapt Your Body: Mitigating Proprioception Shifts in Imitation Learning.

VLANeXt: Recipes for Building Strong VLA Models Adapt Your Body: Mitigating Proprioception Shifts in Imitation Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.927300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:41d98ee139e057aa4908cf167bed7bd01a849dd577dd32d6d70663fb6b6cacf6

Observation 8624a883-38a4-4a8c-a3dc-56efec558575 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

VLANeXt: Recipes for Building Strong VLA Models MolmoAct: Action Reasoning Models that can Reason in Space

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.933214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:7fe3caa5017ab1b8bc0d780cb1b232a94b90863ea8de3a989c35198f6af2f571

Observation 340d6e3b-10d0-4dff-936f-89ffa78a0dec · outbound

This paper cites CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025.

VLANeXt: Recipes for Building Strong VLA Models CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.788929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:3aa2be933ed7062debe55f7dc63489912cfc51abce1e70c3a377cd72e9e718e5

Observation a004a340-b210-4fa3-a532-5fd319d6f2b5 · outbound

This paper cites Mm-act: Learn from multimodal parallel generation to act.arXiv preprint arXiv:2512.00975.

VLANeXt: Recipes for Building Strong VLA Models Mm-act: Learn from multimodal parallel generation to act.arXiv preprint arXiv:2512.00975

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.897990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:0888fff31540b602f27c28646c2b7d7b8fdba32ae2d6e738a013a62e62c6f75c

Observation bb374873-0a04-4ebc-b858-0cd97984f424 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

VLANeXt: Recipes for Building Strong VLA Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.810733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:45812b3cf2fd1c25120e870dc43e664345f74382d41404b26ed5adf4c57414cf

Observation 10ea45f0-60d8-4f14-b5c3-0a5129e41dd5 · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

VLANeXt: Recipes for Building Strong VLA Models F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.839448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:53aa41d652c37ca408808a40138a66e5099204c4a184f219c4dc5b4ee43dc18a

Observation 81c2b160-50b1-46f6-adb6-5307249002d5 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

VLANeXt: Recipes for Building Strong VLA Models A Survey on Vision-Language-Action Models for Embodied AI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.966662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:bf60efaa2c9d0091d75f38f8f44a0880d81f6a81aea2762b93a8a7ece7aa2eb8

Observation 7f42c612-968c-4900-abd2-b075f39f4f38 · outbound

This paper cites Transfer between Modalities with MetaQueries.

VLANeXt: Recipes for Building Strong VLA Models Transfer between Modalities with MetaQueries

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.815233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:dac2fa047d09de977b63c2378e0ed08df401a30092fffd966dede7c9052d9254

Observation c2198c2b-e888-4d1f-b497-b45c68da367f · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VLANeXt: Recipes for Building Strong VLA Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.805192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:f0a4539f749882d970d5e708403b463aa2bda5c7ae549fff5fd835f1f434ea09

Observation 5e275428-fc9b-4194-94da-00872201c09d · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VLANeXt: Recipes for Building Strong VLA Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.974891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:41e2841cb7485d077270ea04ea4a71146ff5b531c2ea037e3e64c00c50abdd07

Observation 3e623b8b-dfec-4782-bb0b-0e1bece95321 · outbound

This paper cites FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies.

VLANeXt: Recipes for Building Strong VLA Models FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.834452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:73475d3da8d2da7fe6446be027db7581dd03af19e3a2a9b28f1830254b6ba77b

Observation e4622c13-61b2-431f-a2fb-aa6d7df9ab69 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

VLANeXt: Recipes for Building Strong VLA Models MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.882340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:171b9c491f9f2e04e08a30b92ca64838136a1025fc289aee76aee052283f349a

Observation f83a9739-9b9f-4adf-ae36-042adf8aab82 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

VLANeXt: Recipes for Building Strong VLA Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.892667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:643f74efd729a52b24f4818fb026b6b198641f58281a1383aa47ddeb121718cd

Observation fee226ad-4cf7-4ec4-8e05-8746dc4aae4d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

VLANeXt: Recipes for Building Strong VLA Models Gemini Robotics: Bringing AI into the Physical World

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.848903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:b994508d07b25b2aee39d2ec1216ce6863a0732ef481b94a987c633697c8a612

Observation bb42da93-9506-4e30-bc3c-1fa1ea1ba714 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

VLANeXt: Recipes for Building Strong VLA Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.970523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:669ad93d86cf408e13c67f1be502bb390ab959861cf64fd569ee6fe4b139ad79

Observation 2d520977-76dd-4e28-b380-d883e3d437b3 · outbound

This paper cites End-to-end Listen, Look, Speak and Act.

VLANeXt: Recipes for Building Strong VLA Models End-to-end Listen, Look, Speak and Act

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.979464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:e49e098b3bbb3a367cc2e7409992eed7a6713fbad53c5fcc31ba9df87f3a8cd5

Observation 71c88fad-a9b5-4a7f-b2be-7ada6c9808c1 · outbound

This paper cites World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training.

VLANeXt: Recipes for Building Strong VLA Models World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.868211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:8cdff27473e5d52ad6fd5982cc70bafda0efd6c2aa08f677f6469deead799ed6

Observation b231477a-b9f9-488e-a402-c4690cd96367 · outbound

This paper cites 4d-vla: Spatiotemporal vision-language-action pretraining with cross-scene calibration.arXiv preprint arXiv:2506.22242, 2025a.

VLANeXt: Recipes for Building Strong VLA Models 4d-vla: Spatiotemporal vision-language-action pretraining with cross-scene calibration.arXiv preprint arXiv:2506.22242, 2025a

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.795112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:aec2d357c835a71457e00a63e62190b09f3471189cf42c9fa349e89fabf549e0

Observation bf431467-507b-4f2d-8d10-0212c85a8db3 · outbound

This paper cites Dreamvla: a vision-language-action model dreamed with comprehen- sive world knowledge.

VLANeXt: Recipes for Building Strong VLA Models Dreamvla: a vision-language-action model dreamed with comprehen- sive world knowledge

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.962649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:55e5502b2510811a9c329d274748dac9911beedfefaff2d15a794531834203a3

Observation 563fcdf1-8af8-4374-a347-18814a6f7153 · outbound

This paper cites Flowvla: Visual chain of thought-based motion reason- ing for vision-language-action models.arXiv preprint arXiv:2508.18269.

VLANeXt: Recipes for Building Strong VLA Models Flowvla: Visual chain of thought-based motion reason- ing for vision-language-action models.arXiv preprint arXiv:2508.18269

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.858979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:9235e5fb07553ae9edb50fe2eead1a9bcf522f1b5a4156060b6969008318e055

Observation 65c5a765-12b4-45d2-82a8-53d1723e8c90 · outbound

This paper cites More Experimental Results A.1.

VLANeXt: Recipes for Building Strong VLA Models More Experimental Results A.1

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-05-21T13:00:10.452298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:992840e962dc86be963becb47738fd131aa5a2c8d46012c58fb5d8fa8ebb3601

Observation bc88e369-620b-4fa8-a040-c67ca39c56a0 · outbound

This paper cites primordial soup.

VLANeXt: Recipes for Building Strong VLA Models primordial soup

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:00:10.455299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:ab4c41eeb0417b589a9df6920e6b187fc6cc410d082ee1b125816dde25220260

Pith citing papers

Observation c136c42d-0514-483d-8f53-d9fc1928c427 · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts VLANeXt: Recipes for Building Strong VLA Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:11.439936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T09:11:21.715023Z digest=sha256:0529ac2b6f5a7f361027d1b88ed00d66107fa66604357ba875949d9c48466ff2

Observation 54fbf5b6-c48d-4a25-815b-fedfe200d0ce · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts VLANeXt: Recipes for Building Strong VLA Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:11.439936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T01:05:07.509291Z digest=sha256:761250ffb8b91662c651d1b3b3f72a9cd4c45657b2aa1dcaa1359b6d0d2693af

Observation 102eb8f3-cd5d-49f3-a700-c087586b9b1f · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR VLANeXt: Recipes for Building Strong VLA Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:11.439936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:7d4b61107036246be498336f9116faf639ed468c4e1b01dbaaf86943f1d632ca

Observation 20be91d9-bff0-49eb-a1b7-848a0b41da9a · inbound

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring cites this paper.

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring VLANeXt: Recipes for Building Strong VLA Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.568474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:39:15.211339Z digest=sha256:f497eb7ffb6f57aaf4f760bb603095af5563d8e22027f827a9af38e1f10ee48d

Observation 94d4f7ab-fae3-4fa6-b308-181e2d207902 · inbound

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space cites this paper.

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space VLANeXt: Recipes for Building Strong VLA Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:58:58.543012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T00:57:58.312567Z digest=sha256:0afc69527701bbf801a23607490c455a462f48925b0760ef6f1133b8eb48fb3e

Observation 5e853214-c132-47ed-87d7-1375423f15f7 · inbound

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection cites this paper.

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection VLANeXt: Recipes for Building Strong VLA Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:09:53.719676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-26T04:25:18.870217Z digest=sha256:26c543f6534c0c7b2fe5093bcd7fef555ceee3befa97f6fabef3971afe6c5d22