Pith. sign in

Paper Citation Record · LEDGER

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 19 inbound Pith citation observations for arXiv:2505.21906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21906 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:04.383484Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:46.040172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.316884Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d21ea7c-2866-435c-973f-a3dd8583ba81 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:56.965076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:56.965076Z digest=sha256:5010f8e3ec855077d7b27f087de3c8c0c792f88c0fb5e909aaf3a3ec4f2bc728

Observation 7d0d4bab-577a-4f16-9761-867515298871 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.008943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.008943Z digest=sha256:9ff234a84db6fcba0467a2a96a87ba458ae923b1b177a58f651cd1d0975dc713

Observation 2eb6be59-bc76-4caf-92a2-b347b7b136a4 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.109866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.109866Z digest=sha256:9bcc0cf844fdff1ff7214408f780af8fdeaba6dd17568beb744ede9264ed6e7e

Observation b7ff553a-31a7-429b-a864-bd232afe9cda · outbound

This paper cites ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.204158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.204158Z digest=sha256:3fc7ab7fdd41f47f48a55c851df59bb3798622fcdd5c53b5305423809193643f

Observation c8acab03-d2af-4d4d-9344-82a6bb1694e4 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.263110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.263110Z digest=sha256:6d3d5c35acb14bbd3df6b4b37f09ab73692666954a909f4beefe93fbcbe0b7c0

Observation 3e18183a-1248-4276-82f8-ec2f5c3291ae · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.337520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.337520Z digest=sha256:ab22b9d1274d15b4225ce2ada3e1ce37ffd8b9249fd854e82e9e74d319c9dc59

Observation 226371c4-8467-4e3d-a7c1-017f7d6f1c8c · outbound

This paper cites ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.438562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.438562Z digest=sha256:8866ef9d6615ffb3905ed6e35564365652d04a03ae63d07c08c503bca39e969c

Observation e2c7eacb-e3b7-463d-8742-a214e14fb596 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965, 2023.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.561396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.561396Z digest=sha256:c0090e7f290a7cb427c5518ee357baae33a82dc438bdcc0b828cba6e7d7b9171

Observation a6947f35-87a2-46e3-a53d-f8d9265340f5 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Reasoning Models Can Be Effective Without Thinking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.600571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.600571Z digest=sha256:618a449a3e67d6184413425d86b648a4dbf3adee9ab056b8a63b0df3d80a24e9

Observation 33193132-135f-45ea-9ddc-72f63e163e57 · outbound

This paper cites Openvla: An open-source vision-language-action model.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Openvla: An open-source vision-language-action model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:08.242054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:57.679874Z digest=sha256:00351322eccc622ea5e659dc84346a5b1aa9e3b1a76625a603735b28c2436fd1

Observation 55827294-2694-4182-b4dd-68de7dac08c1 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.744476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.744476Z digest=sha256:b9b772386d847d83a36579a8c287c09ff923681ceb920ee1389749c1413058e4

Observation 548aeeac-794b-4726-977f-088f427409c4 · outbound

This paper cites Visual reinforcement learning with self-supervised 3d representations.IEEE Robotics and Automation Letters, 8(5):2890–2897, 2023.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Visual reinforcement learning with self-supervised 3d representations.IEEE Robotics and Automation Letters, 8(5):2890–2897, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:07.416228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:57.799763Z digest=sha256:fba11f466525fff29be3fd628e88e97e2a2fc4674b3a824868782ecd3788e0b8

Observation 82d8ebb0-3fe7-4a9c-8a86-7b1672f396dc · outbound

This paper cites Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.876058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.876058Z digest=sha256:d4f6e7eedaa5eae6e49cd58182e4ebcb7276dc22f5d96b6813792cf00f593a10

Observation 746f057e-c8de-4848-a8d1-5f557a1b7376 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.945809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.945809Z digest=sha256:e60c6eba7cb4a28e925a9b8a8a8ca9004746d22b35bba9a05db33d9130a32763

Observation 21d30c9c-8d72-4997-8f6e-6fced690ce5d · outbound

This paper cites Any2policy: Learning visuomotor policy with any-modality.Advances in Neural Information Processing Systems, 37:133518–133540, 2024.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Any2policy: Learning visuomotor policy with any-modality.Advances in Neural Information Processing Systems, 37:133518–133540, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:07.045644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:58.024745Z digest=sha256:9b43fc6cb7720c8250916fb002574aa13e603b3bfa4f5740396cc4eff546e884

Observation 0fc88331-e5bd-4b89-86b6-6607624c4309 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Any-point Trajectory Modeling for Policy Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.104075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.104075Z digest=sha256:2b81085d157e983466223a05887cc5bd14cd2ff783a6be987f050621044f7be1

Observation 695fd35d-8c9e-4a2a-9aff-31d95383513d · outbound

This paper cites Retrieval-augmented embodied agents.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Retrieval-augmented embodied agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.168750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.168750Z digest=sha256:c67d7928fd1224d5f2445a7e17afbe950f472dedb76d56983d75685ebd7a2ced

Observation 52a0c22d-9591-43c9-b18b-14c2a596b3f8 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RT-1: Robotics Transformer for Real-World Control at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.193106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.193106Z digest=sha256:a0b30db670ab3af9fce1664190383e9c5e5fc130c8bcf4bb1837de3059a2d1d2

Observation 91f2e76f-3719-4539-ae8b-4878063b1871 · outbound

This paper cites Rt-affordance: Affordances are versatile intermediate representations for robot manipulation, 2024.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Rt-affordance: Affordances are versatile intermediate representations for robot manipulation, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:06.837973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:58.264752Z digest=sha256:72f43978a72736ad55a9a3fb724cfd7fec6ef819ca973437f7a1b912398a451b

Observation a6aaff8d-86cc-47f4-928e-d0788e7ca553 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.334777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.334777Z digest=sha256:f13e6323fae447618f679298035d525ce055d54fd9c5b7038e57545937c9983d

Observation 7dab7c23-ac78-4632-82d9-6c8438b81bed · outbound

This paper cites Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.416037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.416037Z digest=sha256:070f810eb4e03bf68eb5a0c2261d99f95798daaae1297846239b5b32233bbfe8

Observation b8c7d9c8-7e39-4594-be27-af871733f963 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.459946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.459946Z digest=sha256:8e58b605350738bef8a2e3f25ddf30f0cc3a224e7eac2b4d0c3ad34c915e1ff1

Observation adcd24dd-49c7-46ca-b096-e49bb0e0f592 · outbound

This paper cites Mail: Improving imitation learning with selective state space models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Mail: Improving imitation learning with selective state space models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:06.587781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:58.524746Z digest=sha256:6688de30e2350913ee5139601ac2570a9a11843af2196789078a63addfe862d3

Observation a59c795f-e8ad-4785-b17f-5edc3faddf78 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.704747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.704747Z digest=sha256:fb9210d0eed2d15738dbe4a69a7ca48530462811ef5f64a4bff24101f5413a09

Observation 7b8e407a-0420-49b2-a79f-b764eee9cfff · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.804742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.804742Z digest=sha256:71712023cf476e72fd26a8a7a68a70f8bfdcea1eeb893e20fc66cc21a468fd67

Observation 6abf7fac-59dc-445e-9a07-4e67f2b4dee9 · outbound

This paper cites An embodied generalist agent in 3d world.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge An embodied generalist agent in 3d world

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:06.454384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:58.854746Z digest=sha256:5424c9ddc537958916e11a1a7e50b40c2b7d3bbcf61e3a2ed711e5983d2ad29b

Observation 81db2c8e-c22d-4e86-922d-84ed3e3cf046 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.919499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.919499Z digest=sha256:1155e20290312f99c29ca131d9d36370680fda1fc1b88c5374b5351d052584f1

Observation 36c10730-9e37-41bf-b368-2dfca4b73650 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.994292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.994292Z digest=sha256:9cdb11bebc03cf8a78b5764b0215ef8a37d90c307b6559b2e55bb18fb16c0a04

Observation b796f4bb-a1f0-4644-a7fa-f65ecb5fcd9a · outbound

This paper cites π0: A vision-language-action flow model for general robot control, 2024.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge π0: A vision-language-action flow model for general robot control, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.088384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.088384Z digest=sha256:4509d2ac60fa2a360f186b7d8e1078b768ea6e6b2193af5eed8b0051a807b559

Observation ff4ebc0e-ca74-4e34-8d29-d2164948b1de · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge OpenVLA: An Open-Source Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.190957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.190957Z digest=sha256:a5123bf39fe26c0e0f43ea86d10dd17a07d5295f6c268b0929fd1f5db9eb2eba

Observation 1c4177d1-3793-4444-824e-03d8bf31e882 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.326038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.326038Z digest=sha256:70827cbb93579921603d3a2ef4b9fa1192da54b3c716f74d83f41eba4674465d

Observation a1ddc19d-7136-4002-968a-9f7bb5a87959 · outbound

This paper cites Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.434850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.434850Z digest=sha256:3085ecf2ab1a5d19afa84af01ea8ecd40031cbd8410f4aa043b48a8ecd2b1190

Observation dfc82a78-bea8-472f-b61d-6ecc9759d7a4 · outbound

This paper cites Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.481710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.481710Z digest=sha256:4b61ee2433cb585114f1f8f22d05435a154a8f20097ea5ac6a646fa6de9c2423

Observation c055702b-42d5-45a5-8da6-aebad243287c · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Training Diffusion Models with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.546972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.546972Z digest=sha256:a2ecb19929fae6b0c31d63c3c9771a8281d3807f8889d5d7e53ebaee3a42603f

Observation e7560aca-d665-465d-8b2a-f544313c6dcc · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.631774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.631774Z digest=sha256:e152cad8367bb5e8d3d8260bf4ea2ffff2dcaedd83fdb938220f87d42d0783de

Observation be950bdd-6c22-4ced-8164-733187ea9338 · outbound

This paper cites The Ingredients for Robotic Diffusion Transformers.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge The Ingredients for Robotic Diffusion Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:59.736105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:59.736105Z digest=sha256:92cc5b158504d4cb83731da54931b82c4f6e2f8d617506b6144e078ff171c092

Observation e91f4925-a14b-4183-be26-c6967b0680d2 · outbound

This paper cites Data scaling laws in imitation learning for robotic manipulation, 2024.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Data scaling laws in imitation learning for robotic manipulation, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:06.268768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:59.815324Z digest=sha256:76b920fab4492375de4226f5e7e627320fecd7c4cffd583c8d143754703d824c

Observation 20673b07-97a8-47c7-8149-8c6dcc97bac2 · outbound

This paper cites Multimodal diffusion transformer: Learning versatile behavior from multimodal goals.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Multimodal diffusion transformer: Learning versatile behavior from multimodal goals

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:06.066280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:25:59.947567Z digest=sha256:ff318cca5f9be280ffad6ab8d769b8e2069c92a90bc80cb3b5ad19f89da3ab9d

Observation c40028a2-2dd3-4dc3-9500-aae8fa2059bb · outbound

This paper cites Aloha unleashed: A simple recipe for robot dexterity.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Aloha unleashed: A simple recipe for robot dexterity

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:05.930052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:26:00.024746Z digest=sha256:0d0601ab31bcc45fa2cbab713f4f35c73f4217286828a7211aa503f8a02ed0a1

Observation 33c41de6-8225-42ea-b84a-a0d2332fb78f · outbound

This paper cites Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.168043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.168043Z digest=sha256:772bb581b3485abfd03634dfc59d894d71a306e1d6ed8a1aa25c1129cf64ff47

Observation 7a657ba2-5406-4fa1-a63d-7a7ebe6c9e00 · outbound

This paper cites Feedback Efficient Online Fine-Tuning of Diffusion Models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Feedback Efficient Online Fine-Tuning of Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.311389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.311389Z digest=sha256:5348736ef4d47a930fbb8f45a5e4c27e79596893d5179d5c87bee1fc4370e646

Observation 17512fd7-14bc-4572-bf21-f5f0627d553b · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Gemini Robotics: Bringing AI into the Physical World

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.434117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.434117Z digest=sha256:9652e72afac24e652420234bdafa2a970028350449a093ace3dddf2bdc7e89fd

Observation f5564302-9a1c-45ad-88cd-5fc42a2d1342 · outbound

This paper cites Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.512531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.512531Z digest=sha256:1be5170bb87968b8aee51481e94c10a0844f24a330ef12a3539fa458da099619

Observation 370bbd6c-6ff8-4ac2-b488-f84ef934979a · outbound

This paper cites OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.558392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.558392Z digest=sha256:68d7ac83ffd1ac8f1016a505af39afb168ee152b07a1e72a6bf4d9a201bf025b

Observation 83bca668-757d-4639-8e0b-39db868c162f · outbound

This paper cites Quar-vla: Vision-language-action model for quadruped robots.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Quar-vla: Vision-language-action model for quadruped robots

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.681860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.681860Z digest=sha256:bdc53106e26ffb6fa331f373da745ed8b31e5dfd990c9d712154d45d49ae9d7b

Observation dd9caa4d-7dad-4310-b342-57f2e7c467e9 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.781836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.781836Z digest=sha256:03e890ebac31f2d0598f1fd1ac01241111711ab248eddf4b526c001a946c66cb

Observation b01c3a5a-9fba-4a2f-9007-4fd83e3b3deb · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.871918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:00.871918Z digest=sha256:510e81be9378d124fd257d60df38b4c4cd0ff10c6fc9d81ae6b3634497521996

Observation 3795b51a-6f3b-4b44-b2ac-45e454f1f574 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:01.016059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:01.016059Z digest=sha256:dca96833b5c514d3d8fccea68f3617d3b49a6d1fda8d2634c12a0d93ae4ac958

Observation 13683616-8f83-449b-9bcf-61f3e9db0953 · outbound

This paper cites Robomamba: Efficient vision- language-action model for robotic reasoning and manipulation.Advances in Neural Information Processing Systems, 37:40085–40110, 2024.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Robomamba: Efficient vision- language-action model for robotic reasoning and manipulation.Advances in Neural Information Processing Systems, 37:40085–40110, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:05.810394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:26:01.112089Z digest=sha256:805d5e581d1e5cefdaaf448097729c83d015db41ae608fe84a3247c2486dbf47

Observation aebf1c98-049a-45c9-abfe-6ebad3420eea · outbound

This paper cites Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution.Advances in Neural Information Processing Systems, 37:56619–56643, 2024.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution.Advances in Neural Information Processing Systems, 37:56619–56643, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:01.201975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:01.201975Z digest=sha256:af4a7bd67f6ff25bee146fd0d068538883b773eb1d7e1fac03c1dd8eb96331d4

Observation 0e2c2f74-258b-4012-91f4-fda7d7938d15 · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:01.430507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:01.430507Z digest=sha256:365c3f94a482e312203bfdd5fe1bf07b346d41adef13640154158cc1bcaae86f

Observation d02f426d-edf6-4864-a9e6-ae8c32b90383 · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:01.814858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:01.814858Z digest=sha256:f65236e034a727d460e23e63e5648e1cee0e3ba855460657c4c0d1d92f75ef94

Observation 1bee1918-1f07-4b6a-b3bd-1b580917c248 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:02.337833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:02.337833Z digest=sha256:45172e49436d939a87603dc93c35c9005cac531d6af290001423bfd0eba74492

Observation fc454b8e-d899-488d-9009-45f1e0e41b7c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Chain-of-thought prompting elicits reasoning in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.207138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.207138Z digest=sha256:0a2ecd1755415f8cdeb911f85ab0046e6885f843fcd2e04d84060ea659ad7fdb

Observation 483ba450-c5cf-441a-b022-761b421380c2 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.647132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.647132Z digest=sha256:db79fe9f1cd9d2ef362779cf97d3a6f00f0bb559846b2b57333ff5e4e41a9cc6

Observation e40c096e-f28b-4109-bf45-bfaee6bd7b19 · outbound

This paper cites CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.787816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.787816Z digest=sha256:d5bc13582b05058b01b1bf33f788f6358c2a034ffd73899f6ecb174bba8a89ee

Observation f450f948-70e7-4198-8484-78173268b6a4 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.827165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.827165Z digest=sha256:2282a02c3f836fa3a71febd1109c160f345db55c971c31e5413d3967e2597ea1

Observation 0b71bc9a-7055-4762-b6ba-40edfcfacf0e · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.866819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.866819Z digest=sha256:5aba1e0a03e24eebb5d651bc0e2f1ae6d1edb9bfaa65a7cd9a0e7f7637f0e696

Observation e2f290be-f8a1-46c1-8e1c-654bbba7905d · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.902047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.902047Z digest=sha256:3272096431ee205e964ae68fa8aad8622661d2bff020130c30e976120589cc28

Observation b1a9f423-8651-4a84-80f2-4494d6b44898 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.935938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.935938Z digest=sha256:bc8a7af09a3ee65e94f4f14c41cc8982b13083582ec63d154961232620fef58a

Observation 4971d5ef-374a-449f-8a17-49f31795baf1 · outbound

This paper cites Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.982384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.982384Z digest=sha256:b14499def0e426086b2b271def621cbcfb499caa59beafc5974b887f1e4d37f8

Observation c35bd5b3-604d-40bf-8062-347a8b121928 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:04.068620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:04.068620Z digest=sha256:bc2d18ce9e02a19969237e47faec961246e16eca36253967dd23cefaa230ba6d

Observation dcd4bf35-08f6-4584-8d82-ba1130d069e3 · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:04.124888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:04.124888Z digest=sha256:2fb85927248d22841bf1fcb81ff8f85e593e709c3a6e5d5b4932927f82097abb

Observation 98c2a8d3-5339-4cca-98f9-0c01eec39b24 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:04.195405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:04.195405Z digest=sha256:2e42adad4c812ed831b88434946f956f67640d661199965be3159ccc92249783

Observation 282edde4-87b7-421b-92d2-41e24090bf53 · outbound

This paper cites Microsoft coco: Common objects in context.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Microsoft coco: Common objects in context

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:05.650792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:26:04.228991Z digest=sha256:881c1d42cbb1a6e2b1510c4afaf2c71f0d3844a9c23c9db1ed42a177c0a30427

Observation 04eea78b-1e08-4379-92fc-403f3015ca81 · outbound

This paper cites Towards vqa models that can read.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Towards vqa models that can read

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:05.480969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:26:04.282209Z digest=sha256:200d2f900515bd1365f57fce991c814b64608cf100d68706ec588d26947905fc

Observation 61f29ec0-7e88-4d1d-985f-f887751adcd3 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:04.344657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:04.344657Z digest=sha256:8791638d9000e9e94c2b64840e7374621da7684efcdf567377c9b98a8fdea28c

Observation 4f543ec4-7f8f-4b01-ba79-aae7b6ba8950 · outbound

This paper cites Octo: An open-source generalist robot policy.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Octo: An open-source generalist robot policy

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:05.274977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:26:04.383484Z digest=sha256:2c0051973dce569f6bd22f1aa05a822559b4926d8b9443aa6b98292de5e2d2a6

Pith citing papers

Observation 7b90b889-cb7a-4be7-bbaf-b74925dd9442 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.657904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:9494c4a67d5086dd000153176d1c9f5791e39ae769fd253e9d86985de367d721

Observation a0e2b8c1-28e1-46ef-bed6-456df92c9d11 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.586386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:6b359b55a92ecd01461fbe0cf5189538f61e42b50dfa38b66daf344482a405f9

Observation 3d798a09-016e-4b4a-badf-e51fd4ed9225 · inbound

RationalVLA: A Rational Vision-Language-Action Model with Dual System cites this paper.

RationalVLA: A Rational Vision-Language-Action Model with Dual System ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:46.040172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:24:46.040172Z digest=sha256:1280b978981b1d054a73ec8aaf2276d2453bb3639bb86da910b40b3583ebf5d5

Observation 91bd68e6-817b-4383-90fe-56f24360044f · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.346015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:f0490e35f7c09d51996fc85c696b90ada33978492d665e01a74bf2c3ef80d4ac

Observation 6cb31c1a-69ee-47cd-914e-be08a3b32bd5 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:28.034660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:28.034660Z digest=sha256:6ede8894cf85c1f0777f2a4baedeefb30c8c0cceaabdcaac406dff9c0398c65a

Observation 6a9e238f-c452-44a9-ac56-732d17bd6da2 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.693054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:57787a030201eae8f97387ebcfa0f6a9c73b407859e3799a5f57eaf925f10126

Observation 4cfccbd2-71cf-4a53-9f38-40d77c92705b · inbound

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning cites this paper.

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:17.365848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T16:12:45.211236Z digest=sha256:61703d6fd9f85820ec17c93316dd5316e877a325022f6bab1500eaa15870e875

Observation 51b4255c-f5d1-4777-b590-5064be053c6e · inbound

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models cites this paper.

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:21.135705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T19:45:10.232982Z digest=sha256:fee9ea66496f9137a27adef90b7bd07f2a157e8ac45b40c74c2eff08cd98f91d

Observation 20f8ca47-5f32-4d55-89c2-651193f0b6ff · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:46:04.954056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T15:14:03.865035Z digest=sha256:817322c1653bcfa460f367019423aad8e629f20e065d97f4f576a5d947566e94

Observation 1eb1fddc-ed36-443c-8628-0548704f57c9 · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:05:45.083383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:1cbad00bf85ac2d7724fc407712a11333555de74bcb730265db6718e39b4beb1

Observation 999a750d-1f10-4cbf-95c6-c0e76e5ec01a · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.282356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:94d1b7fbf4072a8ab107456cdb09a761b944bc3777a1522f98e7fce5c90d34cf

Observation c4fedb81-ba07-4d9a-9aef-731750d8ad25 · inbound

PhysBrain 1.0 Technical Report cites this paper.

PhysBrain 1.0 Technical Report ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.958760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T16:34:44.204055Z digest=sha256:8ca190ae087f60322071444296bb4d42f195e385ec22714f392ad05ae821b011

Observation 39aa038f-dc04-48fb-a73a-f3a8e39bc85b · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.242721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:64a9942ee763f1664410105315a285cc8400a5eb9369b05fd88e768d1b135c5a

Observation e79f361a-6613-4805-b7cc-5660a41a7c0e · inbound

Continuous Reasoning for Vision-Language-Action cites this paper.

Continuous Reasoning for Vision-Language-Action ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.541501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T22:07:07.401335Z digest=sha256:cb798ad3e2df16c2e68a737624d4ede7398c1853f5659f1b89d78152b6382b15

Observation 13aab588-7370-48d8-8881-0bc4931c2e1e · inbound

Policy-based Foveated Imaging and Perception cites this paper.

Policy-based Foveated Imaging and Perception ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:17.704616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T15:23:48.688620Z digest=sha256:4c195309eff6ada5dde4d94ebf74669d05e95c4d78cbc2db4f9ad1d5164639a6

Observation 12d49e0d-37ca-4ab3-97e5-d7c91875184a · inbound

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform cites this paper.

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.237371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T09:31:56.289129Z digest=sha256:397a0935c429f03a44cb8744925d840645913b064d858271025b396624102ebe

Observation 668abe37-031f-4b57-9a78-5655687e2540 · inbound

Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation cites this paper.

Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.318532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T08:28:33.173481Z digest=sha256:f140963a50fcb617432a5807414d7afa51ad1723ca296d5b156ed0fe3189c086

Observation 89823bb4-2bfc-4161-b2cc-78d9df11fd8a · inbound

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach cites this paper.

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T19:52:24.975588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T19:52:24.975588Z digest=sha256:5bb6844f7f64818e40b162b39c21e017fb72136b7b6089ed852ac7fb719308de

Observation 731e90b9-b286-469b-b613-037f883a850c · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:43.005296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:43.005296Z digest=sha256:022d7db27daf2ab749842cac424c79e0f81ab8e4e7505fd37d706ef59218ffba