Pith. sign in

Paper Citation Record · LEDGER

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.20988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20988 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:49:50.059578Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 527b2b1b-1a3d-45b8-ada3-d2a399e57285 · outbound

This paper cites Ppu introduction.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Ppu introduction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:46.436457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:46.436457Z digest=sha256:b4ea12f0479d192cc81feead469dc47acb0d9730244073258e8fa867dbcc8183

Observation e210e35e-e839-4cdd-87f1-2b4f3c97f88a · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving GAIA-1: A Generative World Model for Autonomous Driving

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.208845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.208845Z digest=sha256:e784f4c94664b01fbec3d07724274bd2679f30c6f957f3107ebf7ec8a9b5f7cb

Observation 199d8f8d-cf73-4203-a52c-12b3744e3552 · outbound

This paper cites CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.335919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.335919Z digest=sha256:b6c66d97d8fa99307ceab5fa7147dfb8af96e25b340e407d353b853a9a82a34f

Observation 7adb7f2a-b43d-4ebe-aec0-74657411b552 · outbound

This paper cites ADriver-I: A General World Model for Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving ADriver-I: A General World Model for Autonomous Driving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.615397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.615397Z digest=sha256:ddb42dce0157275435022c11ffc3142fcafa4d5c450667f79eae3c6a9513eaa6

Observation 01999687-242a-4c40-9d0e-ebfef0c9b6b6 · outbound

This paper cites AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.783298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.783298Z digest=sha256:e14c1a600b058061a9f0b31538d2675267e358cc3a3e947e4bc8a5d29f88ebd5

Observation 4e999d6c-2cf0-45f0-ad79-1083fb7bba44 · outbound

This paper cites DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.902688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.902688Z digest=sha256:aeff51e05fd401f50eb1010df6b7864cf9451ed4039b85d1fdf767346d63238c

Observation b084b62c-a93a-46d8-80f3-2954eb2eeb94 · outbound

This paper cites Flow Matching for Generative Modeling.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Flow Matching for Generative Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.026207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.026207Z digest=sha256:fdc09e7eb42cc2a38dc34d93a23f3b22d628c339e1fe04661be69abff0d6c0a6

Observation 90980c78-0455-40f7-b57c-a88828014828 · outbound

This paper cites Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.309380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.309380Z digest=sha256:0b8d9d745ad92a58da552607ea9b749e0dc4b7d70363ca80d2994f561b350a5f

Observation bf0f9fd0-9617-4dac-9523-82adb793480f · outbound

This paper cites ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.429071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.429071Z digest=sha256:ec3f02ef46ca2d0fee83bf9cc38d616a046f80ebb9d6461d53b7fa0731f3f08e

Observation 071bc17c-273b-4938-8d51-a4b17dbb3095 · outbound

This paper cites Latent Chain-of-Thought World Modeling for End-to-End Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Latent Chain-of-Thought World Modeling for End-to-End Driving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.558259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.558259Z digest=sha256:b0d5d23a4355289a8651938fff3ed943907d55131dd79c9fd11e0225391c54a2

Observation e94ad116-43ba-4345-a60f-68f8681a30b3 · outbound

This paper cites The role of world models in shaping autonomous driving: A comprehensive survey.arXiv preprint arXiv:2502.10498,.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving The role of world models in shaping autonomous driving: A comprehensive survey.arXiv preprint arXiv:2502.10498,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.668577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.668577Z digest=sha256:2cda2dd3bfbdcda0b6e89b09c84c61f522cbcc80a7def26c3a308e766cbb4ee3

Observation 34b55adb-201d-4a08-99b0-d260433773c7 · outbound

This paper cites Latent-wam: Latent world action modeling for end-to-end autonomous driving.arXiv preprint arXiv:2603.24581,.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Latent-wam: Latent world action modeling for end-to-end autonomous driving.arXiv preprint arXiv:2603.24581,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.777466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.777466Z digest=sha256:e83bdee185ecdf7d4f09532939e1cce718a60b3d8fcc08216ceb849113749068

Observation e305866d-0e6f-487f-b382-fd21c26fbd9f · outbound

This paper cites DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving.arXiv preprint arXiv:2312.09245,.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving.arXiv preprint arXiv:2312.09245,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.925835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.925835Z digest=sha256:2259a60089ce2cdfab7222df290e08bf1acd6e42133755e49af312ba62bdf811

Observation c91c9ba3-ef20-4c8e-a1e2-53fca3c3cfa6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Emu3: Next-Token Prediction is All You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.102938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.102938Z digest=sha256:397a7439669d13bfb1c448dd6c79d3cd22e2dd58acc9320d076d1c1cad7d11d7

Observation 5819cbac-3864-4975-8c75-9336030f2c9f · outbound

This paper cites Large Motion Video Autoencoding with Cross-modal Video VAE.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Large Motion Video Autoencoding with Cross-modal Video VAE

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.239066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.239066Z digest=sha256:d012deebed090a011550b5b9ca9ef334580d839ed23390421b6c9e144439b266

Observation 16267b38-b916-4faf-8bc4-952b1c0f802b · outbound

This paper cites Resim: Reliable world simulation for autonomous driving.Advances in Neural Information Processing Systems, 38:167710–167741, 2026a.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Resim: Reliable world simulation for autonomous driving.Advances in Neural Information Processing Systems, 38:167710–167741, 2026a

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.393759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.393759Z digest=sha256:99b06365bd44776d62c1e1873740bff8cbd4abbdf5111f8969a4496243dd823c

Observation 9eacb8fb-c161-495b-8995-bff457403a36 · outbound

This paper cites FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.521156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.521156Z digest=sha256:9a284e2f42932bc03b26acf60604f66753b30ae8e826ea277bd552813e71f82a

Observation 50f56db1-2421-44a3-bfa8-0ca140526cd0 · outbound

This paper cites Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, and Xing Wei.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, and Xing Wei

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.657500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.657500Z digest=sha256:4e2bdcb7fa71bff7ac1105651c4d10d76fb9b0aa2755b11c05ecddb2201bb59e

Observation 7ed22ca5-ef81-4012-ac4b-ee7e91c1ecb7 · outbound

This paper cites Resworld: Tem- poral residual world model for end-to-end autonomous driving.arXiv preprint arXiv:2602.10884,.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Resworld: Tem- poral residual world model for end-to-end autonomous driving.arXiv preprint arXiv:2602.10884,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.763901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.763901Z digest=sha256:5ddc7d2a26b87ea3dbe8cf6b48f11a368153ca59fbd1b67ac352db1cf9a74f57

Observation dd329218-274c-48e7-a7fe-63efbe1b1e5b · outbound

This paper cites Doe-1: Closed-Loop Autonomous Driving with Large World Model.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Doe-1: Closed-Loop Autonomous Driving with Large World Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:49.882619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:49.882619Z digest=sha256:67e9ce441fba7b6747986adf9b533cbd6029131d8fc8dba1da1f7721144f64d0

Observation 4284fe62-e282-40db-9421-97ac6a85ce86 · outbound

This paper cites Open- drivevla: Towards end-to-end autonomous driving with large vision language action model.arXiv preprint arXiv:2503.23463, 2025a.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Open- drivevla: Towards end-to-end autonomous driving with large vision language action model.arXiv preprint arXiv:2503.23463, 2025a

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:50.059578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:50.059578Z digest=sha256:7367649795029df8c35786265a2c07deeccc42963a3e7e2ea47ac5e29bd629d0

Observation 5b80a77f-b2cd-48ba-ad34-7984a5b662ab · outbound

This paper cites Pseudo-simulation for autonomous driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Pseudo-simulation for autonomous driving

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:46.743096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:46.743096Z digest=sha256:5e3cd86f11304477356bd25cb922aa9cbca37227c2ba0f9fefa1e55560cc7b23

Observation 0f455f9b-bd56-4e93-b47c-568b50e86279 · outbound

This paper cites Driveworld-vla: Unified latent-space world modeling with vision-language-action for autonomous driving.arXiv preprint arXiv:2602.06521,.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Driveworld-vla: Unified latent-space world modeling with vision-language-action for autonomous driving.arXiv preprint arXiv:2602.06521,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.149103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.149103Z digest=sha256:85ffc185232edcdf4c42c57ff107119cde9abde9019284bc711c18dde767da25

Observation f7c32a07-87ab-4ce4-a3e2-2278d4a76736 · outbound

This paper cites NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:46.880137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:46.880137Z digest=sha256:3159d33dc88675614515bddd36f099006c4f591482e30330b48ab92ea2eec688

Observation df31e6e9-5e6d-4cd9-866b-2455a81261a8 · outbound

This paper cites NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:46.573313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:46.573313Z digest=sha256:8e18f7b0696579dc2d0f3fdd86f5b2331e0c5a6d96e3b4f9a80f9f706d690910

Observation 079b8758-c960-45f7-8fb1-e1c89b2d0bf6 · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.035253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.035253Z digest=sha256:9c08b3dbca3335182305071e0178e35ac5e64a51593c3325368cbc7966ea8bbf

Observation 72db3736-993c-4107-8963-e4ea4fdb0f91 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:47.488543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:47.488543Z digest=sha256:d5deb919f9dcf7e3358da9725c44c791d36f67005e9277f00194af9a58b1d8ef

Pith citing papers

No inbound Pith citation observations are available.