Pith. sign in

Paper Citation Record · LEDGER

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

As of 10 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2608.05903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05903 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:48.134898Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4656ede-f50d-4899-a5a7-de2b408c1173 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.922507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.922507Z digest=sha256:bfe9d5ca06fd135b1b1094b2a3c46f7d16c95db6db12b07f2a35d8065bd8aef7

Observation 1222e5b5-f207-49a1-962f-fc8c8e1bc8ee · outbound

This paper cites How learning by reconstruction produces uninformative features for perception.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models How learning by reconstruction produces uninformative features for perception

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.928080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.928080Z digest=sha256:9db42e728cd563be484931eb302f072546e04979a950b7c3f3b30ea5c0d4de9c

Observation 7c434683-3031-45da-8088-35bcbdaf261e · outbound

This paper cites Motus: A unified latent action world model.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Motus: A unified latent action world model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.932662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.932662Z digest=sha256:da3a3f7d1ec3b546a6a986bca6f73b80bc75503b831368f58f7df283133356eb

Observation 861da261-3035-4b17-84c9-1b2650790329 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.937067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.937067Z digest=sha256:d926b7cc8e1de6def773b2343c2098396d8070fb8bec9fde802f880dd313e98a

Observation 02c89345-bc2d-4519-9f3a-8bb185f4ea93 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.941921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.941921Z digest=sha256:04defa3db3866cb5fc89dd7476a52ce8a01a67cf628f9cbbb9132d4729b8c06b

Observation 228d84d0-e82f-4430-9d22-2fbfd362cb2d · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.946682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.946682Z digest=sha256:9b653d6b424e2ad7ed862dfabc8b3976a73e042d8d5baed0ec1c9aef8153b1b8

Observation 5fc537c6-01e6-4577-97df-18a21e9aafa0 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models WorldVLA: Towards Autoregressive Action World Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.951523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.951523Z digest=sha256:38fb82cb9e9581ffd97c26c0df806addc822a02dc2ba3a90a11bcd2273888dc3

Observation be67d509-98b8-48b2-9a6d-ce3413b2d76b · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.956392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.956392Z digest=sha256:e4c8c1dc08946b28880b7fa304973419fc04318516962ada40dee25285bad22e

Observation b807ea57-1fec-4a61-8274-dfefb84668f1 · outbound

This paper cites Lawam: Latent world action models for efficient dynamics-aware robot policies.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Lawam: Latent world action models for efficient dynamics-aware robot policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.961239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.961239Z digest=sha256:04c08437715d1dcaa356eb05e08380d3e4b37fad417cb79a54638d9adb1d377e

Observation bbde95d9-ecfe-43f0-b393-07365fed4b99 · outbound

This paper cites RoboTwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models RoboTwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.965308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.965308Z digest=sha256:736d8bd491dbf29f25320eeebccec04eebf18c797e6faba194bc59d486da1911

Observation 34c671af-598d-434b-b0a0-55d3d5239725 · outbound

This paper cites Learning Universal Policies via Text-Guided Video Generation.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Learning Universal Policies via Text-Guided Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.969390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.969390Z digest=sha256:9fb4284f696f0d1548ff20f04a5f1510c1f9df94ba66b6784a820a4c2c9e43bb

Observation 1a6b8785-226e-4feb-86a5-441be26e6a87 · outbound

This paper cites LIBERO-Plus : A progressive robustness benchmark for visual-language-action models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models LIBERO-Plus : A progressive robustness benchmark for visual-language-action models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.973666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.973666Z digest=sha256:9e090c4f695351d2f151fb06eff38974cc9691b74c8e4e38bb94497126698b6c

Observation 8c8c58b0-839b-40c0-840d-087a0c25bc68 · outbound

This paper cites Prediction with action: Visual policy learning via joint denoising process.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Prediction with action: Visual policy learning via joint denoising process

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.977676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.977676Z digest=sha256:852114b028d4b90b083123e400e9ebcf8e43e09c19c814b08a1ae8eb651b30fe

Observation dde6585f-c910-4a68-b23a-8afe9f7a6c43 · outbound

This paper cites World Models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models World Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.981945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.981945Z digest=sha256:13d71851e79fa0223f3fb1413da149c29a7d7300915f73717f923b6ef258cf8e

Observation ce1ed570-0573-4068-a645-117cf0b45d06 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.986412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.986412Z digest=sha256:cc937bdcc61631ebb870313c3811caaecdde19ecaa13225a01905457e235edef

Observation d2cf1e9c-ae38-4201-9542-e00fb0d69e61 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.990981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.990981Z digest=sha256:425a8af9f011f9c2d8900ec7649f6718e3bf9a26475146afb01096a138f835f4

Observation 41391cf7-aa3f-4014-9dcf-d350faa40c6a · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:47.995526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:47.995526Z digest=sha256:91293c6c8ad8b0a06e487ebe76c67f136d80a5bf424c68123e430549eb7f8b82

Observation 4488247b-6e8c-42bf-a6a1-49782cbee7fe · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.000072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.000072Z digest=sha256:fc87fab62325d081a76c2418ae4b7d062365c6c8701e4ff77b845a4226cd3c5f

Observation ecbc53fe-a52b-4eb0-8909-1b14e036a0d7 · outbound

This paper cites A path towards autonomous machine intelligence version 0.9.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models A path towards autonomous machine intelligence version 0.9

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.004347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.004347Z digest=sha256:d8f39d85b292cd9b523e207d30794dd2e3ad96c3519b114e8babee26ebf08746

Observation c1339ea9-f9f0-49ed-9274-e8db708931bb · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision-language-action model.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Spatial forcing: Implicit spatial representation alignment for vision-language-action model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.008300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.008300Z digest=sha256:0c98cb0295aad6dc508e570d398fe619d67529278fcd39852f2ca247f8d4ac1c

Observation 80e99cfe-4d75-4edc-8e32-744c807af7be · outbound

This paper cites Causal World Modeling for Robot Control.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Causal World Modeling for Robot Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.012576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.012576Z digest=sha256:07920d2fab4cf8f2c3151ad8e7cc853bfc82493d0243a9fe308b2da41c729fa6

Observation 98b8ab2c-f921-4fe1-b943-3b43c8d44166 · outbound

This paper cites Genie envisioner: A unified world foundation platform for robotic manipulation.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Genie envisioner: A unified world foundation platform for robotic manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.016971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.016971Z digest=sha256:d0357e19b35592d60e139d75ba64266f66794867870c6370b4118991f4757dd7

Observation 26a2dc3f-7a77-4b5d-bfb0-b10b0f2351a3 · outbound

This paper cites Chen, Zhenyu Li, Yang Zhao, Sida Peng, Hengkai Guo, Xiaowei Zhou, Guang Shi, Jiashi Feng, and Bingyi Kang.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Chen, Zhenyu Li, Yang Zhao, Sida Peng, Hengkai Guo, Xiaowei Zhou, Guang Shi, Jiashi Feng, and Bingyi Kang

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.020947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.020947Z digest=sha256:c6b4a30cc8e3c7b856c4f8a9d7608e1ac9ae97eb838f9da9c7cd6aad618a62f1

Observation 27e78309-e02a-45f0-bea3-19508a1355a4 · outbound

This paper cites Evo-0: Vision-language-action model with implicit spatial understanding.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Evo-0: Vision-language-action model with implicit spatial understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.025113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.025113Z digest=sha256:49cf96152a6bf885da7dee0d6f18e36d842c57e140104e42fc18771c8b11ad2f

Observation be1497b2-92a0-40b9-85c5-72682119220d · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.029562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.029562Z digest=sha256:78142b884774ec6fec89173f8b0a114e9309409f0927219a629987d12d89e9ab

Observation 0de3e8b0-1087-4239-8caf-ea5cf5cd2aef · outbound

This paper cites LDA-1B : Scaling latent dynamics action model via universal embodied data ingestion.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models LDA-1B : Scaling latent dynamics action model via universal embodied data ingestion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.033852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.033852Z digest=sha256:365ac2ff51db3ca7bffaf10e746faf7829e7198bf66ec9f1ba39648070de706c

Observation 324deabb-cbce-4e2d-ad93-d1774e7a40b0 · outbound

This paper cites Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.038045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.038045Z digest=sha256:adecb423be2e5322537c35dc2e5e5abb2f456f6d9c5a114b1c5441b260dfec2e

Observation 4c245f2a-47c9-43d4-84d2-30910db36a7a · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Cosmos World Foundation Model Platform for Physical AI

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.043315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.043315Z digest=sha256:5a95001dc38b5240949865095c1570b6b5582257dc3841693ff7c36fc0c115a3

Observation 46abd74f-3980-4e42-b8f4-20b8555cfe06 · outbound

This paper cites mimic-video : Video-action models for generalizable robot control beyond VLA s.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models mimic-video : Video-action models for generalizable robot control beyond VLA s

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.047875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.047875Z digest=sha256:6d9abf6af7d040e3c848448a041beb4f6e777fe8a5db20b8d8fa3ec73018f583

Observation de93778c-a4cb-4c4b-b9d2-8f19727bcd0a · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.051995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.051995Z digest=sha256:869fbbadbc1e92b6c202fce9851e7532b86c841320c8b09df2a34368eba3e229

Observation 716a390c-f6e8-46b3-975f-c588f8140822 · outbound

This paper cites World Action Models: A Survey.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models World Action Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.056258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.056258Z digest=sha256:fb5077d7750826a626d193ab61cb0e2e30af7f80f05e477d93cf5ff471e6f83d

Observation 32c32824-b26b-4ffa-955a-e2cf1795d062 · outbound

This paper cites DINOv3.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models DINOv3

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.060849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.060849Z digest=sha256:4b727b7b9e77c9d239f468063f42e7b312eacd9b321cb7eb64383736fa6d0597

Observation 1041b97c-be14-4232-a8e7-8e454e373539 · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.065577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.065577Z digest=sha256:9271f5fd7dabe22d8c65737b4e2fb385a4465996a7a977517ebc0efbc66eb687

Observation 2a7ac604-bd01-4f7c-ad7a-95e2e290ee9f · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.069778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.069778Z digest=sha256:c5d21c5144bde693161abb6bd001b9a0e19766a2275749569acc0b3a6a777533

Observation 183b9e5d-dd94-49d1-ac76-0afbcddac176 · outbound

This paper cites RepWAM: World Action Modeling with Representation Visual-Action Tokenizers.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.074308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.074308Z digest=sha256:c915cd5262b9b74e4d856fe132365be11df98f3b07cb9f6db8a10f5cc91157ef

Observation 0ba81577-a385-4d83-8e11-1fa2e3e91cb8 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.078732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.078732Z digest=sha256:cc7880c7a03bb109d66875c2b48b25a92de46c11f71f093b6ac4821d00b8e558

Observation 5ec1ab36-5b69-4e54-9563-af133ec4d5a2 · outbound

This paper cites FutureVLA : Joint visuomotor prediction for vision-language-action model.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models FutureVLA : Joint visuomotor prediction for vision-language-action model

Reference 37

Resolution
verified exact
doi, observed 2026-08-10T04:29:48.648570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:29:48.082883Z digest=sha256:240c2f2bf2cbb54826f90ce9cce0c4ceb64680029ec2e1fec7a00084215f4cb7

Observation ed0f1e22-d68a-4cea-b41e-d3da08b7c088 · outbound

This paper cites Open-world hand-object interaction video generation based on structure and contact-aware representation.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Open-world hand-object interaction video generation based on structure and contact-aware representation

Reference 38

Resolution
verified exact
doi, observed 2026-08-10T04:29:48.577201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:29:48.087055Z digest=sha256:7ee0a42a6e8e97a96db281a0e8d7da74f81dde5b0a15fccb11663d6a667324da

Observation 1cc128ed-4269-4287-9167-056263ce5d7d · outbound

This paper cites S-VAM : Shortcut video-action model by self-distilling geometric and semantic foresight.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models S-VAM : Shortcut video-action model by self-distilling geometric and semantic foresight

Reference 39

Resolution
verified exact
doi, observed 2026-08-10T04:29:48.513712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:29:48.091183Z digest=sha256:ccb5764a581f864be14c736f9217635215fedae6b7972f3fcbb06fe7e13dc3b5

Observation c4b5b15e-23f9-4823-8839-77fbb5f559dd · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.095417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.095417Z digest=sha256:07a4c156d33270130102171e0a651c16733e84c51e3f7c31744ab8f35043ad83

Observation 121f9558-7f04-4ad5-b05c-72ae241ce07f · outbound

This paper cites World Action Models are Zero-shot Policies.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models World Action Models are Zero-shot Policies

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.100088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.100088Z digest=sha256:5433d7d2e1c38e8141358b27c6eb8cf2e7e6f0feb95dde4d822f152c3e420b37

Observation dc978ee6-8bfd-4a90-a742-8d82ae340870 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.104496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.104496Z digest=sha256:24d82e72a5e7cf8896b842ae824abae6dbb7fb8511b214cc6ff1b831ebaf2edb

Observation 6c7ddad3-3c45-4236-bce6-01599c6d10cc · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.108739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.108739Z digest=sha256:8689dff0a75bc077f332336e05b31d3f8f5ca0ab3fbf123c429bc20db70aefa6

Observation e0f3f362-803f-4634-9317-0c9e4de7af53 · outbound

This paper cites Do World Action Models Generalize Better than VLAs? A Robustness Study.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models Do World Action Models Generalize Better than VLAs? A Robustness Study

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.112964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.112964Z digest=sha256:aae66f35bf9715767d70f029ddf8ab9813a7fda4b0f5b7fd494359e5b95ef13f

Observation 2c6d7765-6ae1-441b-9c43-fdaa48b0bb94 · outbound

This paper cites FRAPPE : Infusing world modeling into generalist policies via multiple future representation alignment.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models FRAPPE : Infusing world modeling into generalist policies via multiple future representation alignment

Reference 45

Resolution
verified exact
doi, observed 2026-08-10T04:29:48.396455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:29:48.118088Z digest=sha256:38268f3e6d475db597a7a22da68d4863563b22e49a666668d91483c6f485dc3a

Observation 35857a76-840a-43df-9679-a9a0136327d4 · outbound

This paper cites FLARE : Robot learning with implicit world modeling.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models FLARE : Robot learning with implicit world modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.122492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.122492Z digest=sha256:348e872a821e875167b731769688005990845b645f945f17d74720042f42bc2f

Observation 96fb89a0-37c9-4ed8-9ec5-ed94ef617ab0 · outbound

This paper cites FlowVLA : Visual chain of thought-based motion reasoning for vision-language-action models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models FlowVLA : Visual chain of thought-based motion reasoning for vision-language-action models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.126464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.126464Z digest=sha256:d9beb223f738e2b24799fb890a529fce08bb4b02205e04c3a79f983c40506e77

Observation 1f828dd3-a70c-4215-b760-a2f89fa6960c · outbound

This paper cites DualCoT-VLA : Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models DualCoT-VLA : Visual-linguistic chain of thought via parallel reasoning for vision-language-action models

Reference 48

Resolution
verified exact
doi, observed 2026-08-10T04:29:48.243625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:29:48.130632Z digest=sha256:a22fe92157665c0f743dd6a22c50efd7e4df1482d4aaee4c7007b1b44827e69a

Observation 337cbdce-6142-4b77-a34a-e47d679b8372 · outbound

This paper cites DINO-WM : World models on pre-trained visual features enable zero-shot planning.

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models DINO-WM : World models on pre-trained visual features enable zero-shot planning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:48.134898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:48.134898Z digest=sha256:7d4cd79cfc78404942e7d8a0622de74dec2edec82985522410042dc8641426da

Pith citing papers

No inbound Pith citation observations are available.