Pith. sign in

Paper Citation Record · LEDGER

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

As of 19 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.09730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09730 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:52:55.674025Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6db37f8-df69-4a5b-a9bc-1bec9f8e7372 · outbound

This paper cites Qwen3-VL Technical Report.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.636885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.636885Z digest=sha256:c426a2c6cc6ae3a5f34aa39d09eec6166b882e3e3438b6e8556f8b43314d0df7

Observation 0dc1c50f-f5d6-4f63-afd5-926905ff0386 · outbound

This paper cites VLA-JEPA: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098,.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling VLA-JEPA: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.642353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.642353Z digest=sha256:e667f50dbce47abd129bf4293dbb0e4e018455a800c3cac058d54f5dc86468b5

Observation a21cfbb2-8175-4df4-9e22-40192f6ccc35 · outbound

This paper cites One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.646630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.646630Z digest=sha256:ca893a08ea9130f98c03bc16f171577716cd5cfd7aaada22321b463d32cbe59e

Observation b75afe70-f9f2-464d-80c5-046d6dc53446 · outbound

This paper cites World2Act: Latent Action Post-Training from World Model Dynamics.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling World2Act: Latent Action Post-Training from World Model Dynamics

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:52:55.836408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:52:55.651035Z digest=sha256:9cc20321ddf6382577cdcc22765a6e5c6cd458f97aa0fcf57b71224292d88f56

Observation f6c4bcfb-0b1a-4d54-9a97-a942ecadba9a · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling BridgeData V2: A Dataset for Robot Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.655345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.655345Z digest=sha256:05cf688870f48687b3db3577013006ecc7e699f2a2dd92fb0bba2c21bbc51df2

Observation 95c30bd7-31d3-4bf8-b858-34080c05bac9 · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.659762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.659762Z digest=sha256:061e48cabac9d3c5904c88dbb9f2b35d103d37b77aa0c3fa41e73c62b92ad64f

Observation 75e865f1-f1c4-4b87-837a-817bd924df4e · outbound

This paper cites 𝛿VLA:Prior-guidedvision-language-actionmodelsviaworldknowledgevariation.arXiv preprint arXiv:2603.08361,.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling 𝛿VLA:Prior-guidedvision-language-actionmodelsviaworldknowledgevariation.arXiv preprint arXiv:2603.08361,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.664989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.664989Z digest=sha256:e111957bea53cc6dfaaa360a5f161770c352f9386e42bbfcf959ecb275aa394a

Observation dabfba20-dc68-409a-b791-160f285c5780 · outbound

This paper cites We use AdamW with𝛽= (0.9,0.95) ,𝜖=10 −8, and weight decay10−8, under a cosine schedule with linear warmup.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling We use AdamW with𝛽= (0.9,0.95) ,𝜖=10 −8, and weight decay10−8, under a cosine schedule with linear warmup

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:56.158365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:52:55.669408Z digest=sha256:1b078d27b036922a54733e9a5561fe51b8fbc38da95800d048bcd86c678e9dad

Observation 101c7be5-a322-4587-99dc-11055849eb3c · outbound

This paper cites an unresolved cited work.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:52:56.142199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:52:55.674025Z digest=sha256:4d14d673fa249105516c96e278f7b715477093c325c2d717da0dd8b2bfcb7c6f

Observation e2b5cf19-cc19-471b-90f2-8f11b8fcc6a7 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling WorldVLA: Towards Autoregressive Action World Model

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.613388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.613388Z digest=sha256:0b2a96d222928f1bbc4f4fade6c57c7860ed27e4e297ee60ac6fe1b6b47988ee

Observation 99eccc8d-ad0a-41e6-8d94-65296ad4ed31 · outbound

This paper cites Robotic VLA benefits from joint learning with motion image diffusion.arXiv preprint arXiv:2512.18007,.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Robotic VLA benefits from joint learning with motion image diffusion.arXiv preprint arXiv:2512.18007,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.618677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.618677Z digest=sha256:01eaa60e45ddbcc94356cb7f96fe976ba623f822a1bfa4f929c0c7fbebead57b

Observation 3cfe59c5-a037-4b0e-95d7-ee56f79def5b · outbound

This paper cites Causal World Modeling for Robot Control.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Causal World Modeling for Robot Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.623005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.623005Z digest=sha256:7347e594363da13d27ef476edce715dc49f50acf65c1fda162ae2fc6ec4b0ad8

Observation 70551851-8387-412a-b9da-29d0052a4284 · outbound

This paper cites DiT4DiT: Jointly modeling video dynamics and actions for generalizable robot control.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling DiT4DiT: Jointly modeling video dynamics and actions for generalizable robot control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.628035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.628035Z digest=sha256:e4017c9903281bc743c82d2950709b6f5a3f85d28f89452a6dd13e50a0ece3ac

Observation bd1375fc-bb65-4083-8e81-a160f95fffbd · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.632302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.632302Z digest=sha256:723a96c1e96bfd31cee2dc668dc5adaabed7e75761ce58297f78c97956348670

Pith citing papers

No inbound Pith citation observations are available.