Pith. sign in

Paper Citation Record · LEDGER

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

As of 21 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2608.08839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08839 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:01.485780Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c4cdf7b-882d-406c-b4ef-208ea7513bda · outbound

This paper cites Motus: A Unified Latent Action World Model.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Motus: A Unified Latent Action World Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.187503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.187503Z digest=sha256:6a88c65fed870f8738d00d5b3cf9d416f8888d3e9ff56e0bb61fb1143203caed

Observation 45bff85b-e0d3-4f0d-a057-d9580aacb76a · outbound

This paper cites π0.5: a vision-language-action model with open-world generalization.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models π0.5: a vision-language-action model with open-world generalization

Reference 3

Resolution
verified exact
doi, observed 2026-08-14T04:27:01.535013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:27:01.213845Z digest=sha256:1a08c690ccffd4e0fe5b4be688df784d036b025da07c5fa8a4dbbef4f3257d6e

Observation 94a0bf8e-e965-4e35-bea9-f4996e576e80 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models WorldVLA: Towards Autoregressive Action World Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.246456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.246456Z digest=sha256:ac0f172925208d58092a387e72e6bbccc09521848c7e940bbbde636fdac66b5e

Observation 45301073-61e7-4430-a514-1d0e169dd403 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.271186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.271186Z digest=sha256:4ac6ac6383bcd7db60e328463558fd503bdf4257d2305bd0808576c6e5004409

Observation 41ef3d8a-b30a-4df4-a51e-6c1d3f278632 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models LTX-Video: Realtime Video Latent Diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.307886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.307886Z digest=sha256:37abe9b7594dc608540e9463b7d2a0a3917067c17a682df11b65ac46758b5a4d

Observation 65d62c5e-809e-4f91-b6a6-4f2e9e8bfd07 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.313720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.313720Z digest=sha256:ff1ef3005db55f7474b79964457c52bf2b6ee5d78f769e937181ad6bad5c5318

Observation 6c097634-f4f9-47eb-a59c-167333b337e8 · outbound

This paper cites Rynnvla-001: Using human demonstrations to improve robot manipulation.arXiv preprint arXiv:2509.15212,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Rynnvla-001: Using human demonstrations to improve robot manipulation.arXiv preprint arXiv:2509.15212,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.318910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.318910Z digest=sha256:c6f2266b0eca866ba7e3d4fc31158f3756ca34df163ad7a804d2cdcd9877ec7a

Observation 5647a053-0c1a-4c39-bb99-f593781aed6e · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-14T04:27:01.322000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.322000Z digest=sha256:fb8f25f9565c7cc7137913abea6dc3a514ca41b85f2893a29c8b107bcec3eb22

Observation 6f83a792-6cb8-4ded-8659-693c7593dfbd · outbound

This paper cites Unified Video Action Model.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Unified Video Action Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.332559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.332559Z digest=sha256:28993b9d6b9b5fb1ba3bbdfca484fcdb510e20887087a50e10dcadff647a3264

Observation f5fa2be7-2580-4e88-9baa-cdbd2005f8e1 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.336990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.336990Z digest=sha256:cc29c5376b55983906864615ba5d0d6a60a2ac1a28a86baf9e6525e17abbd74c

Observation d0d199e1-abd0-4e4c-80c8-412f57e1ed56 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Depth Anything 3: Recovering the Visual Space from Any Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.346136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.346136Z digest=sha256:063751b25717738f4419ad057d2300a24b1a34e9d58037f5e9707f173ddf06a1

Observation 73a90e82-f850-41f5-a373-264846f5716c · outbound

This paper cites OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.356892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.356892Z digest=sha256:8c38470032073cdae578e4ecce650d1bfbfdab66b5eed0eea88fdbec1c5c4764

Observation 416ca311-f861-4115-88b2-4b45c2ff0c0c · outbound

This paper cites Mask World Model: Predicting What Matters for Robust Robot Policy Learning.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Mask World Model: Predicting What Matters for Robust Robot Policy Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.365387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.365387Z digest=sha256:4f5da9ae7f5256bd6beac2f51c3c4b407e241c59d0abea0b1f33c1d7bfb04484

Observation 4f851e50-8e62-45e1-a400-2bce99d00a32 · outbound

This paper cites Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.378434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.378434Z digest=sha256:0fe0b11cd5770899a710c56982209d4d500e17525fd80f989bc99567d4b97bd9

Observation df5a5daf-c291-4eb0-8c32-af1a286e8981 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.382735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.382735Z digest=sha256:0ec6c7a329c43b4964d1a2fefa371a61aae191024adcf05fb0734e69689b1391

Observation 98d7f6f8-d680-4a1a-a310-45b741db355f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.415201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.415201Z digest=sha256:b29b02290bdaf171eb465c0096f718cff946dcb5ef55091233f278d464e8e999

Observation 27bcc08b-6877-4a20-ad7d-935a2ad0db66 · outbound

This paper cites S-VAM: Shortcut video-action model by self-distilling geometric and semantic foresight.arXiv preprint arXiv:2603.16195,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models S-VAM: Shortcut video-action model by self-distilling geometric and semantic foresight.arXiv preprint arXiv:2603.16195,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.436760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.436760Z digest=sha256:a46dd5d740009b464cddf57ef251008e84c2295a21c36671b46d12b91a70fd87

Observation 0c20bc1a-e422-464d-878c-7855c0698104 · outbound

This paper cites MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.450632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.450632Z digest=sha256:3dc0421460b101bc13d35bb0168ea4ad940961238f13c9dbfec0a39112652f3f

Observation f19e0cf9-3b5d-42c8-8b9a-e98d39b885b3 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.455192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.455192Z digest=sha256:f6c36c4072a00d601a8dc83998d16e1043111baa3fb8206227bc2e8803ceaa52

Observation ea6ecb71-d1cd-4df9-b9cb-08a5ea330434 · outbound

This paper cites FlowVLA: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models FlowVLA: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.462371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.462371Z digest=sha256:54f1d7cd6bbbcabc8593b3eca57cf23c0e908ba980f59f20510791a35d86d47f

Observation 01a340b4-28fb-446a-a0ba-9d9f01afd41b · outbound

This paper cites DualCoT-VLA: Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.arXiv preprint arXiv:2603.22280,.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models DualCoT-VLA: Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.arXiv preprint arXiv:2603.22280,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.467787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.467787Z digest=sha256:261f1f7f90c90791411c02c73d6df183337130ecee40820e964d144eedd5f9a7

Observation c3d68a9d-5b10-45b2-81f9-802e6f4f655e · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.479963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.479963Z digest=sha256:ba387250d2eb2334d18127613f113f06af9342075125007ac413b579ced01233

Observation e357fa73-6940-4db3-bb28-3a4ac4759dd9 · outbound

This paper cites DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:27:01.558768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T04:27:01.485780Z digest=sha256:f45d8be40ed8ab87b380e24813b7016f45747f46c42e43cdaea0988407001a86

Observation 6a50ebe8-ed3b-405d-baff-8aefda09b651 · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.402929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.402929Z digest=sha256:712a423dc69fc14c1c3afcce36622a77f1543af3b05843358cb0cb492c2526cb

Observation 7cca0c34-2788-4da8-bdce-8b05e6dbe381 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.238908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.238908Z digest=sha256:8f71abc0cea3d7f9504c5159589e3845db30624713f6be30c86299370931b01b

Observation b1b0234c-48e2-4d0c-8b5f-f201e595275f · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.290804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.290804Z digest=sha256:6b6df6d1828073af5fba53d1625e07c28766d93550a8f14fc5eaa419d8d244f2

Observation 5e761f15-8856-4ccf-af6c-764da0415a20 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.206906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.206906Z digest=sha256:599853627d748655cf848e82c1458e47efaf2887cea1cbc20906968d6bce1690

Observation 3c97199d-61b9-4aae-a5e6-e6f88464c71c · outbound

This paper cites Causal World Modeling for Robot Control.

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models Causal World Modeling for Robot Control

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:01.327148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:27:01.327148Z digest=sha256:cbca02867e3f0e0593ad599c21dfb5a4a8f2e4815db7539b7dd5bbd9c9969c8d

Pith citing papers

No inbound Pith citation observations are available.