Pith. sign in

Paper Citation Record · LEDGER

Masked Visual Actions for Unified World Modeling

As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2607.19343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19343 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:49:09.834338Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e144fd38-614d-4d79-abd4-e49972c8b038 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:57.894759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:57.894759Z digest=sha256:109321b30bdfb5057d871c752902dcacc66db09a2dd28e1d9cbb31a9f51bde10

Observation e153cc02-ae9a-4090-93c7-4e9c44efb71b · outbound

This paper cites Unifying (Machine) Vision via Counterfactual World Modeling.

Masked Visual Actions for Unified World Modeling Unifying (Machine) Vision via Counterfactual World Modeling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.164826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.164826Z digest=sha256:161cf32929cfa8c525345c38b5bc68d365076fafe8914c9e81cf9dd8fc6708f8

Observation d43bb27a-7e9e-4a6c-8454-03d2bb8fd1d9 · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation, 2024.

Masked Visual Actions for Unified World Modeling Track2act: Predicting point tracks from internet videos enables generalizable robot manipulation, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.324753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.324753Z digest=sha256:dd42a8e729dbe38d07b6991267fe23bd6b55c6c11bea4c4f24a2382a1d2b974e

Observation c396244a-b33e-42f5-bf6f-64d27e5ee62d · outbound

This paper cites Video generation models as world simulators.

Masked Visual Actions for Unified World Modeling Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.544990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.544990Z digest=sha256:ef4c83fe2e5e84e242151b847ba67ac69f8e13238cf5ad7ac6a1ac48669c46c9

Observation 92e1cc1d-2170-4645-a044-ddedd9816a4f · outbound

This paper cites Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise.

Masked Visual Actions for Unified World Modeling Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.718932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.718932Z digest=sha256:286e506fd31bad25c9a8e9351b371828a1ec20b8e10810ce077abdd151bcf04f

Observation a06037db-a4bd-4f57-ba84-c9faefb0c543 · outbound

This paper cites SAM 3: Segment anything with concepts.

Masked Visual Actions for Unified World Modeling SAM 3: Segment anything with concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:58.894834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:58.894834Z digest=sha256:fcab02a237663f56eaaba1c412796140e5d846961552fb064b6bed96296a6b81

Observation 7fe0aad0-dfcd-4a31-b296-27d8a4bf916c · outbound

This paper cites Freeman, Jitendra Malik, Russ Tedrake, Vincent Sitzmann, and Yilun Du.

Masked Visual Actions for Unified World Modeling Freeman, Jitendra Malik, Russ Tedrake, Vincent Sitzmann, and Yilun Du

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.064864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.064864Z digest=sha256:f9f7bd91f10e05107ccfdbe7607f6c365f06543f40595812a498a27fa425338b

Observation 26bcfb5d-e03a-4987-8e76-c0ed0e00c3e9 · outbound

This paper cites Learning coordinated bimanual manipulation policies using state diffusion and inverse dynamics models.

Masked Visual Actions for Unified World Modeling Learning coordinated bimanual manipulation policies using state diffusion and inverse dynamics models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.254754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.254754Z digest=sha256:873fa82cdad624455b39c9066f1fc0f650f39503669bae08a1b424a12801088d

Observation 43dbe581-b2c1-463d-9d58-6c9a81cf636a · outbound

This paper cites Tool-as-interface: Learning robot policies from observing human tool use.

Masked Visual Actions for Unified World Modeling Tool-as-interface: Learning robot policies from observing human tool use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.497836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.497836Z digest=sha256:075745aa30d048c0c66b29e58f059028bffb04a4609c015388bbeb0e667cdb9a

Observation bf9c6eef-c02e-458a-9ffa-3599ee0fdc41 · outbound

This paper cites Bridgev2w: Bridging video generation models to embodied world models via embodiment masks.arXiv preprint arXiv:2602.03793, 2026.

Masked Visual Actions for Unified World Modeling Bridgev2w: Bridging video generation models to embodied world models via embodiment masks.arXiv preprint arXiv:2602.03793, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:48:59.894801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:48:59.894801Z digest=sha256:7c37d09b05fbb11bb45c8062b96e5cf47da870bd8bb2f453433dd47dc5c815bd

Observation 1702a4b4-9cc9-495d-96b9-1e3d7d5bd874 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Masked Visual Actions for Unified World Modeling Diffusion policy: Visuomotor policy learning via action diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.023922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.023922Z digest=sha256:f6162f26a1b3c5bb6cf898c2eb88ba1a4f9bbda6d5e0cb8910ca51cfd9138bf1

Observation ffcbb56b-0b93-4263-b34a-091724c31d4d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Masked Visual Actions for Unified World Modeling Diffusion policy: Visuomotor policy learning via action diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.102142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.102142Z digest=sha256:9ccbeab396612a803a3e8c4e2eebde3a0439c2e21322b388759537cdd8acb5cc

Observation 918e5046-42ee-4c96-8962-40cdc95f6386 · outbound

This paper cites Wan-move: Motion-controllable video generation via latent trajectory guidance.arXiv preprint arXiv:2512.08765, 2025.

Masked Visual Actions for Unified World Modeling Wan-move: Motion-controllable video generation via latent trajectory guidance.arXiv preprint arXiv:2512.08765, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.184836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.184836Z digest=sha256:bae88d40dff262bb51df11c1f90cee84edddb6a1cda79457fd51cdfaa004bb8f

Observation 5561a96c-a2c6-4294-87fa-ef81884764bc · outbound

This paper cites Embodis- wap for zero-shot robot imitation learning, 2025.

Masked Visual Actions for Unified World Modeling Embodis- wap for zero-shot robot imitation learning, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.324749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.324749Z digest=sha256:2d5144c850c36a313042813dfc5de477b918c8b7e888a7aca1f688bb416a514c

Observation 5747730c-af1d-45c0-8fff-4325a4508157 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

Masked Visual Actions for Unified World Modeling BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-01T12:49:00.475099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.475099Z digest=sha256:b258d134d68b64061c353c33dc488a8c27aa669d3c9f3706b516a9df5b46005d

Observation 8a7ed1bd-fb04-45ef-9469-912c526c3b19 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Masked Visual Actions for Unified World Modeling Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.624831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.624831Z digest=sha256:52cac62d742401170a61a9e8d3433bad0e7605b5c483e2bc69cf46bb97d6d558

Observation 49ffcd05-92a8-48b5-9401-c21ca9b33844 · outbound

This paper cites Video language planning.

Masked Visual Actions for Unified World Modeling Video language planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.749493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.749493Z digest=sha256:756e128398467b1057d1512f2c92277a73e6ff8889e8331d4e1688d63a460fff

Observation f3f80707-501b-4a75-81ad-66c7d1ae9ecc · outbound

This paper cites Aim: Intent-aware unified world action modeling with spatial value maps, 2026.

Masked Visual Actions for Unified World Modeling Aim: Intent-aware unified world action modeling with spatial value maps, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:00.868780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:00.868780Z digest=sha256:212a06e3581428cca25c27a557d1921f173f8decead45525f98b4a2a1e3e2288

Observation 218a1261-205c-49d1-b6d7-4c2ac1f53fab · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

Masked Visual Actions for Unified World Modeling DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.042267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.042267Z digest=sha256:229f29d8f224fe572f5dc69f36f47a426aef193071b048c8065cb660ed7aa604

Observation 9c26af47-6a88-4cbd-a16c-2d2cdd90d647 · outbound

This paper cites Motion prompting: Controlling video generation with motion trajectories.

Masked Visual Actions for Unified World Modeling Motion prompting: Controlling video generation with motion trajectories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.159661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.159661Z digest=sha256:6f0e09b139286666e9ab66d2d845ef96bb3f1f05d11a3d9d891c2e0ab16f2217

Observation 36716550-3d81-459c-b6ac-995455b6fc26 · outbound

This paper cites Force prompting: Video generation models can learn and generalize physics-based control signals.

Masked Visual Actions for Unified World Modeling Force prompting: Video generation models can learn and generalize physics-based control signals

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.443907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.443907Z digest=sha256:8cd9014968d1f646ff2308477a1515b921a0c70672dbfa1263d2e2869cf5c35a

Observation f7e4cc6c-b7b1-4171-a1c9-16586381e5f6 · outbound

This paper cites Goal force: Teaching video models to accomplish physics-conditioned goals.

Masked Visual Actions for Unified World Modeling Goal force: Teaching video models to accomplish physics-conditioned goals

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.584937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.584937Z digest=sha256:e44b282224b3b8edfbc7ad764332e36836dd1316029d6b6e76594177b3a38e41

Observation 9cfc4af1-0d8f-4a72-934d-6001de896cfb · outbound

This paper cites World models for learning dexterous hand-object interactions from human videos.arXiv preprint arXiv:2512.13644, 2026.

Masked Visual Actions for Unified World Modeling World models for learning dexterous hand-object interactions from human videos.arXiv preprint arXiv:2512.13644, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.704765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.704765Z digest=sha256:4023c222c94d78f5c4b0ddcc43e1e920abd83532716f061f27b7bce93b7ade85

Observation c74a524d-0ea2-4d14-b1f0-f7a703a5db5b · outbound

This paper cites Unified 4d world action modeling from video priors with asynchronous denoising, 2026.

Masked Visual Actions for Unified World Modeling Unified 4d world action modeling from video priors with asynchronous denoising, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:01.844740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:01.844740Z digest=sha256:3298cc273dea0198bff25ccbba0c83351b6cdfeac184481dfad46c857d98147d

Observation befed282-d539-46a9-8cfe-ff39cbb91804 · outbound

This paper cites Ctrl-world: A controllable generative world model for robot manipulation.

Masked Visual Actions for Unified World Modeling Ctrl-world: A controllable generative world model for robot manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.035375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.035375Z digest=sha256:1f12da661259927301b7b86c658aa0fa342ab1c264b14227355b325ec801a48f

Observation d01eb595-33ff-422c-98e1-b41d4659e08c · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations, 2024.

Masked Visual Actions for Unified World Modeling Video prediction policy: A generalist robot policy with predictive visual representations, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.110836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.110836Z digest=sha256:99eb453358c6b9508cf416332fdb7bf36bb5d7dd4a24fa1ab19c5995d2c89659

Observation 7a64bf78-7d54-4a87-920f-a83b8a4d6bb0 · outbound

This paper cites Vid2world: Crafting video diffusion models to interactive world models, 2025.

Masked Visual Actions for Unified World Modeling Vid2world: Crafting video diffusion models to interactive world models, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.216476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.216476Z digest=sha256:72faac122c51d55c8ae8eac231fe6b188fcaa8e6c6341ea881f1decc6fefc01b

Observation 26522a78-c911-4297-b9e6-0d427f2dde09 · outbound

This paper cites Pointworld: Scaling 3d world models for in-the-wild robotic manipulation.

Masked Visual Actions for Unified World Modeling Pointworld: Scaling 3d world models for in-the-wild robotic manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.362222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.362222Z digest=sha256:18942b9bfbb614c713d2a23642d2927e2b6278de84e6f9b4c49410f6ce187a07

Observation b4d59e25-62ab-47a1-82c2-9c5f5436eb9c · outbound

This paper cites Dreamgen: Unlocking generalization in robot learning through video world models, 2025.

Masked Visual Actions for Unified World Modeling Dreamgen: Unlocking generalization in robot learning through video world models, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.461021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.461021Z digest=sha256:63416bc1ba92a3ac4280452cb2be0233384bc16964e1b55bed9e585dc3d96fc4

Observation 3929f513-5414-4dd7-a112-24134706778f · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024.

Masked Visual Actions for Unified World Modeling Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.585802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.585802Z digest=sha256:ecd501dc7dbee5d7cdb0374233b3a006bf2aed44c74506048c3558a35c2442cf

Observation 9b1a1e0b-06b9-41dc-81a7-b2570c4306ce · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.742391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.742391Z digest=sha256:2fd0daf5c0f1d13507654af4f2eaa7e3d45c7c7eda869c8be59f3bd9ff0d42ad

Observation 2f03d7b5-a577-4f6b-a1df-fc7892c172f2 · outbound

This paper cites Dexterous world models.

Masked Visual Actions for Unified World Modeling Dexterous world models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.843637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.843637Z digest=sha256:c131726b7278b8efc54cede9d90dcf1b2055ee32809dc82ba03cc6b32efb8a09

Observation 25ed980f-2583-4209-a8ce-bfa49fa23bc8 · outbound

This paper cites Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026.

Masked Visual Actions for Unified World Modeling Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:02.952171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:02.952171Z digest=sha256:1b7a1bb2349759385c6f844ed9d114c61eeefaa3c6d9d4be233a99d3d43fc3ae

Observation d97b8c95-b61e-4c8d-854f-b13625b305e5 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Masked Visual Actions for Unified World Modeling Learning to Act from Actionless Videos through Dense Correspondences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.044032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.044032Z digest=sha256:2ae7031d7a33a151506b68f3e956d9233c346dbe64bffded139aff8d6889aa0d

Observation 9fe16f9f-8f3c-4a56-86a0-ac38483916fd · outbound

This paper cites World Modeling with Probabilistic Structure Integration.

Masked Visual Actions for Unified World Modeling World Modeling with Probabilistic Structure Integration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.136308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.136308Z digest=sha256:77730441cebdea43b21e1ed2e6fd66aea1599fbe8bfdeb2957888fd4a5781f93

Observation 52ae5854-5dda-4b1d-8ecd-77042020424b · outbound

This paper cites Shadow: Leveraging segmentation masks for cross-embodiment policy transfer, 2025.

Masked Visual Actions for Unified World Modeling Shadow: Leveraging segmentation masks for cross-embodiment policy transfer, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.214635Z digest=sha256:3a24ce3ac63beb32b614bccc1db8ddd4bf469fc4c01b6ae4416233e17bea7d09

Observation a9afdf0e-3d47-4ea8-ac4a-73c7f631f9d4 · outbound

This paper cites Masquerade: Learning from in-the-wild human videos using data-editing, 2025.

Masked Visual Actions for Unified World Modeling Masquerade: Learning from in-the-wild human videos using data-editing, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.274841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.274841Z digest=sha256:66030a9d51dbbe15388946562adcc36864a8a9cb14ab4cb6186cc609c683826b

Observation a8ce6fbf-8a79-44b1-839b-0219aec38154 · outbound

This paper cites Phantom: Training robots without robots using only human videos, 2025.

Masked Visual Actions for Unified World Modeling Phantom: Training robots without robots using only human videos, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.364897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.364897Z digest=sha256:255912bd706aef647ed34633f20b7536f330db5b4e7748f245a7431bce9b30c3

Observation 14ccab1a-39b2-44ea-961e-92fbae02a8c0 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Masked Visual Actions for Unified World Modeling BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.472167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.472167Z digest=sha256:84d096a2982668eb03bd6491fcc481603964fc0050bdd634cb1edcac985f6dec

Observation 64a8cd3a-318b-48f5-a5bd-864f78f11963 · outbound

This paper cites Mask2iv: Interaction-centric video generation via mask trajectories, 2025.

Masked Visual Actions for Unified World Modeling Mask2iv: Interaction-centric video generation via mask trajectories, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.614750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.614750Z digest=sha256:4d8c0d41360d0a67a2168665e7dcef4ddf72bcea4615e26d8bbbb66bc574d5b3

Observation 00decb5e-3f9b-4070-bd75-a0d881ff15b5 · outbound

This paper cites Novaflow: Zero-shot manipulation via actionable flow from generated videos, 2025.

Masked Visual Actions for Unified World Modeling Novaflow: Zero-shot manipulation via actionable flow from generated videos, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.715409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.715409Z digest=sha256:b4c5ee2587aabf7152bdcb424c828418f31aebdcf17987805c056f6f15e0f25c

Observation 2cd019cd-cb77-4e3b-8b63-0d3b591854f2 · outbound

This paper cites Unified video action model, 2025.

Masked Visual Actions for Unified World Modeling Unified video action model, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.844866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.844866Z digest=sha256:b47bdb14b1d922b6e4645ad65e58bc6cc2f6b2e494d45319aba663dbf1ed009a

Observation e3a97ca1-f025-4352-8f6b-92a27f8b75c0 · outbound

This paper cites Genie envisioner: A unified world foundation platform for robotic manipulation, 2025.

Masked Visual Actions for Unified World Modeling Genie envisioner: A unified world foundation platform for robotic manipulation, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.969261Z digest=sha256:48f92da5c0b364118df4d800b7d0118b333cf5c654a335c049a39966fcbbe664

Observation 7074bee0-8bfb-406c-bb61-2d156196acaf · outbound

This paper cites Realwonder: Real-time physical action-conditioned video generation, 2026.

Masked Visual Actions for Unified World Modeling Realwonder: Real-time physical action-conditioned video generation, 2026

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.146958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.146958Z digest=sha256:9e398d82c33cc41c8b7e42a19eeb23c61e9a8ec7b166675793e61294ee8efbb0

Observation aeda653a-28e4-49e1-8fb2-551c6dd612fb · outbound

This paper cites Zero-shot world models are developmentally efficient learners.arXiv e-prints, pages arXiv–2604, 2026.

Masked Visual Actions for Unified World Modeling Zero-shot world models are developmentally efficient learners.arXiv e-prints, pages arXiv–2604, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.313578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.313578Z digest=sha256:859d029c1705d8e9b3c573ac682428b065ac3ca062afce37f4f5b3828d31257e

Observation 0a3626b9-18bd-4032-8342-7b75e430cd94 · outbound

This paper cites Decoupled Weight Decay Regularization.

Masked Visual Actions for Unified World Modeling Decoupled Weight Decay Regularization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.440100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.440100Z digest=sha256:7245ec336f3d9fc7fa7813b18118204999bb0879667a6d6881ab0bcefa2ea6fc

Observation 53ca9292-a352-415c-867c-1d1d2f99feca · outbound

This paper cites Mask world model: Predicting what matters for robust robot policy learning, 2026.

Masked Visual Actions for Unified World Modeling Mask world model: Predicting what matters for robust robot policy learning, 2026

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.517419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.517419Z digest=sha256:9373b7d4db2c3fab5d448a2ffe46ef139904912b5b03cd6cce0e0ec145dff943

Observation 8d819bb7-73ec-48cb-a4fb-3a1113f2eabc · outbound

This paper cites Inference-time scaling for diffusion models beyond scaling denoising steps.

Masked Visual Actions for Unified World Modeling Inference-time scaling for diffusion models beyond scaling denoising steps

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.669230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.669230Z digest=sha256:53ee5d2465db1c61e12f54d0c1b2338bc933fb394b4a58f1e3f7214f98efc31f

Observation 1797b1cc-201e-402b-903f-e5c8efb61ce4 · outbound

This paper cites MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations.

Masked Visual Actions for Unified World Modeling MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.803449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.803449Z digest=sha256:c6565d12ca195381c17dca8a150a3aa6475040cf7a6660b5e58f77a153700791

Observation ba782044-7c09-43b3-ac79-7443413b524d · outbound

This paper cites Motubrain: An advanced world action model for robot control, 2026.

Masked Visual Actions for Unified World Modeling Motubrain: An advanced world action model for robot control, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:04.967163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:04.967163Z digest=sha256:0d4199f4b8f5bd2061fce5d4b45a902684329e4703c0796662747cbf5cf840fc

Observation 9418a5d4-53e5-41d5-a965-49c5dfd9f8fc · outbound

This paper cites s1: Simple test-time scaling.

Masked Visual Actions for Unified World Modeling s1: Simple test-time scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.141732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.141732Z digest=sha256:86d5dcb9fae0566996204012213d24d9afcce938f2eeb58f2d1975d06c16ba77

Observation b115d4ba-4cf2-4981-a1d7-77bdd298e8ce · outbound

This paper cites Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.

Masked Visual Actions for Unified World Modeling Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.213696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.213696Z digest=sha256:95d0b6df4c3995d3c5662e8fdae4841acfd9fa60dae2b06f380474e488d6efcc

Observation 66f9331b-073c-4c7f-b8e1-b45d435d366e · outbound

This paper cites Cosmos world foundation model platform for physical ai, 2025.

Masked Visual Actions for Unified World Modeling Cosmos world foundation model platform for physical ai, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.313907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.313907Z digest=sha256:a5c9962f2ce0f59420dd3bda9803dbec28c3fdda6cd907b2a9e9f3fbe20a5d53

Observation 21d004a3-c5b7-49c1-8d44-8eee52cd6107 · outbound

This paper cites mimic-video: Video-action models for generalizable robot control beyond vlas, 2025.

Masked Visual Actions for Unified World Modeling mimic-video: Video-action models for generalizable robot control beyond vlas, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.491244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.491244Z digest=sha256:f3438fe539cdf6f99cb50bc3eb2ab424c8180c6c9ed6f3fcf4afc8228c0083fa

Observation dd72077b-bc3b-4629-93d2-a02b9ada1f64 · outbound

This paper cites Inference-time enhancement of generative robot policies via predictive world modeling.IEEE Robotics and Automation Letters, 2026.

Masked Visual Actions for Unified World Modeling Inference-time enhancement of generative robot policies via predictive world modeling.IEEE Robotics and Automation Letters, 2026

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.647880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.647880Z digest=sha256:a8c77e0a287306b4366dfecc7a3953f3e2c8af78b12b3f7e1046526ff5f4b9f3

Observation 71b4d916-03dc-4871-b8a2-c00044d164e1 · outbound

This paper cites MotionStream: Real-Time Video Generation with Interactive Motion Controls.

Masked Visual Actions for Unified World Modeling MotionStream: Real-Time Video Generation with Interactive Motion Controls

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.770198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.770198Z digest=sha256:61e0e5e4bbc17ad5f1c63bfbd9d2f3f393b1fece8a29e4d88eff0d33e80697d7

Observation 1ec0c14c-6f66-4171-a489-5ad8b747535b · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Masked Visual Actions for Unified World Modeling SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:05.914908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:05.914908Z digest=sha256:9c456b991961196e75226a2c013bd4499b2ce02fbf19bbe6d9106a24ce7cb2bd

Observation 11d1fc99-ed4d-4540-8640-45f8b08d9725 · outbound

This paper cites Time-to-move: Training-free motion-controlled video generation via dual-clock denoising.

Masked Visual Actions for Unified World Modeling Time-to-move: Training-free motion-controlled video generation via dual-clock denoising

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.115706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.115706Z digest=sha256:bb7a57647fed43432d207044b9af2e8387e978d4b6dc81bb04f2d51b56771951

Observation 7af91829-ba8c-458f-a05f-24b492076649 · outbound

This paper cites Motion before action: Diffusing object motion as manipulation condition, 2024.

Masked Visual Actions for Unified World Modeling Motion before action: Diffusing object motion as manipulation condition, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.248811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.248811Z digest=sha256:609f6d028001ccf382654ff6f9a424334117c432555ec09a4c61ebb919c39744

Observation 5fcd5563-ee5d-4d2f-8558-01d98ef167a4 · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator, 2025.

Masked Visual Actions for Unified World Modeling Evaluating gemini robotics policies in a veo world simulator, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.418490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.418490Z digest=sha256:761e42a92b30fcc2929118148c6a18845daf62f1b76674c631e4585b1b2740d7

Observation 3bd46ac0-233f-4c9d-b740-f7251e10d84d · outbound

This paper cites Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026.

Masked Visual Actions for Unified World Modeling Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.557565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.557565Z digest=sha256:3169a1bebd6aaf606be844ff87dae8e2931cc7f621ff1c1444c17f8033130b02

Observation eec87124-0133-4b11-8def-0642d0540e5a · outbound

This paper cites Understanding physical dynamics with counterfactual world modeling.

Masked Visual Actions for Unified World Modeling Understanding physical dynamics with counterfactual world modeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.659959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.659959Z digest=sha256:9aa22617eec30259b6df3b152660db97c5287f40e47b27c82421ba17ecba163b

Observation 62940561-bdaf-4541-b98a-a626f1390d4b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Masked Visual Actions for Unified World Modeling Wan: Open and Advanced Large-Scale Video Generative Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.746750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.746750Z digest=sha256:b689271e2eb1f65fab9f37d46c41085c3948422fd4fb8da0ac2cc95e169a2d58

Observation 9b9c6d2d-e9a9-4417-8cb2-57c7a9b151d2 · outbound

This paper cites Eva: Aligning video world models with executable robot actions via inverse dynamics rewards, 2026.

Masked Visual Actions for Unified World Modeling Eva: Aligning video world models with executable robot actions via inverse dynamics rewards, 2026

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.915628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.915628Z digest=sha256:072f139be8d8862df57b79628884f251755d31c0dbd1977107dad46fdc02a800

Observation c3dcb25d-8b7a-455e-bf13-25ecb7dc3cbe · outbound

This paper cites Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026.

Masked Visual Actions for Unified World Modeling Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:06.965342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:06.965342Z digest=sha256:e1ebd7fdb6a5f9bcb92c9728ba1baff959de4623c42dc14799b7a73a7abf0617

Observation 741604b0-4a2d-4094-a57e-0a50123789e9 · outbound

This paper cites Precise action-to-video generation through visual action prompts.

Masked Visual Actions for Unified World Modeling Precise action-to-video generation through visual action prompts

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.047797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.047797Z digest=sha256:9055c86ec7bca6d0cfa5979b2ed7287790b87ff0be6ce743df90b03f674230c0

Observation f7cdd53a-fa44-48ef-960c-1a816e2d4875 · outbound

This paper cites Wolpert and J.

Masked Visual Actions for Unified World Modeling Wolpert and J

Reference 67

Resolution
verified exact
doi, observed 2026-08-01T12:53:42.293862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-01T12:49:07.189078Z digest=sha256:69e41bf4f5891f0365b06d197c0f9c40cc61ae03b285adb0f0b9b9c68f7ace55

Observation 85aa13da-c5b9-4d46-801a-221179b3d2c5 · outbound

This paper cites Wolpert, Zoubin Ghahramani, and Michael I.

Masked Visual Actions for Unified World Modeling Wolpert, Zoubin Ghahramani, and Michael I

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.307228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.307228Z digest=sha256:eed6ee8cb0addbfa571f725c0099817021a9d10a675993463143152f14de3da9

Observation cf8c8a53-9b1c-4fbf-9340-eae84e13fda8 · outbound

This paper cites Wolpert, R.

Masked Visual Actions for Unified World Modeling Wolpert, R

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.401627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.401627Z digest=sha256:f5d97628b684ca8b98b0e3dcc691570a4f77baedd7e68d328c9ba566833d663f

Observation 99715a3f-d73f-40a5-a1ab-86533f20e61c · outbound

This paper cites Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein.

Masked Visual Actions for Unified World Modeling Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.439272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.439272Z digest=sha256:971cac279865bf7fcc031d965b07163f76de19e9f5983d90a51c0a05b70d8378

Observation f5d9321b-9944-43d7-abc5-f5748404f8fb · outbound

This paper cites Kinema4d: Kinematic4d world modeling for spatiotemporal embodied simulation.arXiv preprint arXiv:2603.16669, 2026.

Masked Visual Actions for Unified World Modeling Kinema4d: Kinematic4d world modeling for spatiotemporal embodied simulation.arXiv preprint arXiv:2603.16669, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.518129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.518129Z digest=sha256:ffc7fafcfa0940fa5388ff6b2c774d0472c6d04edcd2a477926f2e9169704a8b

Observation 8cfdc1d7-e718-41b2-9505-ac8e1d8a3d17 · outbound

This paper cites RoboPanoptes: The All-Seeing Robot with Whole-body Dexterity.

Masked Visual Actions for Unified World Modeling RoboPanoptes: The All-Seeing Robot with Whole-body Dexterity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.719051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.719051Z digest=sha256:92ab7eb596a5b65f572caf89670cb687760c0f0536b77d38bba73ac80c6c297c

Observation ab6ab918-592d-4345-85df-311e9e6c4ab7 · outbound

This paper cites Learning Interactive Real-World Simulators.

Masked Visual Actions for Unified World Modeling Learning Interactive Real-World Simulators

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.854932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.854932Z digest=sha256:41466f98c8b6e6a61e364f6da140645804040abbf0a6d0f6f0e41d7a5668ceab

Observation 84c4a1d6-799e-470f-8c8a-c27e4f9caa3e · outbound

This paper cites Orv: 4d occupancy-centric robot video generation, 2025.

Masked Visual Actions for Unified World Modeling Orv: 4d occupancy-centric robot video generation, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:07.975262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:07.975262Z digest=sha256:54719b69968397f9b9a9a5559fd3d810afa7c01d7b489b43a421ddab31c914e6

Observation a8e5f30f-c093-4a77-8085-dcaeba235004 · outbound

This paper cites Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions.

Masked Visual Actions for Unified World Modeling Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.072132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.072132Z digest=sha256:bdbbcad35f06762bfbaf0d46a3328c89b09a1f7938721e0ae57b04e75a7f537e

Observation 84f6297c-83a5-4b2f-ac76-2c38fd5fab04 · outbound

This paper cites Veo-act: How far can frontier video models advance generalizable robot manipulation?, 2026.

Masked Visual Actions for Unified World Modeling Veo-act: How far can frontier video models advance generalizable robot manipulation?, 2026

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.184918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.184918Z digest=sha256:d2ecbbcdd16806e78ed9c32f5b7e20cc09bf2136adf70001d87a9dce01a37c54

Observation 7fccbe33-c8ae-4be6-a725-5274c7345b24 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Masked Visual Actions for Unified World Modeling Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.273965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.273965Z digest=sha256:d6bd74ea2ade566088ca89242940dc6040561600d19e66107e58ab4dadd5eab1

Observation 6c349569-4e4e-44f4-a7b6-6acc18cc9ca9 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Masked Visual Actions for Unified World Modeling TesserAct: Learning 4D Embodied World Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.342623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.342623Z digest=sha256:5c4e0fcfcc22b5d46612b862b6b97a87d394285c282dc0e6070fd77aab2003bb

Observation 8d9b496d-c710-48b7-8b9d-8339d9ea3ca8 · outbound

This paper cites Action images: End-to-end policy learning via multiview video generation, 2026.

Masked Visual Actions for Unified World Modeling Action images: End-to-end policy learning via multiview video generation, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.422235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.422235Z digest=sha256:cfdbcba6c7cf8eec69da4321dfeb7f40a1f432090732ca75c42cc0f3a0192ccd

Observation d5264796-98d5-4805-88bd-ef0430366b8d · outbound

This paper cites 3dflowaction: Learning cross-embodiment manipulation from 3d flow world model, 2025.

Masked Visual Actions for Unified World Modeling 3dflowaction: Learning cross-embodiment manipulation from 3d flow world model, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.526653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.526653Z digest=sha256:d2268b8d8821c4c38bf6a38a57179b946d559e0a0d9741e7be1721965ca3416b

Observation 02ec8abe-6c1a-4de1-8892-851c0f196470 · outbound

This paper cites Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets, 2025.

Masked Visual Actions for Unified World Modeling Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.630493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.630493Z digest=sha256:86ee03e6eab1699df5c416876ec5f75dbeb1db00d0c6a36ad4398e622c8351f9

Observation 89f044fc-98f3-4c8f-ace7-cb3e24366785 · outbound

This paper cites Irasim: Learning interactive real-robot action simulators.

Masked Visual Actions for Unified World Modeling Irasim: Learning interactive real-robot action simulators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.710108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.710108Z digest=sha256:1ee32dc41883173f19e39592d667a407c069b75010c822cc86e1bc942ad32c11

Observation 5ab4aa41-44c7-4018-a44f-70630b5e1320 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:08.889648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:08.889648Z digest=sha256:04bb58d4ab86c0c860fcc6b0256d0a7499b7f15e6e039da38fcd0a4a2f81ae59

Observation bf97c2d8-397f-4f77-9d92-8dc248a27450 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.033921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.033921Z digest=sha256:f34983c2e557fb16f1d3450bb98938b6ab905e31e0bebcb0aa65f579a80b29b7

Observation 9b682be1-bf24-422b-a688-e78410cff85f · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.150745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.150745Z digest=sha256:b679438ff318469537149a160c6581d8a2b691f7abdde29783b52f7169b71e16

Observation fe024360-1b51-40aa-a64c-a06c41f6aca3 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.279685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.279685Z digest=sha256:b47282edb8dd9aaacdb3d959789086cc063ce09d3fcea89a0ee22e510726e470

Observation 95f7312e-19d5-42d6-94c0-7ad8072974d4 · outbound

This paper cites Covers Diffusion Policy, ACT, and SmolVLA baselines.

Masked Visual Actions for Unified World Modeling Covers Diffusion Policy, ACT, and SmolVLA baselines

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.427171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.427171Z digest=sha256:62669bf0225b93f0317e785ceae59d9d501d0fb3bff748bf40665cb1a64f778d

Observation 21b2c77b-ec9e-4e58-ab39-bdd522ca28a7 · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.512352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.512352Z digest=sha256:53072272618fe635d29835f932259023481ea221c5723a825aa92dbb9f879856

Observation 40ca33fc-e5b8-44ee-b02f-2fe35df2e30d · outbound

This paper cites an unresolved cited work.

Masked Visual Actions for Unified World Modeling Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.557487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.557487Z digest=sha256:4d981d9dc7ca90e20e941aa07cbbebe4e37539699e17817f948b182aa90d1487

Observation 2d75b15b-efbb-4d84-bf59-8a6507ff0f7c · outbound

This paper cites the toaster door must be flush, latched, and closed by a push from below.

Masked Visual Actions for Unified World Modeling the toaster door must be flush, latched, and closed by a push from below

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.642077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.642077Z digest=sha256:b6568cfe6322dc26762e7fcdeef62b7c240e77db626b893d35ca089f134a8f13

Observation ddfca20e-1655-4c33-ac83-8b263a5b037e · outbound

This paper cites robot_pushed.

Masked Visual Actions for Unified World Modeling robot_pushed

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.706619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.706619Z digest=sha256:fa18cb5495fa53ff1c6f74d02eac6f035d9bdc1df66676dd28e1868270097ade

Observation fdc01ae2-3c49-4693-b456-1c166504bdcb · outbound

This paper cites Outcomes reached after disengagement, via ghost contact, teleport, or autonomous motion are FALSE.

Masked Visual Actions for Unified World Modeling Outcomes reached after disengagement, via ghost contact, teleport, or autonomous motion are FALSE

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.751858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.751858Z digest=sha256:70ab13954431dc73b79bd93a45deebf784ed82d416d80bb196876eca4660c4f5

Observation a6af23b8-78ba-4f05-af0b-73a5d0d03040 · outbound

This paper cites Hovering near.

Masked Visual Actions for Unified World Modeling Hovering near

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.789865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.789865Z digest=sha256:4a3528c30025039c493443c94687a83682f1a4a9a7620b0a19fbde154a27aadf

Observation 7ab1c18a-7e90-414e-93a6-5b710927f094 · outbound

This paper cites Penalize teleporting, morphing, gripper passing through solids, vanishing or duplicated objects, frame-to-frame 18 jumps, ghost contact, post-disengagement coasting.

Masked Visual Actions for Unified World Modeling Penalize teleporting, morphing, gripper passing through solids, vanishing or duplicated objects, frame-to-frame 18 jumps, ghost contact, post-disengagement coasting

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:09.834338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:09.834338Z digest=sha256:494cff96946f2e34d68c2dd4c714fca19a4746a9e8c9ecbb3474d5f49246d5a2

Pith citing papers

No inbound Pith citation observations are available.