Pith. sign in

Paper Citation Record · LEDGER

Native Video-Action Pretraining for Generalizable Robot Control

As of 15 August 2026, this Paper Citation Record lists 100 of 134 outbound references and 8 inbound Pith citation observations for arXiv:2607.08639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08639 v2

Coverage vector

measured 100 of 134 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:53:43.109347Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:52:22.229506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:52:36.169573Z

Reference resolution

100 of 134 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4791112d-ba06-4fd9-a343-451533b727d5 · outbound

This paper cites 1x world model: From video to action.

Native Video-Action Pretraining for Generalizable Robot Control 1x world model: From video to action

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:30.767598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:30.767598Z digest=sha256:867e1ae60fbd5ee76c4b4588c06e4d60ef4b5aa1f6494e33318b9dfd37c446d7

Observation 951a422a-d594-4079-b942-ae85311b4a1c · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Native Video-Action Pretraining for Generalizable Robot Control AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:30.909133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:30.909133Z digest=sha256:1ca2a7720f88714fa1e3a8ec7d287ba4507425129d77aff74cafe94657356aff

Observation bbc95e00-a818-413f-93c3-19e73856b4fa · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Native Video-Action Pretraining for Generalizable Robot Control Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.101633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.101633Z digest=sha256:91cb763f6feafd3fa372b10f5ee1055701bcd56f262b37aeadd4c12353502ad0

Observation 2cd86200-62d3-4548-8ef3-4d81f8605793 · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

Native Video-Action Pretraining for Generalizable Robot Control One transformer fits all distributions in multi-modal diffusion at scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.222980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.222980Z digest=sha256:db88ab98e6855ed342e62942a5f146d9abe1e84b7591de0c5076c6631e1233e4

Observation dbdb3ee5-9578-4c59-9487-76347294ddd6 · outbound

This paper cites Navigation world models.

Native Video-Action Pretraining for Generalizable Robot Control Navigation world models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.412644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.412644Z digest=sha256:83fc1fa8eeec6209402030c055ac6d0ceffd1a5546d3f64c4973888c1d8e7963

Observation 661be493-7e1b-4b4d-ac65-af3f8d91cd85 · outbound

This paper cites A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation.

Native Video-Action Pretraining for Generalizable Robot Control A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.587834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.587834Z digest=sha256:63789399d04667d612ecec822b1c30add089e512ee80a3f20d1563cd029a7150

Observation 7fb6b613-98dc-4318-ad27-8fa5f472e534 · outbound

This paper cites Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation.

Native Video-Action Pretraining for Generalizable Robot Control Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.701876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.701876Z digest=sha256:2b8e3a6e854cb00dfa462dbfee774898062ed0062e5ad0cf0ab3d0eb0c589f08

Observation 4b290628-1729-40e6-8e40-8b147ad68d9a · outbound

This paper cites Motus: A Unified Latent Action World Model.

Native Video-Action Pretraining for Generalizable Robot Control Motus: A Unified Latent Action World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.813726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.813726Z digest=sha256:635d5c9f019a2173936ce0627dc26706c1aaef7a43e1084019291427372e6e6b

Observation 8fe55f92-0a21-4d27-91a8-dc5ecaa0ae52 · outbound

This paper cites π0: A vision- language-action flow model for general robot control.

Native Video-Action Pretraining for Generalizable Robot Control π0: A vision- language-action flow model for general robot control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:31.925337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:31.925337Z digest=sha256:709a49c714184d8e0b74f218c1e7ca857edcf6601815cbb85344fee681b4eee1

Observation 70bc8eaa-17af-421d-9d0e-aecc328444be · outbound

This paper cites Perception encoder: The best visual embeddings are not at the output of the network.

Native Video-Action Pretraining for Generalizable Robot Control Perception encoder: The best visual embeddings are not at the output of the network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.103130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.103130Z digest=sha256:f282e7602bf83d3b84c68d8e33a5a21079d51567d4fba68c0a58c6190b583c92

Observation 173b02f3-5044-43e0-8b79-d68bc036b666 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Native Video-Action Pretraining for Generalizable Robot Control Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.228184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.228184Z digest=sha256:429506e5f7603ce9c174934e4f859678f0b4420b9987a1f6e2ade3d9df355324

Observation f63ad169-2c0f-458d-a837-9b4586a991ac · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

Native Video-Action Pretraining for Generalizable Robot Control Rt-1: Robotics transformer for real-world control at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.353839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.353839Z digest=sha256:7fa6e880e3b198113dcfcb34b283518cd2466d7435bf43869a01c23e897f4c55

Observation df20c81a-5b77-48c3-9385-90bd0e439def · outbound

This paper cites Genie: Generative interactive environments.

Native Video-Action Pretraining for Generalizable Robot Control Genie: Generative interactive environments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.457036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.457036Z digest=sha256:1b40948cb5e1f4afe6e1b8dfc9ca23941c7911971bad9751b234be287f9e6110

Observation 78ba5ec1-d789-4736-964b-e4f74c9096c2 · outbound

This paper cites Univla: Learning to act anywhere with task-centric latent actions.

Native Video-Action Pretraining for Generalizable Robot Control Univla: Learning to act anywhere with task-centric latent actions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.539325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.539325Z digest=sha256:5d3367e865b271782de221e71a08dda78234000d66709ad0fd399a6d504a703b

Observation 4f057200-c315-4afa-a0f6-e928bcf940eb · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

Native Video-Action Pretraining for Generalizable Robot Control WorldVLA: Towards Autoregressive Action World Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.635405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.635405Z digest=sha256:c5e16f45a528d7027914e1e75032b1dacf2ac010f24f6207fc5737f55b1c85a6

Observation b5e47953-cdfe-4354-959a-6ff388a7f1de · outbound

This paper cites Gamegen-x: Interactive open-world game video generation.

Native Video-Action Pretraining for Generalizable Robot Control Gamegen-x: Interactive open-world game video generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.762417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.762417Z digest=sha256:f7c3bef64cdcef7c9f7a9bf68f36cc9db345a700a004a9bd6066d0891fc96424

Observation c542595a-b5d2-4e11-bb8d-d93f4c9c48fa · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Native Video-Action Pretraining for Generalizable Robot Control GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:32.972380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:32.972380Z digest=sha256:ec0fc56d3fd7a9a08baea4a4c359dd148962653f0f151708d85fea11f0dd4e38

Observation 87e76254-e6e1-4f9c-af6d-42d2924ec169 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

Native Video-Action Pretraining for Generalizable Robot Control Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.208987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.208987Z digest=sha256:ae632d78686fa5b50c14bd4b186a4513ee46392c2536fb4efb23b358e09512da

Observation 48a53e59-c3b6-4167-a10d-6352e791a5b2 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Native Video-Action Pretraining for Generalizable Robot Control RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.250157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.250157Z digest=sha256:6938eec58fe57c66e8fd351a15465ff2eafcb3a7673de0c16476efdee29ff00b

Observation 68cd1546-61f5-4e03-ae66-2cb1a4983959 · outbound

This paper cites Moto: Latent motion token as the bridging language for learning robot manipulation from videos.

Native Video-Action Pretraining for Generalizable Robot Control Moto: Latent motion token as the bridging language for learning robot manipulation from videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.343247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.343247Z digest=sha256:a45e388ab54789583ea32a95b90eea4295b8cd0f46490d5ff40f1abeae788241

Observation 5c0b6438-1063-420d-8efa-d80e8f2c297b · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Native Video-Action Pretraining for Generalizable Robot Control Diffusion policy: Visuomotor policy learning via action diffusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.455581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.455581Z digest=sha256:56eade5b0b5e39a8c4cf18c03df905c9208641354660393800e7bfb920d5116e

Observation 0956d37e-3d76-4bdd-9777-5cfe6aa96cf6 · outbound

This paper cites Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots.

Native Video-Action Pretraining for Generalizable Robot Control Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.555398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.555398Z digest=sha256:b645d13f9d14a2a06b608866255373c96504e82812bcd92ee77308219aa2d1ec

Observation e3355251-ecad-4f39-9a14-17c6708576a6 · outbound

This paper cites an unresolved cited work.

Native Video-Action Pretraining for Generalizable Robot Control Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.632405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.632405Z digest=sha256:b3de78176e229a1ce11e61f4e9244e14ed0b5673c4cda0dbf0e16c996c5b8bd0

Observation db9634c2-8b14-4b90-a752-343020cac002 · outbound

This paper cites DeepSeek-V3 Technical Report.

Native Video-Action Pretraining for Generalizable Robot Control DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.701654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.701654Z digest=sha256:7e828a3473fb15b9217f14c6e91ab9822ff99a137f6b81a2a0700258f8823160

Observation 5ee6ee22-5c0f-454c-9818-256c2d565ee4 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Native Video-Action Pretraining for Generalizable Robot Control An image is worth 16x16 words: Transformers for image recognition at scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.811579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.811579Z digest=sha256:e4d066bfaa99e50c08ba41f1ae6c8491c9781754edbba93df8d2d37b585fbcaf

Observation 67203d73-3e83-4168-a274-55deb0870641 · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

Native Video-Action Pretraining for Generalizable Robot Control Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:33.949923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:33.949923Z digest=sha256:0d08e171c32ced96936f257ba1ee07474dc65e63407bf87e0d1c4a6827670da7

Observation a8ee8f03-e6fe-4179-9d38-634cb7936ed2 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Native Video-Action Pretraining for Generalizable Robot Control Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.096363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.096363Z digest=sha256:a14e7e2c2d27b2b0b2778a3f95bb24c1d1460140e8112830f0052356d460cc7a

Observation 02601879-5d48-4577-aa02-056de2bb3ac6 · outbound

This paper cites Adaworld: Learning adaptable world models with latent actions.

Native Video-Action Pretraining for Generalizable Robot Control Adaworld: Learning adaptable world models with latent actions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.281580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.281580Z digest=sha256:42ef77576fb4307fc7c003fae2897faa8bc027fdcbd3e060cf8c33af51954a41

Observation 31414557-3b5a-48df-9539-53361712d2d5 · outbound

This paper cites Infinite Worlds with Versatile Interactions.

Native Video-Action Pretraining for Generalizable Robot Control Infinite Worlds with Versatile Interactions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.387887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.387887Z digest=sha256:fabd12db4dcca662e284473f698ea9a88cf206942bf41ff8176e5c4dbdbc4e81

Observation 5594e99b-04cb-4d29-b8b6-ce12e8f8bc34 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Native Video-Action Pretraining for Generalizable Robot Control Gemini Robotics: Bringing AI into the Physical World

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.471049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.471049Z digest=sha256:f34e3d0126ced05e6e419c4e378d7dfda86bf56b127271818fb0d03024aba828

Observation 377447f0-60c6-49b1-a65e-ea7a9b662a0a · outbound

This paper cites Gen-0: Embodied foundation models that scale with physical interaction.

Native Video-Action Pretraining for Generalizable Robot Control Gen-0: Embodied foundation models that scale with physical interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.624359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.624359Z digest=sha256:d93bb1823350cc631e07e814a7891b79ef853242fc2dc1ba74fe34a87fc11a7a

Observation d456176a-001f-4816-9675-8dc7d9e58a8e · outbound

This paper cites Veo: A text-to-video generation system.Google DeepMind Technical Report, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Veo: A text-to-video generation system.Google DeepMind Technical Report, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.717834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.717834Z digest=sha256:8cbcd96ce395a5684d05781efb139e96586dfff8174c2b7768d42caa4b6dbea6

Observation 11107363-8409-4790-985b-b5408eb670ea · outbound

This paper cites Mastering diverse control tasks through world models.

Native Video-Action Pretraining for Generalizable Robot Control Mastering diverse control tasks through world models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.783414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.783414Z digest=sha256:11ae1588a202ea40beccd59344cb16a1e4fa7bd6db6ae72d2c217cd8d691b510

Observation e531a95a-a05c-4509-856e-d9c1561ec51d · outbound

This paper cites Td-mpc2: Scalable, robust world models for continuous control.

Native Video-Action Pretraining for Generalizable Robot Control Td-mpc2: Scalable, robust world models for continuous control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:34.850606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:34.850606Z digest=sha256:b2b8d7ae5a4832ca709626f417aefba9f3715ffad0dc53964ac8717589499c9c

Observation 71dad1a9-48dc-4050-8fb1-e035209393e8 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations.

Native Video-Action Pretraining for Generalizable Robot Control Video prediction policy: A generalist robot policy with predictive visual representations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.067360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.067360Z digest=sha256:30d1627339a4d6e15cffc4258e7aed8eaaeb8eb09d70319cfc6d00acf6fb72eb

Observation f00659d7-765a-4b10-9aaa-7fb2c145c71a · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.203877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.203877Z digest=sha256:554107a4af54f53063f99638c8af8d7151f9e03fbd377884facdb6b989b7c2db

Observation 0c1cec87-a1c5-4d1f-b45b-3c269cfdd711 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

Native Video-Action Pretraining for Generalizable Robot Control Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.349274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.349274Z digest=sha256:5e4e454474f3c1bb63dfc901f9c6a4e8c712bb00bf1bc210f2fad9667f0872f9

Observation 838f6c65-afa4-4ae8-ac10-d5ad566ea989 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

Native Video-Action Pretraining for Generalizable Robot Control NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.510558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.510558Z digest=sha256:4e19830ee512add9004292027feacb2b6921ae9e2ec443ad20420bfac9e5aae2

Observation a8b627f8-66fe-4e03-93e9-30a69acfd977 · outbound

This paper cites Dreamgen: Unlocking generalization in robot learning through video world models.

Native Video-Action Pretraining for Generalizable Robot Control Dreamgen: Unlocking generalization in robot learning through video world models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.630848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.630848Z digest=sha256:f5928e1f712b401c58421eb5ab6405b87342a70913aee9328b8842ea6cacf7dc

Observation cd018f2c-8dce-4082-9ca8-512938fd69e7 · outbound

This paper cites EgoMimic: Scaling Imitation Learning via Egocentric Video.

Native Video-Action Pretraining for Generalizable Robot Control EgoMimic: Scaling Imitation Learning via Egocentric Video

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.817019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.817019Z digest=sha256:89b7fbea4803984060a202c05b2b14b9f02e1bef12add8a090accc840544e56b

Observation 3715110f-923c-4af3-a946-8688ce375141 · outbound

This paper cites Droid: A large-scale in-the-wild robot manipulation dataset.

Native Video-Action Pretraining for Generalizable Robot Control Droid: A large-scale in-the-wild robot manipulation dataset

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:35.956486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:35.956486Z digest=sha256:bc78f5456d7be52ce10562e0377e2d01b5d9774381a5281dd8f4ea71e4afa91d

Observation e3cd8c94-7709-436b-9103-d8fbb27268a2 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

Native Video-Action Pretraining for Generalizable Robot Control Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.030424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.030424Z digest=sha256:f12b31cefb76643a898f62085a68807aad4d03a375dbeac03e4cc54e3f61d430

Observation db877293-a587-4338-9c94-88fb5b0c5fd2 · outbound

This paper cites Openvla: An open-source vision-language-action model.

Native Video-Action Pretraining for Generalizable Robot Control Openvla: An open-source vision-language-action model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.126332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.126332Z digest=sha256:ece6ba7ed8cdefe7ca785d5cb04603f957f09cab3008394b854712afc9ed8b2f

Observation b9ac1f17-7b7c-4590-b875-da27069294d9 · outbound

This paper cites Auto-Encoding Variational Bayes.

Native Video-Action Pretraining for Generalizable Robot Control Auto-Encoding Variational Bayes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.240349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.240349Z digest=sha256:48f10bdb018cd6d71b181bf923616bbae25f19324cff3dc4d4dcc09cef0fbd56

Observation 35457b15-2c84-4d68-a52d-b6b55a148913 · outbound

This paper cites Kling-v3.https://kling.ai/, 2026.

Native Video-Action Pretraining for Generalizable Robot Control Kling-v3.https://kling.ai/, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.418771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.418771Z digest=sha256:761a83573f4bd1255df0e5e3d95bb555b526558ec678c15cd477ffa2f92be261

Observation 3010070b-ad76-4dde-bf44-81e9e50c57c4 · outbound

This paper cites Tenenbaum.

Native Video-Action Pretraining for Generalizable Robot Control Tenenbaum

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.629307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.629307Z digest=sha256:538d9d2a0e050a76f41de20e76e996d2f569ca3ea13420aa01688c5afa701385

Observation 08d8643c-362b-4d1c-94a0-8a462344ede3 · outbound

This paper cites Deformnet: Latent space modeling and dynamics prediction for deformable object manipulation.

Native Video-Action Pretraining for Generalizable Robot Control Deformnet: Latent space modeling and dynamics prediction for deformable object manipulation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.772489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.772489Z digest=sha256:bcf70bf0b836700546601c427336b71f418357aff0c9397b71d92514752541af

Observation 7b850f8d-3d07-4bdb-99cf-d413de54096e · outbound

This paper cites Cronusvla: Transferring latent motion across time for multi-frame prediction in manipulation.arXiv preprint arXiv:2506.19816, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Cronusvla: Transferring latent motion across time for multi-frame prediction in manipulation.arXiv preprint arXiv:2506.19816, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:36.911336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:36.911336Z digest=sha256:23c0a8a97dec7fb2b159512079b5033157180dac05027114bd4328f2f5ce04c8

Observation 136ccabc-bd75-4dd4-9462-ccba1e9d1ac1 · outbound

This paper cites GR-3 Technical Report.

Native Video-Action Pretraining for Generalizable Robot Control GR-3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.043552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.043552Z digest=sha256:a1b4de07e3bd97bed33400a97454c6af5801115ed9a95e91bb2f2d68f5c7a3e0

Observation adf25337-212b-40a8-8500-0a261b840b74 · outbound

This paper cites Causal World Modeling for Robot Control.

Native Video-Action Pretraining for Generalizable Robot Control Causal World Modeling for Robot Control

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.121799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.121799Z digest=sha256:bddb51502e8ff8a12ea575b1bd5477ecd05a77b5e701496266a386eef3585d4d

Observation 26f40d42-32cd-4570-a5cd-295f3d141a61 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Native Video-Action Pretraining for Generalizable Robot Control CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.201100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.201100Z digest=sha256:f38d0c523a3d8a9b0d6d42610f626fbba7061c0781bded963347720ad04146b5

Observation 2bff0788-78ef-4e41-b829-aaaea7032dde · outbound

This paper cites What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?.

Native Video-Action Pretraining for Generalizable Robot Control What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.282017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.282017Z digest=sha256:b82728bf32ff33a4d8cab42a1589a3ac4a9dbb6358f5c632062ced6eec87f579

Observation 50efdff0-8528-4d59-887c-44622e32f2f4 · outbound

This paper cites WALL-WM: Carving World Action Modeling at the Event Joints.

Native Video-Action Pretraining for Generalizable Robot Control WALL-WM: Carving World Action Modeling at the Event Joints

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.373566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.373566Z digest=sha256:c8b7a3ce0038d44cc2256463458a790104ee5666e9703c80cd6a50d350b2fa39

Observation 079318c0-66b6-42d7-a2ce-cc657409964b · outbound

This paper cites Unified video action model.

Native Video-Action Pretraining for Generalizable Robot Control Unified video action model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.474825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.474825Z digest=sha256:59ae7c969c36ad14b26c597977592f832eb4a763c8a70222bb0a64d63947bdf7

Observation b4d043cd-2f67-4947-abbc-d27ff2982657 · outbound

This paper cites Vision-language foundation models as effective robot imitators.

Native Video-Action Pretraining for Generalizable Robot Control Vision-language foundation models as effective robot imitators

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.538592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.538592Z digest=sha256:d170c520127a1f8a51fb25aadcf420e293e0c488bf6ba45ff556fab241c248e7

Observation 6550940c-3de2-4731-baa7-c75d2de1b457 · outbound

This paper cites Propagation networks for model-based control under partial observation.

Native Video-Action Pretraining for Generalizable Robot Control Propagation networks for model-based control under partial observation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.600551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.600551Z digest=sha256:efb681ac913cde215d0eda51553c0014492fa6d8f23130a922d693782b037625

Observation 213a5310-37e9-4a4e-8dfe-ace2187eca81 · outbound

This paper cites Dreamitate: Real-world visuomotor policy learning via video generation.

Native Video-Action Pretraining for Generalizable Robot Control Dreamitate: Real-world visuomotor policy learning via video generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.703610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.703610Z digest=sha256:43954e0aa05e0fc27fa5f7fadb44d6aede6fcd74e4760a347ef33948e03b266e

Observation c3544f1b-a169-4bf6-bf36-0b2eddcd995a · outbound

This paper cites Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.Transactions on Machine Learning Research, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.Transactions on Machine Learning Research, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.774009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.774009Z digest=sha256:14cabfb8eebc9e2992705737ed1b252fea23ffe5bf3d09490e858d106b542025

Observation 7157e131-888b-40db-80e5-1266bbd3d160 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

Native Video-Action Pretraining for Generalizable Robot Control Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.845066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.845066Z digest=sha256:a582ae8e939b0587354309e97a4d2fe19aa855d4278b2cf667c4255c19c3e078

Observation 989faa58-9c6e-4867-b430-1024be949b9d · outbound

This paper cites Zero-wam: In-context world modeling for zero-shot task generalization.

Native Video-Action Pretraining for Generalizable Robot Control Zero-wam: In-context world modeling for zero-shot task generalization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:37.953851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:37.953851Z digest=sha256:b3e3444af0ab59d6f3e26fa24da86e6f9b4206cdcfb9717e34d8711d46d1d587

Observation 70d988ae-88d0-4c34-b461-6f8c0527ef4c · outbound

This paper cites an unresolved cited work.

Native Video-Action Pretraining for Generalizable Robot Control Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.034197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.034197Z digest=sha256:220f689241d2f7fe586ceca198fe4f5d575f3e8ea2bf624179bc4621a20f59e3

Observation 9d0f956f-6e18-4743-857a-b1275f1c237c · outbound

This paper cites Efficient robotic policy learning via latent space backward planning.

Native Video-Action Pretraining for Generalizable Robot Control Efficient robotic policy learning via latent space backward planning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.149544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.149544Z digest=sha256:1fde8766f046e07a323290ff90b41073d4be6eb7d3a4445d28e64b3ffe447e3e

Observation 42e57d54-56ca-4ae5-8241-ddf09ec27bc3 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Native Video-Action Pretraining for Generalizable Robot Control HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.340680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.340680Z digest=sha256:39157e2aa40f2e4ac43bd91dfc3829f027dafac587f2b120bc3e8ae27cdf7b6c

Observation d0b91c59-d9fc-401a-b559-bf219fad0eac · outbound

This paper cites Muon is Scalable for LLM Training.

Native Video-Action Pretraining for Generalizable Robot Control Muon is Scalable for LLM Training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.463939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.463939Z digest=sha256:d1a112c04a40872251a3173068fdb4729a74f100b1cfb4075cef106ab099bab7

Observation 80357a37-92ed-404c-9954-b500f57a5e60 · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

Native Video-Action Pretraining for Generalizable Robot Control Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.609411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.609411Z digest=sha256:3e3280f4a47ab63c09ff13569e781376b8709b166960c4dc95caac4bd4bf9b16

Observation f742ffaf-db65-4417-84e8-dc83c1b193ce · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.754872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.754872Z digest=sha256:1d385370116c23c4e5d2bc6f0bce34946196adb70ea68120930f0cec33172b60

Observation 31e1959f-f8f6-4323-b3f4-67cbba79ad08 · outbound

This paper cites Being-H0.7: A Latent World-Action Model from Egocentric Videos.

Native Video-Action Pretraining for Generalizable Robot Control Being-H0.7: A Latent World-Action Model from Egocentric Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.870441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.870441Z digest=sha256:e773894ee922a55ce5b335c2e0f361828d050a54da401f23f8683b2a9a9bc7d3

Observation 016127eb-9c9c-4361-9303-2beffbf4ce19 · outbound

This paper cites Deep learning for universal linear embeddings of nonlinear dynamics.

Native Video-Action Pretraining for Generalizable Robot Control Deep learning for universal linear embeddings of nonlinear dynamics

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.007360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.007360Z digest=sha256:b5a7d7f97f57bff05720dabc98a34411a82a050bb4eb30472aaabe3f8ff08a25

Observation 82731db9-f58d-4d75-abb7-ea7391f4bdd7 · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

Native Video-Action Pretraining for Generalizable Robot Control F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.114561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.114561Z digest=sha256:cf4e005264733b469906fd3994a063328a8c7d103e42787d202a4ae4e6ed09d8

Observation a760df59-1c6b-4b7a-b322-088f2b59a3c9 · outbound

This paper cites Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence.

Native Video-Action Pretraining for Generalizable Robot Control Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.262642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.262642Z digest=sha256:099cf6280dbd4652f16fe048119b4778aa07bc0a7422ecc0f7ec46f3d6ba2b8e

Observation 7f7965a7-7177-4bf5-86d3-14e3cfc22207 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Native Video-Action Pretraining for Generalizable Robot Control GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.357615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.357615Z digest=sha256:f079ec9c1af3db344eed659b3401d5637a789c1201cc3be81bd6cf8011c7b018

Observation db996ed5-1aae-4989-a311-ac1f516cfd3b · outbound

This paper cites Octo: An open-source generalist robot policy.

Native Video-Action Pretraining for Generalizable Robot Control Octo: An open-source generalist robot policy

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.613696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.613696Z digest=sha256:c0257268240a8f37b543b1807ae1d0ce368b395a06243da2b68b542481d14f55

Observation df05b6f2-a795-4cc9-84fc-fc67e77e1147 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Native Video-Action Pretraining for Generalizable Robot Control Open x-embodiment: Robotic learning datasets and rt-x models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.782758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.782758Z digest=sha256:db0b0459cd3ed6a17fa156b18d5018890e35023c3c4f1830df0b9357cb5122e9

Observation a57a7258-0d4c-4acc-809d-7563ea038bfd · outbound

This paper cites Video generation models as world simulators.OpenAI Technical Report, 2024.

Native Video-Action Pretraining for Generalizable Robot Control Video generation models as world simulators.OpenAI Technical Report, 2024

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.940731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.940731Z digest=sha256:3c0e38233d5ae286744acc5c556fd51b731d8a0db1ceaa2a4da16db01657c7e7

Observation 3b474e31-cf5c-4a69-9e3c-6971a874a2d4 · outbound

This paper cites Genie 2: A large-scale foundation world model.

Native Video-Action Pretraining for Generalizable Robot Control Genie 2: A large-scale foundation world model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.075561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.075561Z digest=sha256:cc3c2a111b0704517273531b6d3f8cf21a85f41605d348905c07cbdcdc314d41

Observation 620f651e-dbd9-44f5-904a-03fc0f8512ea · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

Native Video-Action Pretraining for Generalizable Robot Control ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.223709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.223709Z digest=sha256:44aec9e7468a07c752094946b8c7ab7f098f92b80cb1b040b1473a61b50f8243

Observation 067d523a-a812-44f4-9827-541fb04706b0 · outbound

This paper cites an unresolved cited work.

Native Video-Action Pretraining for Generalizable Robot Control Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.273393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.273393Z digest=sha256:bc138285aca39f7e7776e13380ad583268562d2d88364865781f2192c0106121

Observation 8239e981-044c-45e4-b994-364b48dccf44 · outbound

This paper cites Mv-umi: A scalable multi-view interface for cross-embodiment learning.arXiv preprint arXiv:2509.18757, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Mv-umi: A scalable multi-view interface for cross-embodiment learning.arXiv preprint arXiv:2509.18757, 2025

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.361416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.361416Z digest=sha256:70b58c3b53dbd1fffca36f41b00eaf3a54985d49f018ae19946e04978705ed0c

Observation 51c7d500-7bf0-4d62-b789-b5fab20716c3 · outbound

This paper cites Advancing Open-source World Models.

Native Video-Action Pretraining for Generalizable Robot Control Advancing Open-source World Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.455544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.455544Z digest=sha256:1096aaf59d7d5de27d79c8e74ed834cb652d64c1c22cada76cc1574d802957e2

Observation bbd73e22-516e-466e-8214-5e87c11c6397 · outbound

This paper cites Learning to act without actions.

Native Video-Action Pretraining for Generalizable Robot Control Learning to act without actions

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.536693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.536693Z digest=sha256:00200dafc8b6382de88e1434a88cc415297fba8a1dbd5c8664c6e6932ad152bd

Observation 81493de0-104e-4a03-94b3-386f92f95158 · outbound

This paper cites Action-conditional implicit visual dynamics for deformable object manipulation.The International Journal of Robotics Research, 43(4):437–455, 2024.

Native Video-Action Pretraining for Generalizable Robot Control Action-conditional implicit visual dynamics for deformable object manipulation.The International Journal of Robotics Research, 43(4):437–455, 2024

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.616470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.616470Z digest=sha256:cf2594a181f47e89b6ea14b322b69a7a67e4aeaf8c63225095f9bc2f8737120c

Observation 8dd98248-6ba5-47c1-a330-e251b58143f3 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

Native Video-Action Pretraining for Generalizable Robot Control MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.697488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.697488Z digest=sha256:189781d8b4913b9595d4deecce077600ae43957d2b536ea2b832b0f836d9fd70

Observation 5050c9b9-fbab-4fe4-a951-5af4a2183d50 · outbound

This paper cites Robocraft: Learning to see, simulate, and shape elasto-plastic objects in 3d with graph networks.The International Journal of Robotics Research, 43(4):533–549, 2024.

Native Video-Action Pretraining for Generalizable Robot Control Robocraft: Learning to see, simulate, and shape elasto-plastic objects in 3d with graph networks.The International Journal of Robotics Research, 43(4):533–549, 2024

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.781791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.781791Z digest=sha256:4bc19ca8e15d03a628bc2b70e359337e544f07f36fda16bc5f446a0de24ce4cc

Observation 7743bd04-6843-40a3-9be4-6d6e366a7489 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Native Video-Action Pretraining for Generalizable Robot Control Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.859548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.859548Z digest=sha256:0435eba0f9f62e41a3e42d7b97652885e692df72e6a2037b42cb7e2cfe123a07

Observation dc66e10a-b886-4ec7-800f-8a1b7940c302 · outbound

This paper cites Consistency models.

Native Video-Action Pretraining for Generalizable Robot Control Consistency models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.958158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.958158Z digest=sha256:f6787e08b3e7348d021a7304bcfde13bb771d3aadcfeaf89e014b867074968ce

Observation 83a27aaa-6b36-4ba9-9c6c-000e835946ed · outbound

This paper cites Memer: Scaling up memory for robot control via experience retrieval.arXiv preprint arXiv:2510.20328, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Memer: Scaling up memory for robot control via experience retrieval.arXiv preprint arXiv:2510.20328, 2025

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.099524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.099524Z digest=sha256:d352960b85e104a97d0531d06ce2e3b2005a4cc011d9dbf393b2cbf46b59cc57

Observation d9b72329-d27c-4458-bc99-b92b35339719 · outbound

This paper cites Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs.

Native Video-Action Pretraining for Generalizable Robot Control Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.215461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.215461Z digest=sha256:fa54e1db7f82c676a32d5b63b87e6a3fcad16f76d426c00737eb3db47bffde8f

Observation 5e5fb8d2-8c54-482b-b4e8-a7a0de2aaa46 · outbound

This paper cites Application of a particle-in-cell method to solid mechanics.Computer physics communications, 87(1-2):236–252, 1995.

Native Video-Action Pretraining for Generalizable Robot Control Application of a particle-in-cell method to solid mechanics.Computer physics communications, 87(1-2):236–252, 1995

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.358203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.358203Z digest=sha256:cde5ccd17f5f9018fdb61d1b822b7ac434ad2e74a11badab60874b0eddcda19d

Observation ee6b3130-daee-486f-8e24-78cc2faf161c · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.485457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.485457Z digest=sha256:b974257b3d7cf48f18fa7b28f4752592ffccb2d1ffde5ee452ef639a54756f8c

Observation 8ddefbd4-669a-4e69-97ff-ebd79ff697a2 · outbound

This paper cites Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026.

Native Video-Action Pretraining for Generalizable Robot Control Causal video models are data-efficient robot policy learners.Rhoda AI Blog, 2026

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.677791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.677791Z digest=sha256:456922e640afb4ee50705e48139f4799c2b3cd1a72a45f4e893e3743617bc31c

Observation 26cdb36a-6763-4cef-a2ae-90a695e11eba · outbound

This paper cites Gemini 3.1 pro: A smarter model for your most complex tasks.

Native Video-Action Pretraining for Generalizable Robot Control Gemini 3.1 pro: A smarter model for your most complex tasks

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.823564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.823564Z digest=sha256:5f1bf9713a88805ea87832a6f19b55d676dae1f54b3c8c89638dae55a6aeac21

Observation 58245a4f-cd90-404d-87ad-35e7a1ab0f69 · outbound

This paper cites Predictive inverse dynamics models are scalable learners for robotic manipulation.

Native Video-Action Pretraining for Generalizable Robot Control Predictive inverse dynamics models are scalable learners for robotic manipulation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:41.975646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:41.975646Z digest=sha256:09f741e94609162329328aaf81381597e45166648d98d913f2cb98533b03e69b

Observation 6d550953-393d-4d40-8911-971133e44638 · outbound

This paper cites Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651, 2025.

Native Video-Action Pretraining for Generalizable Robot Control Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651, 2025

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.073860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.073860Z digest=sha256:dc2534b7074d0e5dd62572be6d8ecb41fc92972f9bd82cc4cd24524315ccf632

Observation faf7bff2-7eaa-46d0-9391-627f988a76c7 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.Transactions on Machine Learning Research, 2024.

Native Video-Action Pretraining for Generalizable Robot Control Improving and generalizing flow-based generative models with minibatch optimal transport.Transactions on Machine Learning Research, 2024

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.262525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.262525Z digest=sha256:22ebae40709c18a7aeb33e7ed36f8948a0a05a8c16368986cb04c74ba3f8d582

Observation 37b92bd0-663f-4b0c-b5de-319d721355a6 · outbound

This paper cites Wan-2.6.https://wan.video/introduction/wan2.6, 2026.

Native Video-Action Pretraining for Generalizable Robot Control Wan-2.6.https://wan.video/introduction/wan2.6, 2026

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.444882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.444882Z digest=sha256:cfbfc9a0d0986cfedcc9d5d4d931abee5e76819e53550ab6e1d53f190de8a0cd

Observation f62ed543-6f13-49fa-a962-7001878f9cb1 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Native Video-Action Pretraining for Generalizable Robot Control Wan: Open and Advanced Large-Scale Video Generative Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.538533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.538533Z digest=sha256:b9c9dcc724473f762ff230e456dae2cf4566c15ec4ae28bae468b73bfc94b6d2

Observation 0db74f32-95e0-4f38-a473-5cc412fd5aaf · outbound

This paper cites Mimicplay: Long-horizon imitation learning by watching human play.

Native Video-Action Pretraining for Generalizable Robot Control Mimicplay: Long-horizon imitation learning by watching human play

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.695660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.695660Z digest=sha256:c37364b188fbab19f5ae4ed7ddd41e6da25f6fb93773a8d62999018aa582f549

Observation aad1b8e6-edd7-483e-8485-9b6d7919d4f3 · outbound

This paper cites Omnitokenizer: A joint image-video tokenizer for visual generation.

Native Video-Action Pretraining for Generalizable Robot Control Omnitokenizer: A joint image-video tokenizer for visual generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.855514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.855514Z digest=sha256:ee98fef438ec436be65783bf181a8c16c29e5fb8057ae47dab2e7430140cfb82

Observation c6f6253c-93ed-4fcc-a289-8e2b50587317 · outbound

This paper cites RepWAM: World Action Modeling with Representation Visual-Action Tokenizers.

Native Video-Action Pretraining for Generalizable Robot Control RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:42.995964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:42.995964Z digest=sha256:c0d1952fd16805a421e80dd93bb9f743d796573d388ac34612a3489c41af35d8

Observation 3f495a8d-c245-4a17-b45c-985ac738d9c0 · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

Native Video-Action Pretraining for Generalizable Robot Control Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:43.109347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:43.109347Z digest=sha256:7da2d6a361d8e661724fd0f45cb346d366e7d42a0e4ab6571134a45a9e4cfd4a

Pith citing papers

Observation 213f53be-0816-4ac8-8652-76f621b0e5f1 · inbound

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories cites this paper.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Native Video-Action Pretraining for Generalizable Robot Control

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.493281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.493281Z digest=sha256:b5af9e4e9eced2208bdb6cf6af61c3f04c9a37792f3bb0e13599a64c9db951b9

Observation 8db9f672-2b78-4965-bdc9-ed2e07dc608c · inbound

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory cites this paper.

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory Native Video-Action Pretraining for Generalizable Robot Control

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T14:13:45.477522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:13:45.477522Z digest=sha256:09324f835bbed9de4a23c9fa9481bbce742972e3af928c87194fc350602f19b9

Observation 5f426447-1146-4b8f-aa3f-02a2ae4b3f2e · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control Native Video-Action Pretraining for Generalizable Robot Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:27:05.790751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:27:05.790751Z digest=sha256:4b06f20ab43bdad94b70bfa22373a8b63ce59ac1f939c8f3e59c2347a041e7ec

Observation a06b5ef1-b49c-47a9-9332-77450d8f000d · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control Native Video-Action Pretraining for Generalizable Robot Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:47.769490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:47.769490Z digest=sha256:6e96f46d632a3b0d3007da2c33a478fc3584aa6c1da3731842708506f8b5685b

Observation cc3e1df8-ad08-4ef1-8d2c-83ae62681b9d · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Native Video-Action Pretraining for Generalizable Robot Control

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:52:36.172453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T14:52:35.397549Z digest=sha256:0063d6c8698c68e31ac922e2db99df74f4b4dd238dbfa6f8e8e3d28dc43df5c9

Observation 5d226029-f286-4af2-9c6a-c2cd1dcc6a45 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Native Video-Action Pretraining for Generalizable Robot Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:52:22.229506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:52:22.229506Z digest=sha256:a3c529ad14b4ac5ccdb14fda35591582672d7b8a7d745b8a63458b22b61064f4

Observation a9a3a9a8-6615-4ec5-b52a-6bad37bb2b98 · inbound

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility cites this paper.

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility Native Video-Action Pretraining for Generalizable Robot Control

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:38.447840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:38.447840Z digest=sha256:f0dfdcc6f7fcada04ee8a9e58756ff372e866b26b809ae6446fe79651989d800

Observation ab61b01b-3d4e-49bf-ad36-3df8fd7f856f · inbound

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility cites this paper.

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility Native Video-Action Pretraining for Generalizable Robot Control

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:54.863870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:17:54.863870Z digest=sha256:0fdb952bc5a448c25d4fe1975e768db32cb6e5fc3184dfff3ea0f1f4ab9f0d92