Pith. sign in

Paper Citation Record · LEDGER

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

As of 18 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 7 inbound Pith citation observations for arXiv:2607.15330.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15330 v2

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:04:11.999325Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:17:24.951028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:52:36.313517Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11bb3548-07d5-4973-9d2d-123330fe7fc0 · outbound

This paper cites GPT-4 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.117791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.117791Z digest=sha256:87294308ccae6d4b152bca6aea081414ad3cb6084452e062a630804eb06b8991

Observation 605d0126-85e7-43ba-99c7-d5c89b8e032a · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.225496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.225496Z digest=sha256:159a4a16fe7aaa10eb3194a81a982998e2cc282a4db44e32846dfce93169f723

Observation 7fdce7e2-72af-4e1a-9e39-af8369972acb · outbound

This paper cites Qwen3-VL Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.355811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.355811Z digest=sha256:a70a8b84720e1e6abfec63d13f16e29c0829b24aecc3782afdc2c6d202587fde

Observation 9a52982e-ab77-4977-b06b-2cd22399abde · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.504446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.504446Z digest=sha256:f39d56513c86ee900c6b64ba0a767a39419edcae12576753742c37718e582702

Observation 9a11f851-67fc-4cb7-9c2d-616f550a2ab6 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.709806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.709806Z digest=sha256:612fb2f29caabcef32a84a98c697e8fef1da29bf7a155b7a21be5b69681c2010

Observation a06ac054-a993-41e9-8ea8-a60638044087 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.879682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.879682Z digest=sha256:cd1266247dd46aa324b6481990cda7990b5981c96214a1dd758397768140cc1f

Observation 3955dccf-29e0-4ce0-be99-4b7659549687 · outbound

This paper cites Language models are few-shot learners.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.955958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.955958Z digest=sha256:017339781813ebef925bba47e36587a8fa76df4f093df9417ac58633da8555d7

Observation 94428611-91ed-4eab-9408-eac62345b40c · outbound

This paper cites Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.063560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.063560Z digest=sha256:e67e3cbf75792dcab874d3f084a507fa32caaebd2a4270d132bb43d38e35e749

Observation de8021b7-cd0e-472e-bb1b-de885e8a0e0b · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.143342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.143342Z digest=sha256:0c35467504b1a2160986454ef8fbb76920cdbc34875a853a4f24623069d267b6

Observation 5aa6b5db-92b7-4149-8ded-902f997da321 · outbound

This paper cites GR-3 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR-3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.224041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.224041Z digest=sha256:91807be350c6d5eebbab4cba964c975c827536d9e033d0188abbefe46ea6bd18

Observation 01b1a5b9-9e36-45b9-900f-ca28e94475f6 · outbound

This paper cites ABot-M0.5: Unified Mobility-and-Manipulation World Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.292824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.292824Z digest=sha256:9b60766acb0547396045b0880e71ecfb71e44108a28da9ca9fd2c97f205bc654

Observation a91aa2f1-0c5a-43bc-9af8-de7a746a5b36 · outbound

This paper cites RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.390410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.390410Z digest=sha256:85bb415c9692d7a6e84c55026e43a1069c630a0f22bb63bb51082e91cceecc47

Observation 400c7946-e052-4d43-89aa-478c300bdf5e · outbound

This paper cites Training Strategies for Efficient Embodied Reasoning.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Training Strategies for Efficient Embodied Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.468016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.468016Z digest=sha256:ec540bd4641c0bdb2abb5232a05b81eaaa0d0ee2bb93d2de0db4eff7101cf5e5

Observation a71c2dc8-c7fa-425f-ba20-905e005dd208 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.537030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.537030Z digest=sha256:361876de00c97d23ee17ae4808f82b283fadc24e3614cbaa540d110af60bca09

Observation b72ce5ab-7741-470f-975e-0e8edfac7f1c · outbound

This paper cites Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.633844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.633844Z digest=sha256:ea9194aee51640e01574b4aaf99167143750b722cdc9a32c676b2f3ca42b2e4d

Observation d1e7ee89-44c8-4c46-a6e3-27c88216a173 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2024.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.778343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.778343Z digest=sha256:13970915bd9ca522b78164f790a8d033ea8ccbf6a601ebbbfbaa530ed989dbcc

Observation 88f2b9c8-6bb2-4ea0-868c-c7995c5825d6 · outbound

This paper cites Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.022781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.022781Z digest=sha256:4115b57406ce5cb43a9fb2e94702a4a0912e6a14a6955bee04c66653eb97bf8d

Observation 27d05d05-5854-4299-a7f5-2f22d7fcd3a8 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.112802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.112802Z digest=sha256:f7146d3dd5128f1e6e399360b2c200112069251ce7b62115a04207ed1bc919f5

Observation b0517684-f747-44da-b962-a8232099456e · outbound

This paper cites MolmoAct2: Action Reasoning Models for Real-world Deployment.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.196872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.196872Z digest=sha256:4be9389cb3d2a40c278beed40eadc7b221b546c95b216e450bf50728d39d0185

Observation 35698623-2e34-4d0e-9955-6d5a6314550b · outbound

This paper cites Galaxea g0.5 technical report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Galaxea g0.5 technical report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.316861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.316861Z digest=sha256:d042dd170dc8c49a700d567b504cde9d9d1f425bddf6ec1629e263db84123a7c

Observation 3e06fcd5-8fcf-4926-83b2-6d7765aa8b22 · outbound

This paper cites Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.407672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.407672Z digest=sha256:af4dc63c8c8a719492f4fa95301ddd525f905e107207ce6b5bd18e682740a838

Observation 52eb2fbd-0f66-4138-98fb-9e64b0d22fa6 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Training Compute-Optimal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.478584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.478584Z digest=sha256:da168e8324b5c2d1ffbcc6d17128982dc09906faa6f0ef2f186e6f829708fa39

Observation 7378915d-97d3-498f-b507-32b18b84e27d · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.537382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.537382Z digest=sha256:33f78087068861bf69f2310fe0cfa01a92ee6fb94b074ad8cadb3ddb56935ae2

Observation 58fa531a-0a35-45ce-b0ba-8f1298fe14ec · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.650486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.650486Z digest=sha256:c34ab2d30b31d5616d046b06fbafcce513e1d41818a6e953858c6c6178e7ab60

Observation a937599b-d17a-48bd-9544-644636988650 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.761066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.761066Z digest=sha256:dd3df88df04f2834c71fd7f28230c4e024556a9fb447c13d1f4a39ac83c0b0e3

Observation 32c109cc-2363-4441-9bc7-5cf88537232a · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.858008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.858008Z digest=sha256:89f69b4a233dcd30d0ecc35ac2b74ed90598e76c9e93465bf1e527566473b2a5

Observation ce86acd5-91d1-4d46-9ab4-d3612851450b · outbound

This paper cites Scaling Laws for Neural Language Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scaling Laws for Neural Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.934636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.934636Z digest=sha256:18521106f3b7255416a703336590c4c1b99004f9f5c7cc272b8f2a782023107f

Observation ec8145c4-eb6d-4480-b9c1-d94e833b9b27 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.006714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.006714Z digest=sha256:49e26225acffb41ca9ce4ee8a924c12bda645fe60a0c6b7057151153cd65e10d

Observation 81b2b861-086a-422c-af9b-769064d5ff03 · outbound

This paper cites RLDX-1 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RLDX-1 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.063691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.063691Z digest=sha256:2bf2425e0e6f0f63314116c6eba6230def9eb6164f259b4d35fab39b7d989e83

Observation 4e68fd07-b6ec-4942-850c-7059404b9e2a · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories OpenVLA: An Open-Source Vision-Language-Action Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.186263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.186263Z digest=sha256:5f05c3fe49badfc9985fa585f4a241c3dbf2eb343d83b3c45fa6862f9321448d

Observation af3d625c-3303-4af3-a624-b3297f6c60d1 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.287670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.287670Z digest=sha256:111b7cfe15985b19c02ab10d74c05f0996667695f8b0e5aae28c6a265170ec8e

Observation 2464ccfb-0554-4edd-90d8-cd32a7430533 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.398602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.398602Z digest=sha256:236b7355860b3ca306b705fcbe97749bcbf6edca9c6092848c5ee20696ceea3d

Observation 85a13f53-493d-4e77-a9b1-3d46e3270576 · outbound

This paper cites Learning to act from actionless videos through dense correspondences.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Learning to act from actionless videos through dense correspondences

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.581358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.581358Z digest=sha256:76ab11bb7564407e91defd84cc23d8d1040a1e0715125b9dfd03d126f2a55e59

Observation 72bd1631-89b2-4849-8c71-2a2b582af427 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct: Action Reasoning Models that can Reason in Space

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.790902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.790902Z digest=sha256:ab614457356e595f49fab64aed476a4c7c03443a2039fa51e9124e78abdd9e60

Observation d545d019-ecb6-415e-94db-c9ccff4e2e03 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.943489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.943489Z digest=sha256:cab6426980d1ee0d17b15b730d8644abe1fcb709beed30b0745514b5fcb55a2f

Observation 3804a4ed-233c-4b10-9cda-bec18a3197da · outbound

This paper cites Causal World Modeling for Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Causal World Modeling for Robot Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.115676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.115676Z digest=sha256:f5cbc1d623fc1d9bc6ad510d891ad2b1b391ba424209f9b8b692a377eba78855

Observation c8149403-1b63-49f5-ad55-3c20cf0b6c71 · outbound

This paper cites Gr-mg: Leveraging partially- annotated data via multi-modal goal-conditioned policy.IEEE Robotics and Automation Letters, 10(2):1912–1919, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gr-mg: Leveraging partially- annotated data via multi-modal goal-conditioned policy.IEEE Robotics and Automation Letters, 10(2):1912–1919, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.322392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.322392Z digest=sha256:ca8405f887b257cb077c1f982ec0f8e07a94cad1de29bd48ead80d3a477d4722

Observation 9e33eb75-6072-430c-a5a5-df4184f466bc · outbound

This paper cites Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.480239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.480239Z digest=sha256:57783ed0d94ea768c0b77d0aaa6076ae051ba51fa6191fc8f23a4b5f7eedf795

Observation da8e04ef-f768-449a-8b04-c065b67a56a0 · outbound

This paper cites SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.686284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.686284Z digest=sha256:bba8b9018988650fcfe71b1704f97820a846b8cf38db927f051b70ffaacbcbed

Observation c26cbcde-3016-4a5e-85ad-d0bc8ccd912b · outbound

This paper cites Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.836866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.836866Z digest=sha256:6ac2f6fd952db4b0b12bf460f059d9a37d2e2a21686b4d9ec695a91f6c2cd137

Observation 9275040c-f375-47c1-bde7-ebf1aaacd914 · outbound

This paper cites Unified Video Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified Video Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.991822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.991822Z digest=sha256:0feae005268f2e857c1927e647e5cb76f9f8839f77002bf903d6fdc5c5d4c469

Observation 9dc5e2df-7a63-44a3-9226-42faad756336 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.138071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.138071Z digest=sha256:0ac3f67b9d89bf25315ab3275dcee72509c1213846157618c460ffd5e88d8026

Observation 88dc05cd-c5fc-4fa2-9d21-a6d5acb375d6 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.293371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.293371Z digest=sha256:7b53df94c162d79d2491db6c844efa4e160dab59db7510b056d3a07a1d84636b

Observation 27963403-1a86-48a4-9baf-38ec2c91e2b7 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.419047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.419047Z digest=sha256:f7025c6467aaa672175bcfa7a4fc9cab56eefd4e06277fc2c0153efa7b7ac4dc

Observation 65f7c628-ed11-4a50-ba6c-8172d8fdea8f · outbound

This paper cites DeepSeek-V3 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories DeepSeek-V3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.489716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.489716Z digest=sha256:b44adc4a171ad94ac097114f6d059a641133a275591426833da0623ac917fdbc

Observation d605c88b-0a22-41c2-9d22-5eec95c27e62 · outbound

This paper cites ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.559029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.559029Z digest=sha256:315c172ad3f0aeeaa5b364dc7ccb33abe1e919500086de4bf9f8731484053048

Observation ddf314b0-bdca-44ef-99b7-2eac4c1d3ad9 · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.695444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.695444Z digest=sha256:1fb348656e18a45ef4049fea024342f2d2614797dd13a0e589ff055f7ec81295

Observation d0da9d55-b206-4473-bd3b-509c8dabff88 · outbound

This paper cites Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.891865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.891865Z digest=sha256:1c16b82798fa5e8233efeaaa00a61cec15542db5780ab92e20eb9dd6d0b00a1a

Observation 2ae21ebe-a60d-4db7-aeda-c7554c421de5 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.963758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.963758Z digest=sha256:3ebb1f8e6e4d5c666f851e284f50b500117d988a6acf6b01cb6485b6dbff259d

Observation 5efc2714-0f33-4327-81ef-d8033792d4ec · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.119524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.119524Z digest=sha256:5119ae5125f366af71803b0e896d957729b7f063da8bbb70d36fb9971a5851c9

Observation d9301cf6-8c57-4712-b8bb-df456ef42c4e · outbound

This paper cites Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.287415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.287415Z digest=sha256:b7aa0e5ea4d764182daa50fe8bbaf1331d96bc30f9057a740dafd364d4d5440d

Observation 4fbd697c-610d-4dc3-95ea-c91b01898eb0 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.391932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.391932Z digest=sha256:625ca1729201e52ec1a6efb2dcb7c43c094607fae47b1357fa1ca8488756a3a3

Observation b1b7906b-00a3-40b6-9ede-df5d6eabe242 · outbound

This paper cites Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.541501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.541501Z digest=sha256:3d40f9c424853f5b8a3b255c7f19a10498bc178fec3c8b1d557bc45d5fb129c3

Observation 19d1673e-9e66-4725-88a3-376012e51931 · outbound

This paper cites GR00T N1: An open foundation model for generalist humanoid robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR00T N1: An open foundation model for generalist humanoid robots

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.652775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.652775Z digest=sha256:cce4f54e45e99f8ba400a8f682c21792f46d31148ae96c35c23a557ff1f52ecc

Observation 4f1eb89b-d28b-4701-bb23-c184692f9990 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.791080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.791080Z digest=sha256:b85d1ad0505525cace7ac85f9ccacd7f787c9249fd79eb2c362db4fc8c9ee17d

Observation adfdcd9a-76f3-4940-aaa1-6db85fa6e05b · outbound

This paper cites mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.928505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.928505Z digest=sha256:5b12950a752e9f0d5814155752a28d86c0cf532681348d79c4d3ee3f8a19ae17

Observation 3821579b-4c34-4dcd-b8ec-3a79210a916b · outbound

This paper cites Scalable diffusion models with transformers.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scalable diffusion models with transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.037779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.037779Z digest=sha256:a92853e83c602503b5c3a4c5e81601749246a41bbbb4e36993681706c3cd00fb

Observation b5b906f0-94ed-410b-b654-8cbc2c256d97 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.230214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.230214Z digest=sha256:bdf19b2b9da6cfbd977738628836d713d352a4ceb6611f6069f0c8790a214c13

Observation be4dad51-fdcd-4991-89a6-6f2bf3f0efbf · outbound

This paper cites Coordinated humanoid manipulation with choice policies.arXiv preprint arXiv:2512.25072, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Coordinated humanoid manipulation with choice policies.arXiv preprint arXiv:2512.25072, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.351663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.351663Z digest=sha256:27229265f39bf54818044c3122b39b671031d35f50484e02760a4943c347cf59

Observation c7ec629f-4625-414a-9ba2-4c4780ea2f4a · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.479669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.479669Z digest=sha256:5b48913155e9a59fe17ae0e92ac825410ae48f05dd99b2ea15340c1282444aed

Observation c3ce2c8f-bff1-4071-9e9f-8baf8f6cca98 · outbound

This paper cites Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.613967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.613967Z digest=sha256:ddf18e6acdb6ec290735e259e6179308ce534cf757b426eb79c9994bf9fa17c6

Observation 526ff66a-1512-4c10-b841-31128f6c2545 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini: A Family of Highly Capable Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.798973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.798973Z digest=sha256:591b812b524da2e7776872cedf0a5b46dd7cf4093ec66e42894f3dfd48df0060

Observation 73fbf405-d6d4-4024-bcbc-18c48348837f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.931913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.931913Z digest=sha256:f117e38e603c3d75bb74dc82f53d057fec038b0d52ea8d87159e9f6cac8163e4

Observation 54d0cacf-9905-4812-9370-8b6541f3ad95 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini Robotics: Bringing AI into the Physical World

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.086788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.086788Z digest=sha256:1068d9d3b3f0c48c07d2bc8dffe7e3bbf9f4e2d74a84c14c52e06cf6afdd0887

Observation ea493c74-4e01-414c-85ec-437ffb62a7fa · outbound

This paper cites Gen-0: Embodied foundation models that scale with physical interaction.Generalist AI Blog,.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gen-0: Embodied foundation models that scale with physical interaction.Generalist AI Blog,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.181000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.181000Z digest=sha256:6e883aebf9dbbccf628538786ea6c457e1ea77b2f43f02ff1c136a5d0725c753

Observation 2e7129a0-07f2-4ea4-b51d-240ba3d2c36e · outbound

This paper cites Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.426810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.426810Z digest=sha256:0e0be30bfa1857decdd2528ee159b01f61ae3339ad465263e2b594e1107340e7

Observation fd335565-bdce-4b16-94cf-dafbcb829c71 · outbound

This paper cites Gene-26.5: Advancing robotic manipulation to human level.Genesis AI Blog, May 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gene-26.5: Advancing robotic manipulation to human level.Genesis AI Blog, May 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.591352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.591352Z digest=sha256:9099008fb015183bf9a2a5db36f98acf2e21360f03c88735792ba860316163d3

Observation 6408f5d5-3515-41ab-979e-bda2709f60eb · outbound

This paper cites Motubrain: An Advanced World Action Model for Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Motubrain: An Advanced World Action Model for Robot Control

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.705109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.705109Z digest=sha256:b7300c43a6168b52fed7597a03177da612ce5b5031a4de51be9ed34c3479130f

Observation 05a07c5b-9198-4593-b574-79a2746497dd · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Octo: An Open-Source Generalist Robot Policy

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.846442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.846442Z digest=sha256:2eb763cef79dc518b73277372e8dd635be7b81f2e869a19a975e5857341c8530

Observation 3dc7a77b-da8f-4386-9f13-545297db9c9b · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen3.5: Accelerating productivity with native multimodal agents, February 2026

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.970461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.970461Z digest=sha256:68b3f3ab142834da5e9008f5b4edb670a2570e478ffb920fab06c157c06abf3b

Observation 37642b6d-bc80-4668-8e71-19881e850fcc · outbound

This paper cites Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.076726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.076726Z digest=sha256:1618c9647d7cfe62f3092bb5cff16009dc2d802e32b0ec7167275bb133b4e74d

Observation 44c93c5e-71fb-43eb-a527-f39338861437 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.222348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.222348Z digest=sha256:d20d31d605e3c54460a272d3247e7e876434d010d9ffdc714a0f77755e620fa2

Observation 1cb546a4-854a-4567-80b1-5f9eb81dba9b · outbound

This paper cites World2Act: Latent Action Post-Training from World Model Dynamics.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories World2Act: Latent Action Post-Training from World Model Dynamics

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.361341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.361341Z digest=sha256:21defe6ceb13a8d77c887dcaa902e85288a0b0ae9d7218f3e76e468a86fef8bd

Observation e10fc9c9-56a0-4063-b656-35fbd7d67500 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Bridgedata v2: A dataset for robot learning at scale

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.546668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.546668Z digest=sha256:3eec90bd7dc3deb9c4e48087716221b10f7f4aac5f2f1973b86b15b984820d26

Observation 5bbb7262-ef3d-492a-be20-e2d399360cd3 · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.689523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.689523Z digest=sha256:31eb86c90db3b40c4a637593fa2a532cca53c5b8551d0aa573295df948ab6e37

Observation 5ff7846f-3914-4da3-abbe-e85fa1fba95c · outbound

This paper cites A Pragmatic VLA Foundation Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories A Pragmatic VLA Foundation Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.883631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.883631Z digest=sha256:6b4951659be1414e7375ca551d9e3927df4c5150fed522941428d372da173e5b

Observation 8847c4a0-c647-44d2-ba2c-5483b46002be · outbound

This paper cites Dexumi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dexumi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.062173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.062173Z digest=sha256:b58d858244364a3f89175b2025a5e9174b50b5804e149280a1064079c09a8129

Observation 571a8b58-02ec-4e07-8ad5-b643aa9fdfb4 · outbound

This paper cites MemoryWAM: Efficient World Action Modeling with Persistent Memory.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MemoryWAM: Efficient World Action Modeling with Persistent Memory

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.236077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.236077Z digest=sha256:999d1a982f4d448c4bdc4589fa7884e1f29c74c0c4419f3e8c41baece860f534

Observation aeee5494-a31e-4895-9005-01475331ad58 · outbound

This paper cites Gigaworld-policy: An efficient action-centered world-action model.arXiv preprint arXiv:2603.17240, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gigaworld-policy: An efficient action-centered world-action model.arXiv preprint arXiv:2603.17240, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.365590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.365590Z digest=sha256:8fde3fecee35c25db5bc7dd8476dbe3d71ef438df11464e2c644f7c624022144

Observation 7419d939-de71-4609-96b3-1baa5e0f54ab · outbound

This paper cites Starvla-α: Reducing complexity in vision-language-action systems.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Starvla-α: Reducing complexity in vision-language-action systems

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.530949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.530949Z digest=sha256:900e1f7d0c662fd60403fa4b9cf008fd0a5d4ee9c4e11041b7c53b3cb15e9c97

Observation 88207e1d-9378-44e1-b8a2-f40c32a463f2 · outbound

This paper cites World Action Models are Zero-shot Policies.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories World Action Models are Zero-shot Policies

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.676504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.676504Z digest=sha256:d2eba3e8c6a13143ba33c754dfbaf8e5ca9853a5a8cf9cdc36e14ef64c1c6e52

Observation 628babad-8bf5-49a2-b9e3-d20096c08eeb · outbound

This paper cites Wall-OSS-0.5 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Wall-OSS-0.5 Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.858319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.858319Z digest=sha256:bb9eacf64b2db272b9d3b579f6b7d54a84c7f5889c5c524f8f8b34808ae0f9d8

Observation 181cbaaa-febe-41b0-b66a-e702b18d0f09 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.006144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.006144Z digest=sha256:9c5111cf3c7e91ecd2f669974ee424515f5deea154518be4d243bcf68ad89e78

Observation 395ff754-9e63-40d0-a1d6-14444ec8c58b · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.138148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.138148Z digest=sha256:4bb6513a9b596fa87885aefbd08f2fccd93c9fde84373cc9ffca908d8a0b79a3

Observation 20e0fc5f-94c4-43fc-ae96-5be679b4161b · outbound

This paper cites Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.301551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.301551Z digest=sha256:b3884a46731b6b4aa0ac6a38fbc2f7fb23689bb7d262f0fc4d4c5176e04c6275

Observation 213f53be-0816-4ac8-8652-76f621b0e5f1 · outbound

This paper cites Native Video-Action Pretraining for Generalizable Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Native Video-Action Pretraining for Generalizable Robot Control

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.493281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.493281Z digest=sha256:abc346708086124709226d0836a1a9d71d0ed270e48aea9b718043686381e440

Observation 1a1a8114-70f9-4b9a-8a6b-382661e04863 · outbound

This paper cites Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.653923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.653923Z digest=sha256:eacf12b00cbe2eaea451ab158dc4c78ad66b6d7c968ccc8c3be86b008b57d10f

Observation 78a64214-cc8c-46c0-8ab8-8c27ff66a29d · outbound

This paper cites RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.767303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.767303Z digest=sha256:a4fb62d23bb168a391907c7a9d03a8d36081bddaafab39845a960a8bf0e6cd9f

Observation 8662e36b-e228-487d-864a-51bc0c18689b · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.975158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.975158Z digest=sha256:5a9e7a5eca3a2511034e7e874b6522e0fa6d4e33a9e41cde2012e3033f69e72d

Observation abd2b7e2-65a0-4fe7-9342-6116df8470fb · outbound

This paper cites Fastumi: A scalable and hardware-independent universal manipulation interface with dataset.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fastumi: A scalable and hardware-independent universal manipulation interface with dataset

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.125595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.125595Z digest=sha256:a2e23900d1740989748d57fb87c97a4889969dcb6beaff7264561e27681b8ede

Observation 7574aabc-55ee-4055-8bc9-af23deb11fc9 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories TesserAct: Learning 4D Embodied World Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.217040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.217040Z digest=sha256:dad5f24da6dcb224910daacf5746aeba1d1ec0975cb7b4286f7b229062fa23c8

Observation 8fe5f4eb-12bf-4b5c-bf0c-183af1e848f3 · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.291258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.291258Z digest=sha256:5b342d7eb683cbcdcff57626b863ca057808929e751d2789bc0285e4e93f43ce

Observation b17054e2-dfe6-4555-aa1a-2ee1aaba470b · outbound

This paper cites Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.427801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.427801Z digest=sha256:c639b7ed9762c3c66a9bcd85441b4abcdfd4c44187f62d4e3a538115e5f17dc0

Observation b9a015f2-28a9-445d-9777-5fd99f0a73e9 · outbound

This paper cites Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404, 2026

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.570801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.570801Z digest=sha256:eb22a67c7bf40a02db97459e2fa0a73e6af67b8e8ba68be78fac48ffecff087b

Observation 720928f3-040e-4442-b1ad-e10b739fc6e2 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.705089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.705089Z digest=sha256:4cbc64ffb2e6d05e1481a40558d038cca8dd55d46e8b41a4ff2c8c248da3e7f5

Observation 6f9c9bc2-b575-4938-a372-d6cf0c9f1af9 · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.845223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.845223Z digest=sha256:d259a77f933833990e79f40e9f1b6a467f2cb6607867540185bd7d59c822c76d

Observation 9eb94489-b3e8-461a-985b-bc8690f8338e · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.999325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.999325Z digest=sha256:9cb1bada9e8a9ff952bb2131c4dc714856a51f1bba032c33800dc4b09fb9e5b1

Observation d79aef79-964d-4da3-96ab-35e32f332ef8 · outbound

This paper cites an unresolved cited work.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.331395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.331395Z digest=sha256:34ece02997841798a83a3169197c72080aaa6a1f33264d49ea6d995c3c149c5e

Pith citing papers

Observation a0aac408-4465-4ad5-a4a1-9ef4b06df91e · inbound

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens cites this paper.

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-30T12:43:44.703075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:43:44.703075Z digest=sha256:43a7223fbb69fa15b0ead0654afda8b869201a765a93ca064b40c6d4c25051ba

Observation 4807410a-d1a0-4b83-b8e9-f36242eb9ed9 · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.603531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.603531Z digest=sha256:8cf386f8ac8bdba30a0e85fb9056e3b22a6d565b96c1dee6843d7e44286f81c9

Observation 93313b23-be72-440b-899f-91968f484b7d · inbound

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation cites this paper.

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:17:24.951028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:17:24.951028Z digest=sha256:41bf51c733a61b7e971776ad6ec46b2a40955a6d6d58e49fd08d014a144e45ec

Observation 26a539bc-f8cc-403b-84fe-ecb4a2e103a7 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:52:36.316445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:52:35.371109Z digest=sha256:50deaa91ae27561868483f1953175182e468f1a48beb64c63794266255c18097

Observation 35032d6b-ae97-4ea2-bd59-3487dbbc969a · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:52:22.178974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:52:22.178974Z digest=sha256:9632fdb3584bd29e83672e963523519b2216d161c975de41809af747bf74ba32

Observation ba3a3086-a292-43c5-88be-b9a14fafaf44 · inbound

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment cites this paper.

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:54:24.134310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:54:24.134310Z digest=sha256:8d782842e68985d7773bae8d68a9dcdd5b1859f4033632b4ad03cc27a125e915

Observation f11a420e-44e5-456c-94a5-f8161ea423b4 · inbound

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment cites this paper.

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:47.516551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:17:47.516551Z digest=sha256:845290d3d314ede43e9b02b34e371b0565c303aa8ff493738d23993d9c6b3e49