Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:33:46.663820Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 100 of 161 outbound references and 63 inbound Pith citation observations for arXiv:2507.15597.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:33:46.663820Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:10:14.892153Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T04:16:48.668605Z
100 of 161 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6f1216af-6b17-4952-b576-49cb6f98da3a · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Review on human-like robot manipula- tion using dexterous hands.Cogn
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213d718d-a32c-4af2-9165-83202181cc38 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human- like dexterous manipulation for anthropomorphic five-fingered hands: A review.Biomimetic Intelligence and Robotics, page 100212, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fba0bd0-220a-4636-88c9-994dbfdec581 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RT-1: Robotics Transformer for Real-World Control at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a33fae2-b40e-46a0-9aa4-30f4344c74ec · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c761ec84-fc8d-4f83-90a1-a10fa9f9c1b3 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos OpenVLA: An Open-Source Vision-Language-Action Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd76e2bb-41aa-4466-9b3b-2a46499d2114 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3755938-f091-4b94-bf74-424a6f6857db · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos A Survey on Vision-Language-Action Models for Embodied AI
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 083d9651-4bab-4508-bd8c-d5c12ae2c2ce · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0167b96c-8cbb-478e-8f81-72518a88ed38 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15bf4cc7-8ccd-4bf6-85d5-88b7d152e3ca · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d16fe479-8fab-4ff1-bafc-196a245d4018 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Octo: An Open-Source Generalist Robot Policy
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6b63b9f-f0e5-4395-9dc3-3f997dcb7cb8 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Benchmarking Reinforcement Learning Methods for Dexterous Robotic Manipulation with a Three-Fingered Gripper
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 85fbd045-d31a-4aeb-8290-96b34abab5cf · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Dexterous manipulation through imitation learning: A survey.arXiv preprintarXiv:2504.03515, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca9382e-bac0-4847-b679-5c588f9d4f30 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491785a3-b285-43a1-84bd-44dd8f44e829 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Scaffolding dexterous manipulation with vision-language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184e660e-cae0-459c-a0b8-4082e56b2734 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6029deac-11f4-4de2-8f49-475f5190838a · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9012a185-501a-4aa7-b078-aa123dc4d97a · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c96aaab-977f-4a81-9243-af342b15fcf0 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DexVLG: Dexterous Vision-Language-Grasp Model at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc36849-c43d-428b-b6ad-a01e4bd224c4 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Efficient residual learning with mixture-of-experts for universal dexterous grasping
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d52d96-73ba-4d36-b6ac-1a52ae2cc874 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos R3M: A Universal Visual Representation for Robot Manipulation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c462fb1-22bf-4e43-905f-15c05c20b67a · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Real-world robot learning with masked visual pre-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 135da890-1f3d-4ffa-a1a0-76593fa951cf · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d704d37f-98f8-4199-a36d-79442114c205 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Improved baselines with visual instruction tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2097f02f-9b4a-4ef1-aa62-ffc549579545 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Humanoid policy˜ human policy.arXiv preprint arXiv:2503.13441, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741b91a3-801d-4279-96aa-da1d43cc7a09 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Integrated linkage-driven dexterous anthropomorphic robotic hand.Nature communications, 12(1):7177, 2021
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b02f06-166e-421e-ad27-2187dbe4f505 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f3022c-e89f-4ee9-bf75-12b6dc9d4e00 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Autoregressive image generation using residual quantization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acbb01f9-ab9e-405a-9aaf-dc55f47d8d26 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea0e8ec-bf17-403b-accf-af76bacde7dc · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1988342e-245c-4e46-b539-17e5a62fccee · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Improving language understanding by generative pre-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed2d5c7-ac6d-4597-bfd7-852ab5fb44e2 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e4cffd-b982-4108-a1c4-743565f197b9 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language models are few-shot learners
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60782969-dfa1-4753-9d7f-d8d02d600f00 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4374b8c0-d9c7-49ff-b2cb-24044a597685 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0f0ddd-94f6-4668-a86d-3723dcc5bfa3 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos UniCode: Learning a Unified Codebook for Multimodal Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation deb167aa-01ee-46b0-9fbd-898b7204f23c · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos From pixels to tokens: Byte-pair encoding on quantized visual modalities
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf0c352-4347-4667-bd28-45f795d9c910 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unified multimodal understanding via byte-pair visual encoding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066c0331-e95b-4295-b5a3-562d86f09264 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos LLaMA: Open and Efficient Foundation Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0342fd5e-004f-42f7-ac47-d2370206bec7 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad47a498-55cc-41c1-8ab6-3ddd7213b88d · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16384af9-1a5c-4ebf-8af8-82e3d0a36ef1 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Learning transferable visual models from natural language supervision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa01e2e-e516-487d-b31b-24f3b03e132d · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Sigmoid loss for language image pre- training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795e4fc5-e425-4936-b631-b248f5c0f6af · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Flamingo: a visual language model for few-shot learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73eea413-02bd-4805-9e29-ebe6079d4b4b · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4c7e95-305b-4de5-a648-8eb8351a6023 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c79e090-2cd3-4c00-8792-c604cdfd504e · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Otter: A multi-modal model with in-context instruction tuning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae659049-9bdc-4c7c-a9a9-ae018ab13db5 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a8bc03-cdfb-4104-b4c7-46b0ec370c46 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1b8b98-c779-4bba-ba14-eacaaf7e8583 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos VideoOrion: Tokenizing Object Dynamics in Videos
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38d25635-d466-4879-9c91-7c7a075b8501 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b40deb-04fd-4eac-9224-ba01623ceca6 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Gemini: A Family of Highly Capable Multimodal Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4144850f-7008-4400-b3f4-71b67d53e5f1 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7303ebdd-6754-4c31-abb3-8adc9def0a10 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos GPT-4 Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d05e1ec-5676-4ae6-8442-5ff22a4731f8 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen2.5-VL Technical Report
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2359026e-9846-4af3-baa0-3e9031579db4 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbd68bd-fb7c-474d-b25b-d70e32d48fb4 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9fe530f-52cd-4c1a-99ab-7acedadf48b7 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edecf294-eb4f-4955-bde3-9c9d01281c77 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos The kit motion-language dataset
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6faaf8af-74e8-4ad0-a2a7-ccd77fdca5b5 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Amass: Archive of motion capture as surface shapes
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53026f1-1d85-4640-81e7-8aba6ff3d6ff · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Generating diverse and natural 3d human motions from text
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60971193-2d7b-4864-a387-f741b9783ea8 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Babel: Bodies, action and behavior with english labels
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee93827-7900-4bec-87d3-aa21d56eef54 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advances in Neural Information Processing Systems, 36:25268–25280, 2023
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e470422-15ee-460f-a820-a5e001100406 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Egobody: Human body shape and motion of interacting people from head-mounted devices
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c5d703-6a66-4c4d-aba1-7872741a2f59 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Nymeria: A massive collection of multimodal egocentric daily motion in the wild
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a028043-3caf-45a2-948c-f1761def8109 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Scaling large motion models with million-level human motions
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5e0376-e134-4548-85f5-bed4f8b76961 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Smpl: A skinned multi-person linear model
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bfadaba-b22d-4fbc-9e10-e9213c25b9d1 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Expressive body capture: 3d hands, face, and body from a single image
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7f0419-75a2-4b7a-8313-4339200e13ea · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human Motion Diffusion Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64da1b0-6537-49c6-ad33-c267df7b40fb · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1483f9c-a7ff-4d5e-9aaf-a978a0a6bfe8 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Executing your commands via motion diffusion in latent space
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e06db4b-fddb-477f-9953-fcfc76856df7 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Physdiff: Physics-guided human motion diffusion model
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0acc39f5-fa8f-453b-8df1-d67450ec27dd · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Remodiffuse: Retrieval-augmented motion diffusion model
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59819418-f3cb-4b2b-a203-e1ee2f99d6bc · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 682795f0-2e56-4979-9ad7-04e47722f7b2 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Large motion model for unified multi-modal motion generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3781b643-41af-4e61-9e88-dfb34367abfe · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Neural discrete representation learning.Advancesinneuralinformation processing systems, 30, 2017
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ca6c3c-e873-4be6-8e5a-76ab3fc3d47e · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Generating human motion from textual descriptions with discrete representations
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faaf87c1-c4a7-4048-ae4f-ba422cd2cc71 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Momask: Generative masked modeling of 3d human motions
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c50ce6-706c-492e-82a5-fb1296c08074 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cabb6102-fbb4-441b-a064-d4dd6e18247c · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos HumanTOMATO: Text-aligned Whole-body Motion Generation
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fde13bd-7364-46ce-a8b4-bf498e37734b · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Finite Scalar Quantization: VQ-VAE Made Simple
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a3761b-fb1e-4cd7-9020-57caeba12fd1 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e1c9c6-8b7c-45e9-8f6f-66736c3c6272 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motiongpt: Human motion as a foreign language
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7acbec-d0c8-4121-9b14-ad743b74c024 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cd73b6-a9aa-41d4-97c9-144cc1f960c5 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionLLM: Understanding Human Behaviors from Human Motions and Videos
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfed5262-2915-4001-b16e-cc2be9cb474e · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Avatargpt: All-in-one framework for motion understanding planning generation and beyond
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba2656d6-4ed0-42c8-92a0-fb815962cbee · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motionchain: Conversational motion controllers via multimodal prompts
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b8a8fc-126e-426a-913b-64b2385a169d · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d245ff8-9b14-45cf-bd77-9f159ac67b35 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human motion instruction tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a7212e1-8afc-45b2-a49d-cba95fb4f3e0 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9fdfc3-20bd-4aa5-a976-0eb87aef1e2c · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Perpetual humanoid control for real-time simulated avatars
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b6f6754-fb9c-4da2-ac3d-db188e302e23 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ExBody2: Advanced Expressive Humanoid Whole-Body Control
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dad6a80-a450-43de-b8e7-b6141c91e102 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Reindiffuse: Craft- ing physically plausible motions with reinforced diffusion model
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8d7db9-e877-4a54-8368-084b19507e00 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900fbd39-23f3-426f-9049-b2f95871ed97 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hand-object contact consistency reasoning for human grasps generation
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a3701f-955f-4d51-b7f4-43fe3e148d98 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Joint hand motion and interaction hotspots prediction from egocentric videos
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079824de-be87-4b1b-97ba-a373f732fc17 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hot3d: Hand and object tracking in 3d from egocentric multi-view videos
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19ece4c-6d3d-4fda-867d-6e11e371444c · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9bcd271-54de-4248-a716-a2d7b3e2b4a0 · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Oakink2: A dataset of bimanual hands-object manipulation in complex task completion
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d286b6-3268-443e-8562-e447361b5c5f · outbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ego4d: Around the world in 3,000 hours of egocentric video
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b59d164-430f-4446-9e55-2a5eac4ef086 · inbound
Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea14750-e58d-4048-b56d-b8b74d8d3eb6 · inbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f85282-66ee-412a-84da-64e4d1894b45 · inbound
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 228ccdd7-920f-40ef-bcab-b572bd8fc53a · inbound
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14292983-4c79-4716-84d6-caa3591e71cf · inbound
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3dc394-c737-4508-8cc6-ddc2e5d46e0c · inbound
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation df45d407-7e35-4e6c-ae92-b93786349036 · inbound
AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4bb248b3-fd46-4cde-a901-5771d349e93c · inbound
LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e903d4-96b3-4591-9df8-7feafe535f33 · inbound
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 995409de-51e6-4c60-8313-a0d448e7afc8 · inbound
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1506ab-6795-42ac-a247-a675d0621889 · inbound
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0d1f7b55-f0f4-4ffb-a777-628617c0a3dd · inbound
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792b1e76-3bab-4a37-bf0e-e6655d513d82 · inbound
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64153462-9782-4f60-8a31-79fdad4c8103 · inbound
A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a2c6c44-858e-4db3-8980-fbea3cac73d0 · inbound
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fdef3828-756b-4c33-8992-1d37bb0ebb4b · inbound
EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8fceb112-ea8a-4399-9ffb-052c51f8bd42 · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 17e17457-c4f2-43ba-8fd9-251ff72234fa · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ced91d16-378a-404b-8909-be2b52314578 · inbound
Being-H0.7: A Latent World-Action Model from Egocentric Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13379a2f-4b85-4ade-9d08-19553221a7e8 · inbound
World Model for Robot Learning: A Comprehensive Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe5defcd-459a-4c6d-9907-f07a38569958 · inbound
HumanNet: Scaling Human-centric Video Learning to One Million Hours Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a47c22ac-14c5-4628-abbb-cb6baa225d67 · inbound
World Action Models: The Next Frontier in Embodied AI Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f15d5d9b-626c-4120-bb49-7f6ebdfa7de5 · inbound
Towards Robotic Dexterous Hand Intelligence: A Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6258c08-aa48-4067-98fd-59a673b365a8 · inbound
Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51ee9ed2-a371-49ed-b5cc-a6fd44d4c33b · inbound
Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7b6ac8ad-7cde-43e6-9872-23eaf06c23d0 · inbound
Dexora: Open-source VLA for High-DoF Bimanual Dexterity Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b88443ab-14a4-4068-9482-5c141b4c9267 · inbound
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b4d04cf-4c1d-4503-9490-b72a87ba08e3 · inbound
BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7ab91d45-1471-4c74-b056-ea02fcfc11e3 · inbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd4dd992-0dc2-4880-be98-213785842ab2 · inbound
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e1fb98e0-3d75-4e42-bbdf-70f7cd08c5c7 · inbound
Unified Video-Action Joint Denoising for Dexterous Action and Data Generation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9bb1aec4-8dd7-4f1e-8274-da739a390b14 · inbound
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 389ecd41-aead-427a-bbcc-05e33a2970c7 · inbound
RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79d0f620-2bb8-42da-992b-afc11cad27bc · inbound
LARA: Latent Action Representation Alignment for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ca0da9aa-21c0-490f-bec8-32def5ec9703 · inbound
LARA: Latent Action Representation Alignment for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 462e6594-bd69-49ee-bb7c-2f578b6c608b · inbound
$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 18ae0df7-1766-4f47-9266-4fe04473699e · inbound
Next Forcing: Causal World Modeling with Multi-Chunk Prediction Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bfee454c-5b31-4cba-9671-a3779d34a43c · inbound
LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 67082a9d-78e1-4134-b360-31a571b052fa · inbound
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b1983b03-9216-4236-8dcf-73e1abbe4232 · inbound
Do as I Do: Dexterous Manipulation Data from Everyday Human Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 25489053-9cea-4a60-8674-87b32178ba29 · inbound
ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6ed30d62-4531-400c-b6a4-28aeb16a039e · inbound
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 73810ba6-6a3f-488b-922b-b89b9011c5fd · inbound
Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e24bd8d-ec6c-4acb-8ff3-dd8b1f91a3b6 · inbound
Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d836be21-89c4-4eb5-b6bd-f59877630a17 · inbound
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a6366124-29b5-4006-a58d-fff533a68f5b · inbound
Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f5181d8a-4432-4e9d-b577-41cd6ea9a2ff · inbound
Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c898bc55-29e6-4ff2-b04c-63542be34257 · inbound
From Foundation to Application: Improving VLA Models in Practice Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation caedafdc-980f-48df-93e8-8e0073c6f14a · inbound
Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f742ffaf-db65-4417-84e8-dc83c1b193ce · inbound
Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2766f5a6-6e0c-4c25-8cc5-2eea594df53e · inbound
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2138be29-39f8-4103-9629-8bd6580c280d · inbound
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efc2714-0f33-4327-81ef-d8033792d4ec · inbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58695d96-07e2-4795-85c2-02cc499f5d1d · inbound
HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf6ea43-c392-463e-9ea5-59d2c684839b · inbound
LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd9e4a2-c9cb-44ad-a3cb-036e28beb184 · inbound
Data Pyramid for Embodied Manipulation: A Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 237
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012798cc-eb2f-46ec-a04b-3a3a4f9b30b8 · inbound
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 227
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719960a9-0211-4c33-8590-b2815129ed21 · inbound
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 208
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a809f7-8b0d-4924-99dd-b9e9b18e3015 · inbound
PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6434a5c5-d393-40ae-b61a-1c6cd850d6d0 · inbound
PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67c89dc8-ab09-47eb-9697-ebc544dd42f5 · inbound
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8ecbb9-fb0a-4255-93ec-bd6f50059c1b · inbound
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d4981d-d08b-4f6a-85bd-92b8aed13ce4 · inbound
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.