Pith. sign in

Paper Citation Record · LEDGER

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

As of 21 August 2026, this Paper Citation Record lists 100 of 161 outbound references and 65 inbound Pith citation observations for arXiv:2507.15597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15597 v1

Coverage vector

measured 100 of 161 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:33:46.663820Z

measured 165 of 165 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:52:38.381035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T04:16:48.668605Z

Reference resolution

100 of 161 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f1216af-6b17-4952-b576-49cb6f98da3a · outbound

This paper cites Review on human-like robot manipula- tion using dexterous hands.Cogn.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Review on human-like robot manipula- tion using dexterous hands.Cogn

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.030744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.030744Z digest=sha256:ccc3f89f24f5eddc70482832fa655f2a30e4306454cabe0e85dafe4ef3cb2276

Observation 213d718d-a32c-4af2-9165-83202181cc38 · outbound

This paper cites Human- like dexterous manipulation for anthropomorphic five-fingered hands: A review.Biomimetic Intelligence and Robotics, page 100212, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human- like dexterous manipulation for anthropomorphic five-fingered hands: A review.Biomimetic Intelligence and Robotics, page 100212, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.117746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.117746Z digest=sha256:02532db162946404811a7d3ab55999ff347481d9eaf8bb0ef30c911d52e0c741

Observation 7fba0bd0-220a-4636-88c9-994dbfdec581 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.186546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.186546Z digest=sha256:77a3c4dd10e4d1f65a171247268ee131abf3f41f16ac66f1ce1459569f4a8dc9

Observation 6a33fae2-b40e-46a0-9aa4-30f4344c74ec · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.258306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.258306Z digest=sha256:e484f994c50ef44a02133711005ad0412c95fc550d481bac6bc8975174088652

Observation c761ec84-fc8d-4f83-90a1-a10fa9f9c1b3 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos OpenVLA: An Open-Source Vision-Language-Action Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.333345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.333345Z digest=sha256:6f0c0af94f1b25e1448e76a516a79473833638c557b2aa54667a8278663ac65e

Observation fd76e2bb-41aa-4466-9b3b-2a46499d2114 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:33:38.444159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.444159Z digest=sha256:b7119315eba9ec88c41db70b3b14318ddd8becc6c840211b27cb4eb95526d035

Observation c3755938-f091-4b94-bf74-424a6f6857db · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos A Survey on Vision-Language-Action Models for Embodied AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.504555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.504555Z digest=sha256:1db223fe7be4bdd403f0172fb4eab9e0e8938ae405a4162d00c6585290aa6f72

Observation 083d9651-4bab-4508-bd8c-d5c12ae2c2ce · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.565294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.565294Z digest=sha256:48b1d9f718f1bf3fbf7b1386261621e32ad54205b47818e4c2711bc5b7c0f169

Observation 0167b96c-8cbb-478e-8f81-72518a88ed38 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.669232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.669232Z digest=sha256:2f42a3d0cc276d96be3fabadc0f56a5c9ba3187aa4872d007bd1acca9ecc079d

Observation 15bf4cc7-8ccd-4bf6-85d5-88b7d152e3ca · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.782713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.782713Z digest=sha256:590f903d9f42e38afda734964f64ed5e4833f4a6a22d839e184c6507932f6774

Observation d16fe479-8fab-4ff1-bafc-196a245d4018 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Octo: An Open-Source Generalist Robot Policy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:38.850217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:38.850217Z digest=sha256:4ff529aede5059e95a41d0f0ff53ad401aa9e50cd0c63724bf85075971e18065

Observation f6b63b9f-f0e5-4395-9dc3-3f997dcb7cb8 · outbound

This paper cites Benchmarking Reinforcement Learning Methods for Dexterous Robotic Manipulation with a Three-Fingered Gripper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Benchmarking Reinforcement Learning Methods for Dexterous Robotic Manipulation with a Three-Fingered Gripper

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:52.228371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:33:38.935193Z digest=sha256:267593fefa264a639fbebb3f14947ecd3947b8b5d2bcf00ee65ba3451d204c9e

Observation 85fbd045-d31a-4aeb-8290-96b34abab5cf · outbound

This paper cites Dexterous manipulation through imitation learning: A survey.arXiv preprintarXiv:2504.03515, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Dexterous manipulation through imitation learning: A survey.arXiv preprintarXiv:2504.03515, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.037596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.037596Z digest=sha256:4918555696dfb1b5b75208cf4b5e0fc892f226ce8598aad52fd5eec097e682e2

Observation 9ca9382e-bac0-4847-b679-5c588f9d4f30 · outbound

This paper cites Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.082418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.082418Z digest=sha256:ed583f867ad5f31c3739b0a8b02d1eac0438a76b25c085893efe7e73a258e5c2

Observation 491785a3-b285-43a1-84bd-44dd8f44e829 · outbound

This paper cites Scaffolding dexterous manipulation with vision-language models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Scaffolding dexterous manipulation with vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.147164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.147164Z digest=sha256:3428aadf75e38651e1009183a9f9ccc97a2341017f93d8325a4bcc41f6b29d11

Observation 184e660e-cae0-459c-a0b8-4082e56b2734 · outbound

This paper cites Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.204910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.204910Z digest=sha256:c199992c294449fa60316e7e558aa0a34df55e0839f628c8a3032526df891349

Observation 6029deac-11f4-4de2-8f49-475f5190838a · outbound

This paper cites Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.260785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.260785Z digest=sha256:1410b43f0a287b82cd0fdcac8bd0257dfd7f39b9562e7c20ad2d0d3211826323

Observation 9012a185-501a-4aa7-b078-aa123dc4d97a · outbound

This paper cites Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Dexgraspvla: A vision-language-action framework towards general dexterous grasping.arXiv preprint arXiv:2502.20900, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.321889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.321889Z digest=sha256:3ae4f1fd7b44aa2402eb47d56e59f95982e804dd0a73161f6e3031b7d8a49392

Observation 9c96aaab-977f-4a81-9243-af342b15fcf0 · outbound

This paper cites DexVLG: Dexterous Vision-Language-Grasp Model at Scale.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DexVLG: Dexterous Vision-Language-Grasp Model at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.445246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.445246Z digest=sha256:2c037fdcc07685ad7a5592db6b1d6591a338acc9253a456f7be02ae8b3b45e1d

Observation ddc36849-c43d-428b-b6ad-a01e4bd224c4 · outbound

This paper cites Efficient residual learning with mixture-of-experts for universal dexterous grasping.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Efficient residual learning with mixture-of-experts for universal dexterous grasping

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.511611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.511611Z digest=sha256:ac66394616e2617a3eba5fdcdc314711c6be0d2220cda11e72eb244dfc70c840

Observation 13d52d96-73ba-4d36-b6ac-1a52ae2cc874 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos R3M: A Universal Visual Representation for Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.565406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.565406Z digest=sha256:2855695b35ca2035add33308fcf7eadd5e3a505f005d97efaecbe57dd4a8f521

Observation 1c462fb1-22bf-4e43-905f-15c05c20b67a · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Real-world robot learning with masked visual pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.683285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.683285Z digest=sha256:a7f595b61417d48960c1c7ce3fb31b922d66ed67c43fb5b0fd0256515aa738c9

Observation 135da890-1f3d-4ffa-a1a0-76593fa951cf · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.742191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.742191Z digest=sha256:59b054a84504befa149fb1fcf1dd1f29e587e7a426c864fe65fab457bf852784

Observation d704d37f-98f8-4199-a36d-79442114c205 · outbound

This paper cites Improved baselines with visual instruction tuning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Improved baselines with visual instruction tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.817847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.817847Z digest=sha256:01fab808886c00d495006f57489087b45b17b19ba3a846997e2d4c4c897eecf1

Observation 2097f02f-9b4a-4ef1-aa62-ffc549579545 · outbound

This paper cites Humanoid policy˜ human policy.arXiv preprint arXiv:2503.13441, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Humanoid policy˜ human policy.arXiv preprint arXiv:2503.13441, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.886667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.886667Z digest=sha256:ca432d179037ce373265c45cf3e060a17bfe400c741acde2f04f1a665e75deb4

Observation 741b91a3-801d-4279-96aa-da1d43cc7a09 · outbound

This paper cites Integrated linkage-driven dexterous anthropomorphic robotic hand.Nature communications, 12(1):7177, 2021.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Integrated linkage-driven dexterous anthropomorphic robotic hand.Nature communications, 12(1):7177, 2021

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:39.964864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:39.964864Z digest=sha256:c19f10f97d56cb3e9afff15421fcc858e45ab7ac3c32f5d732f21e52b72f9e37

Observation e8b02f06-166e-421e-ad27-2187dbe4f505 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.066734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.066734Z digest=sha256:30be23583910b1e1a937995f3839fd0c144c9b5fa8da02be91d588894a504374

Observation 19f3022c-e89f-4ee9-bf75-12b6dc9d4e00 · outbound

This paper cites Autoregressive image generation using residual quantization.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Autoregressive image generation using residual quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.128808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.128808Z digest=sha256:57c372613efd4e995af9599884f09b59b6e173932b2c86b4c75a61ef92c6703a

Observation acbb01f9-ab9e-405a-9aaf-dc55f47d8d26 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.225699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.225699Z digest=sha256:073ffcdd1f5c6e3111089d37f37c2a1f4679ba3d38fc4534343a92d321d3d8c0

Observation 2ea0e8ec-bf17-403b-accf-af76bacde7dc · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.309722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.309722Z digest=sha256:0b64952a6fb2ca396556be34f4b632e10f45ed92d7131426420bd5356920d5fc

Observation 1988342e-245c-4e46-b539-17e5a62fccee · outbound

This paper cites Improving language understanding by generative pre-training.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Improving language understanding by generative pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.390951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.390951Z digest=sha256:4abcef859fbc7b086cd95e1ae854e31d4d9c886df5e2917ddf3b79659ab49bdb

Observation 5ed2d5c7-ac6d-4597-bfd7-852ab5fb44e2 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.524143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.524143Z digest=sha256:9c3026b9b5dbc8b8455d8c3b3b3cd65a8228466f732634d90164e07aa98b94f4

Observation 97e4cffd-b982-4108-a1c4-743565f197b9 · outbound

This paper cites Language models are few-shot learners.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language models are few-shot learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.627678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.627678Z digest=sha256:f74c8daa2e4983ff90d7b16ec7b72bb5c8d3eb158bcd3e1574c66c1f75f94c56

Observation 60782969-dfa1-4753-9d7f-d8d02d600f00 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.742365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.742365Z digest=sha256:f4d2c051b1dcaadaa68ceb272ac4332248653a04830dc0660aecb05fe4eb34f8

Observation 4374b8c0-d9c7-49ff-b2cb-24044a597685 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:40.877665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:40.877665Z digest=sha256:42b6446353318567cde1430ab38f0beaa12134a832e65edfe831195f13311746

Observation 8c0f0ddd-94f6-4668-a86d-3723dcc5bfa3 · outbound

This paper cites UniCode: Learning a Unified Codebook for Multimodal Large Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos UniCode: Learning a Unified Codebook for Multimodal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:51.856771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:33:40.966946Z digest=sha256:ded39ef7e5e85331ece8c137c748f3656309747943cc3d51a0892c175e403ca7

Observation deb167aa-01ee-46b0-9fbd-898b7204f23c · outbound

This paper cites From pixels to tokens: Byte-pair encoding on quantized visual modalities.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos From pixels to tokens: Byte-pair encoding on quantized visual modalities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.028548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.028548Z digest=sha256:c39cc8a3790fbc742bf2fb6682142e2b6c5ce60f709765ca6eab01525c7bda08

Observation adf0c352-4347-4667-bd28-45f795d9c910 · outbound

This paper cites Unified multimodal understanding via byte-pair visual encoding.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Unified multimodal understanding via byte-pair visual encoding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.136954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.136954Z digest=sha256:efafa801cd49b9be39a4e0af1a90bc195c6c072b166cc0b091ab44de8a87137e

Observation 066c0331-e95b-4295-b5a3-562d86f09264 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.212733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.212733Z digest=sha256:a8534a5872ecb717feb447fa1d958f447b479c70100b90b72fe2b1e3863ceb42

Observation 0342fd5e-004f-42f7-ac47-d2370206bec7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.297194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.297194Z digest=sha256:c0f9d443f493ff56f084871ceb29ebb0b7fd4bd9a4d1869edbbfbfdb0618b786

Observation ad47a498-55cc-41c1-8ab6-3ddd7213b88d · outbound

This paper cites Qwen Technical Report.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.390792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.390792Z digest=sha256:bfd3d7cacb0c53988722e62b5edd24975c5ee7115757cfb4e49c37725af7a5cb

Observation 16384af9-1a5c-4ebf-8af8-82e3d0a36ef1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Learning transferable visual models from natural language supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.519049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.519049Z digest=sha256:c419f14860fedda0b0760bfd872b2242b8695748632f6b5e77e9894d8a8948d7

Observation 4aa01e2e-e516-487d-b31b-24f3b03e132d · outbound

This paper cites Sigmoid loss for language image pre- training.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Sigmoid loss for language image pre- training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.612755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.612755Z digest=sha256:813a9f44c0cf6ebf25fe071f6e892a1dc0bbef6c6d9306b5216950b2702cb4a6

Observation 795e4fc5-e425-4936-b631-b248f5c0f6af · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Flamingo: a visual language model for few-shot learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.653621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.653621Z digest=sha256:e93af1a998fac95e7204e5fb1fc02fa0c94a9bf8cd086407f0753170c185825d

Observation 73eea413-02bd-4805-9e29-ebe6079d4b4b · outbound

This paper cites An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.744204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.744204Z digest=sha256:bc82e1d9151968645f515081666dee59367c707140a0eb57c4231a6af400e90d

Observation 4e4c7e95-305b-4de5-a648-8eb8351a6023 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.848263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.848263Z digest=sha256:5e1423f4f377b226a55fbf89e572ec6b3128454a1d2bce59905907a850f2f254

Observation 2c79e090-2cd3-4c00-8792-c604cdfd504e · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Otter: A multi-modal model with in-context instruction tuning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.991890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.991890Z digest=sha256:fc5a80b5b8eeb6f2aec6147633e430a7ba7e5f55f0606aa8858cddd3872afbaf

Observation ae659049-9bdc-4c7c-a9a9-ae018ab13db5 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.059251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.059251Z digest=sha256:3a2f2edf6f9e3989198df952b6e86a0842c3a7a87d36fa129f56fbf0a51d73c3

Observation 27a8bc03-cdfb-4104-b4c7-46b0ec370c46 · outbound

This paper cites Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.174500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.174500Z digest=sha256:ee024e62dd8bc5c7e8342099067cb5291d273d9a1c80252571080c2dea069bbf

Observation dc1b8b98-c779-4bba-ba14-eacaaf7e8583 · outbound

This paper cites VideoOrion: Tokenizing Object Dynamics in Videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos VideoOrion: Tokenizing Object Dynamics in Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:51.694013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:33:42.245773Z digest=sha256:3e54c23b11f2cd3bd31b7a149db733cd385c9ee46bd381739856dcd046ffbec5

Observation 38d25635-d466-4879-9c91-7c7a075b8501 · outbound

This paper cites Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.364931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.364931Z digest=sha256:e620621bee04c2fcdd1e5b1c76f44280affba98feff83f567e497d237cf70e08

Observation f2b40deb-04fd-4eac-9224-ba01623ceca6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Gemini: A Family of Highly Capable Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.485380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.485380Z digest=sha256:eb4e5ae35acb0dfd1902836053185f46e63fe25cc3e849956e288baaea9a91f2

Observation 4144850f-7008-4400-b3f4-71b67d53e5f1 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.589633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.589633Z digest=sha256:a46be507915c761fc603b8331272f6111fcd5e4f0fcf1d136d6835787793c5e0

Observation 7303ebdd-6754-4c31-abb3-8adc9def0a10 · outbound

This paper cites GPT-4 Technical Report.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos GPT-4 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.732915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.732915Z digest=sha256:d3e7193e90121d3b2adea66cf91bad1524fe92e2ea156272808b12773cf709cd

Observation 0d05e1ec-5676-4ae6-8442-5ff22a4731f8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Qwen2.5-VL Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.826045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.826045Z digest=sha256:e777f0fa7890fd5be808a9830adb9cf1705e324516ac6700fb6a211a6ad00f2c

Observation 2359026e-9846-4af3-baa0-3e9031579db4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.934724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.934724Z digest=sha256:582d76eeee81993265cc958d4ce19ea18ff36a0d46db32087bfc98fc78554131

Observation bdbd68bd-fb7c-474d-b25b-d70e32d48fb4 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.026591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.026591Z digest=sha256:a70a4ef218b8936e47ddae70c36090b5a68e00a8de630cb65651732a8438f976

Observation d9fe530f-52cd-4c1a-99ab-7acedadf48b7 · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.146808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.146808Z digest=sha256:c77d639ee651cd2270c8133a1e386f77300c35c5b1ea74ce93a589664efbca1e

Observation edecf294-eb4f-4955-bde3-9c9d01281c77 · outbound

This paper cites The kit motion-language dataset.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos The kit motion-language dataset

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.239350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.239350Z digest=sha256:ee08eea2408cc50310021e045a7d1009ba75963396f8269f9c8ea97366ca524e

Observation 6faaf8af-74e8-4ad0-a2a7-ccd77fdca5b5 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Amass: Archive of motion capture as surface shapes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.309159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.309159Z digest=sha256:ee7d22389d9489142be848ffe39cdab6d31f7aa7921baf0af9c963a2bdf40b8d

Observation d53026f1-1d85-4640-81e7-8aba6ff3d6ff · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Generating diverse and natural 3d human motions from text

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.383085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.383085Z digest=sha256:bb67d6717598d480387d305aef350cf99c2eabcf41f1615952cf3e02a36b803a

Observation 60971193-2d7b-4864-a387-f741b9783ea8 · outbound

This paper cites Babel: Bodies, action and behavior with english labels.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Babel: Bodies, action and behavior with english labels

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.482132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.482132Z digest=sha256:d0221bbe83d98649605544ea802305131ca44a899295a559c1b014b1b363b423

Observation 5ee93827-7900-4bec-87d3-aa21d56eef54 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advances in Neural Information Processing Systems, 36:25268–25280, 2023.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advances in Neural Information Processing Systems, 36:25268–25280, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.571970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.571970Z digest=sha256:ace50d5429061db7ede87a7ace145ab9c3df5197488859c17c5d32f09e7ba6f0

Observation 1e470422-15ee-460f-a820-a5e001100406 · outbound

This paper cites Egobody: Human body shape and motion of interacting people from head-mounted devices.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Egobody: Human body shape and motion of interacting people from head-mounted devices

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.663928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.663928Z digest=sha256:963462e32d1bbef3c4bb066877b10e151a72279d2a30d21a2e1d60a158654a39

Observation 18c5d703-6a66-4c4d-aba1-7872741a2f59 · outbound

This paper cites Nymeria: A massive collection of multimodal egocentric daily motion in the wild.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Nymeria: A massive collection of multimodal egocentric daily motion in the wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.756773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.756773Z digest=sha256:35edec38163d5233e07344f3838916cfce10188457f4d932ac62d31963210b29

Observation 5a028043-3caf-45a2-948c-f1761def8109 · outbound

This paper cites Scaling large motion models with million-level human motions.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Scaling large motion models with million-level human motions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.859592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.859592Z digest=sha256:68c4170a5d756f8c6221e0c8eb2d0a3d173e5bdeb16b9d5d193cddb35aea0581

Observation 9e5e0376-e134-4548-85f5-bed4f8b76961 · outbound

This paper cites Smpl: A skinned multi-person linear model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Smpl: A skinned multi-person linear model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:43.946231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:43.946231Z digest=sha256:b5e38c01fcfa30b625c2ec1e0cd574f5a77309b3e3dfa476c4501747f3f4f3ea

Observation 8bfadaba-b22d-4fbc-9e10-e9213c25b9d1 · outbound

This paper cites Expressive body capture: 3d hands, face, and body from a single image.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Expressive body capture: 3d hands, face, and body from a single image

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.086686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.086686Z digest=sha256:0c967b49235a5fdae4948837aecd478c422a0ee73dae7b88737e47faf4ef1d0a

Observation 1d7f0419-75a2-4b7a-8313-4339200e13ea · outbound

This paper cites Human Motion Diffusion Model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human Motion Diffusion Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.187290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.187290Z digest=sha256:6e5a4fa4336d02ba75d1194295e79785d6277a15868babe94fb205deb3d9ddb6

Observation f64da1b0-6537-49c6-ad33-c267df7b40fb · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.341154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.341154Z digest=sha256:1dc676be5149589822cc34b0c9da1d02c764dfffbdb0318c9d2e682c9c42c52a

Observation b1483f9c-a7ff-4d5e-9aaf-a978a0a6bfe8 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Executing your commands via motion diffusion in latent space

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.449291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.449291Z digest=sha256:d219f38d8529339d64088aeb3133831363de894e4551d5b6e6d30a4c016e2de7

Observation 4e06db4b-fddb-477f-9953-fcfc76856df7 · outbound

This paper cites Physdiff: Physics-guided human motion diffusion model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Physdiff: Physics-guided human motion diffusion model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.595272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.595272Z digest=sha256:d0af835e31385241c14ac10ae1f384ced9eb411c93764babd4bcb7c7d747aa5f

Observation 0acc39f5-fa8f-453b-8df1-d67450ec27dd · outbound

This paper cites Remodiffuse: Retrieval-augmented motion diffusion model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Remodiffuse: Retrieval-augmented motion diffusion model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.691840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.691840Z digest=sha256:fb142b48e4422639114a8f3b48b4634855d4a10da491bf8611f20f6b26270901

Observation 59819418-f3cb-4b2b-a203-e1ee2f99d6bc · outbound

This paper cites DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.817249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.817249Z digest=sha256:92380f0241989bcddcf7d3a6cbfb8557d5f0b0469b232a2c1634222f68458648

Observation 682795f0-2e56-4979-9ad7-04e47722f7b2 · outbound

This paper cites Large motion model for unified multi-modal motion generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Large motion model for unified multi-modal motion generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:44.965632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:44.965632Z digest=sha256:e16ec2ea7e418861d3cd3b347846adb9dd18b3afc71948a06549c06bbd658789

Observation 3781b643-41af-4e61-9e88-dfb34367abfe · outbound

This paper cites Neural discrete representation learning.Advancesinneuralinformation processing systems, 30, 2017.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Neural discrete representation learning.Advancesinneuralinformation processing systems, 30, 2017

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.087540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.087540Z digest=sha256:13df18f7574c1445173d2cca955bab9e27761b7ca2142cf3bcc55a761cd47eb9

Observation f2ca6c3c-e873-4be6-8e5a-76ab3fc3d47e · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Generating human motion from textual descriptions with discrete representations

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.136041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.136041Z digest=sha256:a0e250ffade46bdf71b5efc9f2965b91a4d0fb3ae415aa4bea0e9d8942fe8970

Observation faaf87c1-c4a7-4048-ae4f-ba422cd2cc71 · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Momask: Generative masked modeling of 3d human motions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.196778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.196778Z digest=sha256:4b604ebd1bc56ab13c352db464e17e2b81192769b4b0ee3e2b79c1153dc26f4f

Observation a6c50ce6-706c-492e-82a5-fb1296c08074 · outbound

This paper cites Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.323096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.323096Z digest=sha256:fcd1124b049e73e4bdca78e9274bf6eec6fdd41760322ba3d303523dd18250cb

Observation cabb6102-fbb4-441b-a064-d4dd6e18247c · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.392908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.392908Z digest=sha256:41ff5ae6e9c7829c36350b8768cdce23390ea4e5949a5d8d2f17b608476477f8

Observation 9fde13bd-7364-46ce-a8b4-bf498e37734b · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Finite Scalar Quantization: VQ-VAE Made Simple

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.489989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.489989Z digest=sha256:74c90ea81ea170a023d77262857e0aff0be16ef469872e80164b0a2de91f7efd

Observation 42a3761b-fb1e-4cd7-9020-57caeba12fd1 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.572085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.572085Z digest=sha256:1d4ff6ec621a450f321f20a7d09ee0836e3320240bec8c7ad631658fea4fe6c6

Observation c9e1c9c6-8b7c-45e9-8f6f-66736c3c6272 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motiongpt: Human motion as a foreign language

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.640936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.640936Z digest=sha256:9f253e55c2f226e19afa1c759e7fc6377f91a4e51c06d795c77aa3c1fdaa5aee

Observation df7acbec-d0c8-4121-9b14-ad743b74c024 · outbound

This paper cites MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.729880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.729880Z digest=sha256:a3908e9dfbd822cb8da08b536c6c4cfc1fd3e547fe0a5ee29fd745fb340cbf09

Observation f7cd73b6-a9aa-41d4-97c9-144cc1f960c5 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.839617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.839617Z digest=sha256:8700fd2d5099f2fec22cfdc107f62cc52698732b34075bab054ee76ea7e3af03

Observation bfed5262-2915-4001-b16e-cc2be9cb474e · outbound

This paper cites Avatargpt: All-in-one framework for motion understanding planning generation and beyond.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Avatargpt: All-in-one framework for motion understanding planning generation and beyond

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.967329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.967329Z digest=sha256:28fc0e2dd98f91b98e7f42d8bf6d9f52c06b580a0e3fd8a9612408011cef2d9f

Observation ba2656d6-4ed0-42c8-92a0-fb815962cbee · outbound

This paper cites Motionchain: Conversational motion controllers via multimodal prompts.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Motionchain: Conversational motion controllers via multimodal prompts

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.024326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.024326Z digest=sha256:eb03e5487448fde1ed794179fa8a1c98baa382a4ef1cd01bac9a221f1c76784c

Observation e9b8a8fc-126e-426a-913b-64b2385a169d · outbound

This paper cites LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.059708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.059708Z digest=sha256:c450ba1f1b0dede6d083f511d3ad01cce4b528c52e14e8523dfccbc179426a45

Observation 5d245ff8-9b14-45cf-bd77-9f159ac67b35 · outbound

This paper cites Human motion instruction tuning.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Human motion instruction tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.130864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.130864Z digest=sha256:c9a643cd2c07994eb99ca7072c5d25bb7ecf523e54f6b08d25b669b7de519ff5

Observation 8a7212e1-8afc-45b2-a49d-cba95fb4f3e0 · outbound

This paper cites RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.196891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.196891Z digest=sha256:d6df1e59865a20c6d068a768d8cce824a34fa04da5af5014debdf3e677afd2ea

Observation 4a9fdfc3-20bd-4aa5-a976-0eb87aef1e2c · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Perpetual humanoid control for real-time simulated avatars

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.255945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.255945Z digest=sha256:5e2a07fdb7e8b27e0743ccd53444092d802554fdcba7514be81d60888e47f822

Observation 2b6f6754-fb9c-4da2-ac3d-db188e302e23 · outbound

This paper cites ExBody2: Advanced Expressive Humanoid Whole-Body Control.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ExBody2: Advanced Expressive Humanoid Whole-Body Control

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.336690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.336690Z digest=sha256:f68390264aec73adaf90137fcb5bbacd268dfe96757537979cafde6e9c11bf4b

Observation 6dad6a80-a450-43de-b8e7-b6141c91e102 · outbound

This paper cites Reindiffuse: Craft- ing physically plausible motions with reinforced diffusion model.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Reindiffuse: Craft- ing physically plausible motions with reinforced diffusion model

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.386963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.386963Z digest=sha256:7dd5c07e3dc4a6005316e111eede8cfdc09fc09449c96d7449365d036622ff30

Observation ae8d7db9-e877-4a54-8368-084b19507e00 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.451798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.451798Z digest=sha256:87274d2286753e3a13f6d8d25ac7da903afff22e79c25c96d2619bf7cb83ca35

Observation 900fbd39-23f3-426f-9049-b2f95871ed97 · outbound

This paper cites Hand-object contact consistency reasoning for human grasps generation.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hand-object contact consistency reasoning for human grasps generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.508374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.508374Z digest=sha256:0a0876bde5baa6e7180790417279291c6f211231f8abb8758f9d1b5ccc17495b

Observation 91a3701f-955f-4d51-b7f4-43fe3e148d98 · outbound

This paper cites Joint hand motion and interaction hotspots prediction from egocentric videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Joint hand motion and interaction hotspots prediction from egocentric videos

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.541539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.541539Z digest=sha256:2baf0e7ae23a22f660771ceca2a31743296a84693a23c48ab21e8a6ad68d989b

Observation 079824de-be87-4b1b-97ba-a373f732fc17 · outbound

This paper cites Hot3d: Hand and object tracking in 3d from egocentric multi-view videos.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hot3d: Hand and object tracking in 3d from egocentric multi-view videos

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.564839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.564839Z digest=sha256:87c51d87b01752ae2403c90fe368635ac50fcca0c969c969c8425bcb7c2707e6

Observation b19ece4c-6d3d-4fda-867d-6e11e371444c · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.641803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.641803Z digest=sha256:9ec1bf120e19d24d8be84ac7adacc4bb89db263d025a314100e62e4e459b25e2

Observation e9bcd271-54de-4248-a716-a2d7b3e2b4a0 · outbound

This paper cites Oakink2: A dataset of bimanual hands-object manipulation in complex task completion.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Oakink2: A dataset of bimanual hands-object manipulation in complex task completion

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.644330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.644330Z digest=sha256:f678af4deba36624baf0c8a0d4cd7165c4a9373d1864b554f40bc4cba1356f55

Observation 21d286b6-3268-443e-8562-e447361b5c5f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:46.663820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:46.663820Z digest=sha256:3c21e0e7c91066b7fb9240f614b0d45474a03812a4df1b41ea032b7806ef1c0c

Pith citing papers

Observation 8b59d164-430f-4446-9e55-2a5eac4ef086 · inbound

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration cites this paper.

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:16.766190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:50:16.766190Z digest=sha256:4af9d1fd99346fb815460725e28d6e544b6209114cebd80bec0074f0ed544a3b

Observation 456959d4-8736-4618-9dfb-3b2ebb6116e2 · inbound

BiNoMaP: Learning Category-Level Bimanual Non-Prehensile Manipulation Primitives cites this paper.

BiNoMaP: Learning Category-Level Bimanual Non-Prehensile Manipulation Primitives Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:52:38.381035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:52:38.381035Z digest=sha256:cd12a048204da871747ec7a9aea9ff439b8d870feb669a16f31c78426873531e

Observation 8ea14750-e58d-4048-b56d-b8b74d8d3eb6 · inbound

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model cites this paper.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.531852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.531852Z digest=sha256:a0a2058e56fd34d528c7d1c284be27a0e363f9b62f90df569b7931891fb2b836

Observation a8f85282-66ee-412a-84da-64e4d1894b45 · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.265883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:68db08b2dd86a890c05432b9e9cd35e39e206e519fc2bac83106f27a1fa7dcc1

Observation 228ccdd7-920f-40ef-bcab-b572bd8fc53a · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.870351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.870351Z digest=sha256:cf4091485de65d8df1f23f6828dd48e9dff99c0c646fb8398f3173c2848c52c4

Observation 14292983-4c79-4716-84d6-caa3591e71cf · inbound

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter cites this paper.

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:18:55.692685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:18:55.692685Z digest=sha256:038e88b4a77a20ebe4d3d103839bf9196631df4b5cb14c4ab57ba2efbb422ecc

Observation ce3dc394-c737-4508-8cc6-ddc2e5d46e0c · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:02:34.356382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:67f4de4d09e039ffb3df7170a883d49e35c4c2b2e0c48f12235542a77156c3b1

Observation df45d407-7e35-4e6c-ae92-b93786349036 · inbound

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models cites this paper.

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:02:24.979717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T06:01:30.803128Z digest=sha256:30adc0ce7f3bdc04189c592200893871282c180bef3ed38e8778352bc6ec8abf

Observation 4bb248b3-fd46-4cde-a901-5771d349e93c · inbound

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion cites this paper.

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T23:57:44.645995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:57:44.645995Z digest=sha256:b5d4b566b5fdcd0ac336d331d30a4b253e04a05c8c1753a18cfd3c766353c887

Observation 66e903d4-96b3-4591-9df8-7feafe535f33 · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:30:20.893877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:f3655273e215d49641bd613cdecb291ddc13bf2526c03d341fae46d293dce2d3

Observation 995409de-51e6-4c60-8313-a0d448e7afc8 · inbound

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves cites this paper.

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T21:11:31.022938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:11:31.022938Z digest=sha256:bdfbe0d22a45d911c1abddc92550c109931cd0c23cd69280db1481c4d4cce925

Observation 3f1506ab-6795-42ac-a247-a675d0621889 · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:57.002389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:07:41.489995Z digest=sha256:ad5d17438588bd575bcf04600f69e7990558b7d4ef798b2d2eaceac5230c7a43

Observation 0d1f7b55-f0f4-4ffb-a777-628617c0a3dd · inbound

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World cites this paper.

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T08:25:22.011013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:25:22.011013Z digest=sha256:ad2203730dac0529cbb98e7ff864d2029c7b7315827eb2a70c4cbbfbbc458854

Observation 792b1e76-3bab-4a37-bf0e-e6655d513d82 · inbound

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment cites this paper.

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:03.890345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:36:23.197843Z digest=sha256:ca63566a1be154729ef2f1c8e13791ce078e303efc72c63bfbb85f03f3f60a42

Observation 64153462-9782-4f60-8a31-79fdad4c8103 · inbound

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies cites this paper.

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:30:22.689485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T12:29:38.306670Z digest=sha256:a5ff9a433a2dd06fe6e708bf69e605707dea3c224048de5504441d9c746929c6

Observation 2a2c6c44-858e-4db3-8980-fbea3cac73d0 · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:29.409174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:5a6680e303b9112731330a1ad7968bc956be7706dae99d0340b1ef3a03c4ef2b

Observation fdef3828-756b-4c33-8992-1d37bb0ebb4b · inbound

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks cites this paper.

EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:16:11.740209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T06:19:34.743194Z digest=sha256:972d8ff6fe656968301998881c616b9f6622ca3276d49239cdca1f37d18d8dc5

Observation 8fceb112-ea8a-4399-9ffb-052c51f8bd42 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:11.725302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:e1416e14ea8b16a0d4dd8d1bb7a25f8a8565c8776a6cfd574ce098e489d415d5

Observation 17e17457-c4f2-43ba-8fd9-251ff72234fa · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.156747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:a28db5caca2302d1fe7d9afcb965d91e0554422ee9930d49ab8fd5f7bd24ef82

Observation ced91d16-378a-404b-8909-be2b52314578 · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:56:08.488051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:b9deae8baf64e9a99a655ec5a8d91133c4da828242d976a3949fa729b5e95406

Observation 13379a2f-4b85-4ade-9d08-19553221a7e8 · inbound

World Model for Robot Learning: A Comprehensive Survey cites this paper.

World Model for Robot Learning: A Comprehensive Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:11:04.931247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T20:38:12.709629Z digest=sha256:afd971023210026e7eda1c3b61df772b8233d71e958a44f02c51b5904950abbf

Observation fe5defcd-459a-4c6d-9907-f07a38569958 · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.148923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:6396747ede4803e8443713ac4019885b62d24430bf2413317b173b4362473ee3

Observation a47c22ac-14c5-4628-abbb-cb6baa225d67 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.846258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:cda06b7402abb105b70ecb996608d97e6a2d478aa1eb1d0727cc7008bc092ec4

Observation f15d5d9b-626c-4120-bb49-7f6ebdfa7de5 · inbound

Towards Robotic Dexterous Hand Intelligence: A Survey cites this paper.

Towards Robotic Dexterous Hand Intelligence: A Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.046963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T05:44:23.945064Z digest=sha256:87c3144d25f438f7722a551f43a8151b9c54ace884248707abb7a7b8d533c917

Observation e6258c08-aa48-4067-98fd-59a673b365a8 · inbound

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention cites this paper.

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:09:42.122084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T03:09:26.858170Z digest=sha256:d22343bb96d56462cfa05e5b02c5726c85e395d4d7e243329361a68755256b5c

Observation 51ee9ed2-a371-49ed-b5cc-a6fd44d4c33b · inbound

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention cites this paper.

Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:34:05.565221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T08:31:57.077346Z digest=sha256:5c309059efbb54d5eae2991b2dabad0900e1865b41928e3791d3f79f217e844b

Observation 7b6ac8ad-7cde-43e6-9872-23eaf06c23d0 · inbound

Dexora: Open-source VLA for High-DoF Bimanual Dexterity cites this paper.

Dexora: Open-source VLA for High-DoF Bimanual Dexterity Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:48:11.778910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T09:43:53.032153Z digest=sha256:07b9b44b0cd26e2d18c51faff6aaf07c70eb98d1f489d6d2387c35c622e76170

Observation b88443ab-14a4-4068-9482-5c141b4c9267 · inbound

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models cites this paper.

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:24:04.188694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T00:18:33.150920Z digest=sha256:168276b24380bb5a95e558857e60af8ad02421a31adcd5240fc4afcce845adef

Observation 3b4d04cf-4c1d-4503-9490-b72a87ba08e3 · inbound

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models cites this paper.

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.768238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T07:18:59.266263Z digest=sha256:5500eaac00b9d32c55200834d282673ba04dd30513084285a0af620801b189b4

Observation 7ab91d45-1471-4c74-b056-ea02fcfc11e3 · inbound

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data cites this paper.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.210434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:2cf9578d42f4040ed1b906a62f93e4f4793cb646d724c2f46316a8d55eef7fbf

Observation fd4dd992-0dc2-4880-be98-213785842ab2 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 220

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:27.773712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:f57ff971dd2c6d42df97c1620494c31cb70eb597295de2eec41e3d56119b97db

Observation e1fb98e0-3d75-4e42-bbdf-70f7cd08c5c7 · inbound

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation cites this paper.

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.363432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T10:28:22.430294Z digest=sha256:0f57dcdce651f9f7f76b82cc956a636a97189c5f85782c4f987d25bfc209da29

Observation 9bb1aec4-8dd7-4f1e-8274-da739a390b14 · inbound

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis cites this paper.

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.909752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:05:26.382826Z digest=sha256:5ed600752438f1416e58506b70ba5f047a061cc3405055ea1e687e0a21e02557

Observation 389ecd41-aead-427a-bbcc-05e33a2970c7 · inbound

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning cites this paper.

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.100854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:44:32.716179Z digest=sha256:4bc52805bc3b9d68ec7c3a46b4b64316c6e1c7205ae510f11657c938b162a2b2

Observation 79d0f620-2bb8-42da-992b-afc11cad27bc · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:09.505386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:25:17.522240Z digest=sha256:2e19d445a1dc7ccb9c1feee7b1c12c11e0063470e9474db14dfae0caf949c339

Observation ca0da9aa-21c0-490f-bec8-32def5ec9703 · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:35:34.493805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T07:17:42.045939Z digest=sha256:76ce0ebfc4c44d27d95b3c543f5ff0caa9466ab1a73be33609eb52056bd4b93e

Observation 462e6594-bd69-49ee-bb7c-2f578b6c608b · inbound

$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models cites this paper.

$\omega$-EVA: Envision, Verify, and Act with Latent Interactive World Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.742396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T16:10:02.176202Z digest=sha256:429120bc1840448c1239ea7058a14d40fc024d7127bf57ae97d761f032e4c323

Observation 18ae0df7-1766-4f47-9266-4fe04473699e · inbound

Next Forcing: Causal World Modeling with Multi-Chunk Prediction cites this paper.

Next Forcing: Causal World Modeling with Multi-Chunk Prediction Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.513373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:31:53.905704Z digest=sha256:f70f2cfb27f734e253ab231d8de2db7cf467ea1c5563c958bead9204ce0f2242

Observation bfee454c-5b31-4cba-9671-a3779d34a43c · inbound

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition cites this paper.

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.220598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T10:02:54.183839Z digest=sha256:18e89b11f5d405a47b66a1ddc190f890c7e0853e1eebba2fc03a5ff544ad4c0c

Observation 67082a9d-78e1-4134-b360-31a571b052fa · inbound

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models cites this paper.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.855989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:7f3a976fa0d1020bb1c38813dfae7775c6e10f0e62bebacf1477ff16acbd6c8f

Observation b1983b03-9216-4236-8dcf-73e1abbe4232 · inbound

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos cites this paper.

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:19.497076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T20:51:21.209882Z digest=sha256:51d3a2963dd1a1193060eefbb75346a4e67af8fcf72ced70b6b0789f67fa0162

Observation 25489053-9cea-4a60-8674-87b32178ba29 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.363664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:8fb9e89f29c207265fa37743498ae592186c3e67152aa14e78a2d0b1006605d7

Observation 6ed30d62-4531-400c-b6a4-28aeb16a039e · inbound

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining cites this paper.

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.683310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T17:53:24.431287Z digest=sha256:ca3dc8b4c3c16a493abab46a799623c28ea4e43e2a07fb9a6221342c353697fc

Observation 73810ba6-6a3f-488b-922b-b89b9011c5fd · inbound

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data cites this paper.

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.736280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T11:47:54.951618Z digest=sha256:82a8af8813f248c87148b6a93c057462d60eae3936c2ab8a5d128d85257f7c95

Observation 8e24bd8d-ec6c-4acb-8ff3-dd8b1f91a3b6 · inbound

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos cites this paper.

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:58.480120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T23:57:48.395172Z digest=sha256:758d5f14680723b8f87c257eee6207c3a120c02c5951dfd14d11bcf2bc9522b0

Observation d836be21-89c4-4eb5-b6bd-f59877630a17 · inbound

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots cites this paper.

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:55:51.368357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T04:23:04.622902Z digest=sha256:609f4905f656675370c3d94482a4a7672a2d2de1ec50435ffbf6fd5f8a75fbf5

Observation a6366124-29b5-4006-a58d-fff533a68f5b · inbound

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments cites this paper.

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.344117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:02:17.092213Z digest=sha256:419b232dce3a04b710795d13a1174cff220f6b095787c4095d906092a4bb4372

Observation f5181d8a-4432-4e9d-b577-41cd6ea9a2ff · inbound

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation cites this paper.

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:26:54.003775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T11:18:32.259684Z digest=sha256:7bfb88f71906516f15396fad1b14b5fcd6e25c6da0199f570169de057b09730f

Observation c898bc55-29e6-4ff2-b04c-63542be34257 · inbound

From Foundation to Application: Improving VLA Models in Practice cites this paper.

From Foundation to Application: Improving VLA Models in Practice Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2b91211b50495cd1dcee08b1b0070accfc4dee5619c8357ad1ce16eff65740b3

Observation caedafdc-980f-48df-93e8-8e0073c6f14a · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:16:48.670270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T04:12:29.153764Z digest=sha256:15f34d8b5e8f60f588fc8b22e9b0a7519a32935e10ebceda2c8ee745f5e0cace

Observation f742ffaf-db65-4417-84e8-dc83c1b193ce · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:38.754872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:38.754872Z digest=sha256:d4fb6d86f727797467baa3b241f2f79bf4ff48634e75722bf85ac3fa6f5e82df

Observation 2766f5a6-6e0c-4c25-8cc5-2eea594df53e · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:3d9f6aabbe4e2f7a35754454c4d7c7b851439a6612969780f9c39b3d1ba12c87

Observation 2138be29-39f8-4103-9629-8bd6580c280d · inbound

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning cites this paper.

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:25:28.115858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:25:28.115858Z digest=sha256:8497f109a2f545c913e4a4871a422678de650f8715230b4b019792c3142a0902

Observation 5efc2714-0f33-4327-81ef-d8033792d4ec · inbound

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories cites this paper.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.119524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.119524Z digest=sha256:b87d4327586332f0c0f57f7447aed7145f01624702a7cc746570ba05b108f0c2

Observation 58695d96-07e2-4795-85c2-02cc499f5d1d · inbound

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis cites this paper.

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-01T19:06:43.487945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:06:43.487945Z digest=sha256:49c00ac4112037af6cc6468c6d6e154d059ab4b7c5f870c10a011d079e83c177

Observation baf6ea43-c392-463e-9ea5-59d2c684839b · inbound

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments cites this paper.

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T23:29:10.712814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:29:10.712814Z digest=sha256:949dc1797b5a17f52cc8e7a0aa9ca4e70d9521b982b41cbee5af46bb9d5b873d

Observation fcd9e4a2-c9cb-44ad-a3cb-036e28beb184 · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 237

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.854691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.854691Z digest=sha256:1a2e90189112ed547a3a0d4dd1f5a2721fe91ee89e398da9176642ef3f0cf0f0

Observation 012798cc-eb2f-46ec-a04b-3a3a4f9b30b8 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 227

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:29.293605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:29.293605Z digest=sha256:b04e6117b68154dbb0119ace56855f4100ef14611b46ebacc6e3d416b639ed5b

Observation 719960a9-0211-4c33-8590-b2815129ed21 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:10.898271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:10.898271Z digest=sha256:2a2abe7a680c2daee8ff7e90ad50ff02bce6d11d8a77a37d0454f0d012fe56d3

Observation e0a809f7-8b0d-4924-99dd-b9e9b18e3015 · inbound

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph cites this paper.

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T18:35:26.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:35:26.658756Z digest=sha256:ff7311e0bb4b18d71bea97124a46e77d8cfe32e1c65aba389f01b1a6b852b326

Observation 6434a5c5-d393-40ae-b61a-1c6cd850d6d0 · inbound

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph cites this paper.

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:10:14.892153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:10:14.892153Z digest=sha256:0637622586d45aebf1ff943a1fb69fe95ee7c0399cba0f8a4ec4b538504f65a9

Observation 67c89dc8-ab09-47eb-9697-ebc544dd42f5 · inbound

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data cites this paper.

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:25:52.097084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:25:52.097084Z digest=sha256:68d5eab4d7ecf1225b0948d17ed26886bbcf567e0a031693bd416414dcd7fb14

Observation db8ecbb9-fb0a-4255-93ec-bd6f50059c1b · inbound

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units cites this paper.

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:58:35.689675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:58:35.689675Z digest=sha256:d1c4531467dcba53024a508a75720cda01088589f9a96351eb750269065c3af1

Observation 2731ebb9-479e-4c9f-bfcc-0eda9e53a124 · inbound

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation cites this paper.

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:39.903042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:39.903042Z digest=sha256:7f96a93b56dbdc7168d71adab1b63462450aac059ed3c22c2cbd5f71abc86986

Observation 36d4981d-d08b-4f6a-85bd-92b8aed13ce4 · inbound

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation cites this paper.

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:32.515534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:32.515534Z digest=sha256:78b6b11b44ac1d057dd5e5b04f19cb0862b2abe91f453578823bfa948f183b70