Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:55:03.686435Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2508.07863.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:55:03.686435Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:25.303896Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T13:47:57.512626Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dece0443-f9d0-4aed-8a4b-fb9ee5b6a50c · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Momask: Generative masked modeling of 3d human motions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f6a5fe-b310-4021-88d5-e9b0c9f65b11 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Human motion as a foreign language
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff48fa4-f852-4f2c-843c-a6a25f92051e · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating diverse and natural 3d human motions from text
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0611e9-d169-49ac-8e34-ebb202c2a7a7 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b3db17c-7c61-421d-81f4-80040385add2 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c21bda0-ecc1-410b-b15d-9684fde545d4 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Improved baselines with visual instruction tuning, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be59e44c-9ed1-4551-8228-2bee0d92fc49 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A large-scale rgb-d database for arbitrary-view human action recognition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5113d07b-6e8d-48e6-b842-1a78fd74b251 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 74225f25-b5f1-48d9-a88a-24eecdcaf43f · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Scaling large motion models with million-level human motions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2fe7aa93-f7bd-4160-a6be-71415b12f3f5 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Autoregressive image generation using residual quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01277ddf-1995-4ea3-b713-52a20ec98685 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Temos: Generating diverse human motions from textual descriptions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c7e44568-127d-4782-b635-1021a5a3175c · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language2pose: Natural language grounded pose forecasting
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 41af36a7-045a-4bf5-b95d-ddffd17f0d17 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Finetuned llms are general-purpose motion generators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f82544c-813a-4c9c-b787-9e4c43a67498 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5109d80-5552-4ce4-8338-eef799e9e219 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recurrent network models for human dynamics
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 34812253-a8e0-43bf-ad8c-09a39ec1b6b2 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A neural temporal model for human motion prediction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation add6134e-aa7e-4f2e-996d-651df9ac534f · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A stochastic conditioning scheme for diverse human motion prediction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f2a32cee-4d38-4827-9e1b-3a6acd8c1de7 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Learning diverse stochastic human-action generators by learning smooth latent transitions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a35875b-133f-408f-856e-523ea6a1330b · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionChain: Conversational Motion Controllers via Multimodal Prompts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c0d45bf0-6888-4ddc-82c2-125238f37983 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8293c10-4cb8-43e5-b917-088f5d8a3fb0 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Large motion model for unified multi-modal motion generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1dda8c6c-d1f5-4128-9e8f-1cc6bfa07b58 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Flamingo: a visual language model for few-shot learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341ddcb9-d858-4868-aabf-0b7c45401c19 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b49001d-75f1-408a-997f-aed91bd5b1ed · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20663abe-3cae-4f82-a4be-36d4943a4bcd · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e480fb9c-08b8-4a63-ac25-0bbcf0bddad0 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Lisa: Reasoning segmentation via large language model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa3ce64-3ad0-41e3-b8c0-f154ae49a476 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce96e7aa-98de-4a25-b0fe-a707dd9a63cf · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 085c8fec-7879-4d83-9ba4-3c5a4fa9332b · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c21b9c-1ed6-46bd-a882-91d9cc60f152 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Finite Scalar Quantization: VQ-VAE Made Simple
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19116fe6-42eb-4ba8-9981-4744899a9c44 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37024144-556a-4f56-b275-8c5f52378f5e · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model HumanTOMATO: Text-aligned Whole-body Motion Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed91865-da65-4d6e-a1db-a804263ec9bb · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1bf1a21-a541-42d9-9561-b1a77cfae242 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd5e36c-bab4-44e7-aca4-8ca3ae7de1e4 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07474f84-7df4-418b-832d-84842521fee3 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Sigmoid loss for language image pre-training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0cc928d-8f53-4146-9b96-1e97f4d9a744 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831fad93-aa76-4d93-82c8-7cf4f39ea056 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Smpl: A skinned multi-person linear model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4ae44264-f948-4cca-937d-02e8bf0a804b · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9f424e3-a602-4eb6-bb9f-bff04676e6e3 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model You only look once: Unified, real-time object detection
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16abd518-9938-4017-989e-04af108538c1 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Wham: Reconstructing world-grounded humans with accurate 3d motion
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef57ef7-c289-471c-9781-0d3376cbafff · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Perpetual humanoid control for real-time simulated avatars
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e663f2-0d61-4a5d-acaf-797a31be13c3 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 519ee99e-1e62-4771-9393-ceb389115477 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e52c3ca-4481-48e3-b569-1d8350ff632b · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f031606a-9685-4d35-967b-9da7383f2d0e · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26d89571-6974-453b-96e8-826bd48f83a2 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The kit motion-language dataset.Big data, 4(4):236– 252, 2016
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c88c3d4d-616c-42ec-b458-2c6cfc015bd6 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Amass: Archive of motion capture as surface shapes
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e600edf6-3a56-4126-bdd1-9dcb55995bc4 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recovering accurate 3d human pose in the wild using imus and a moving camera
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 95a54321-6915-4f40-b11e-4f536679466a · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ac2b0286-e91c-4598-8878-12aa803647ed · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b1eb614-d23a-4f42-9002-67d3e549c225 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Executing your commands via motion diffusion in latent space
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2938cc62-ac40-4c9e-b2d1-53d03875db61 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159b2500-41a7-49cc-ac83-ec77fdb8669a · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating human motion from textual descriptions with discrete representations
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 63b8103f-e63b-4970-bbd4-3bcc41f92c0e · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c5c879b-afdd-482a-875b-d6ae4e2c782f · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Avatargpt: All-in-one framework for motion understanding planning generation and beyond
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 63da06a9-a3db-4858-89f5-3acdbde51454 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model llamacpp.https://github.com/ggml-org/llama.cpp, 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 60cd66d7-ba14-42f0-8102-68c5edc2b3df · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Parco: Part- coordinating text-to-motion synthesis
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b5ff2380-7174-4532-9f71-1442afa78375 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f4fe6324-23d1-49cd-8c34-8ab627fa4359 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8dd1cb8-d25d-4708-83cc-06817a31f57f · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Microsoft coco: Common objects in context
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359e48f1-c9e6-47d1-8b3e-316920fdb2b0 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posetrack: A benchmark for human pose estimation and tracking
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4d4680ae-0b05-493f-b552-7d70fa057295 · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Resolving 3d human pose ambiguities with 3d scene constraints
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a18f99fb-b686-4ff0-9033-a97193bf97dd · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Behave: Dataset and method for tracking human object interactions
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a0088ee7-8b45-4f9a-9ff4-ac15d6d8754c · outbound
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model the left hand is positioned below the right hand
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f3f0f2b1-ba20-4233-91d7-cb8d54d2cfd2 · inbound
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76dac94f-3c6c-4fa7-b748-124078c28352 · inbound
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5fa7384a-329a-43d7-9e3c-f6db91017093 · inbound
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.