Pith. sign in

Paper Citation Record · LEDGER

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

As of 14 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2508.07863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07863 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:55:03.686435Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:25.303896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:47:57.512626Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy28
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dece0443-f9d0-4aed-8a4b-fb9ee5b6a50c · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Momask: Generative masked modeling of 3d human motions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.061147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.061147Z digest=sha256:be8767ec8cb9fe58e6d1ea3abf1330d1f97afb81997471291d88a7da43475526

Observation e7f6a5fe-b310-4021-88d5-e9b0c9f65b11 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Human motion as a foreign language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.194279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.194279Z digest=sha256:c8bbc3992a93407d8a02015594c8ebf8386cf65895735f146e34ad597342d8ca

Observation dff48fa4-f852-4f2c-843c-a6a25f92051e · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating diverse and natural 3d human motions from text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.356942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.356942Z digest=sha256:6440256f6be43e2f34edaf1a7a83390099c87b1fee340c64fca34902ae511ee5

Observation fa0611e9-d169-49ac-8e34-ebb202c2a7a7 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:11.071203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:57.480981Z digest=sha256:3ada19220c57f364374cc6796916e098eec5f8849f99a30fab60bb57fd7c1278

Observation 8b3db17c-7c61-421d-81f4-80040385add2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.627224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.627224Z digest=sha256:21db331bf46fc8142b98ef662d30d3eafd0a560be148d490fae113a57375d9c0

Observation 1c21bda0-ecc1-410b-b15d-9684fde545d4 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Improved baselines with visual instruction tuning, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.739675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.739675Z digest=sha256:8d8804ef2b8b982c3b2397a721491587e5bf0dc79569c5ba81f0e063de8aafd4

Observation be59e44c-9ed1-4551-8228-2bee0d92fc49 · outbound

This paper cites A large-scale rgb-d database for arbitrary-view human action recognition.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A large-scale rgb-d database for arbitrary-view human action recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.900308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:57.899801Z digest=sha256:a703f4c069af7e6b7458708e1888f38a5fd6546947caabd396146d9e8cd21e65

Observation 5113d07b-6e8d-48e6-b842-1a78fd74b251 · outbound

This paper cites M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.531949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.024126Z digest=sha256:cca0aa21a0e97b2395141419b9b4132e6015c1d0aaab425ff07410d308ea67e1

Observation 74225f25-b5f1-48d9-a88a-24eecdcaf43f · outbound

This paper cites Scaling large motion models with million-level human motions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Scaling large motion models with million-level human motions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.589979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.170476Z digest=sha256:4192e0c89c4866718b7b4df898f0ad2972b4254adff1743e464db671a5d68d00

Observation 2fe7aa93-f7bd-4160-a6be-71415b12f3f5 · outbound

This paper cites Autoregressive image generation using residual quantization.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Autoregressive image generation using residual quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:58.286611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:58.286611Z digest=sha256:62c66d615a93c36bf0aa4642c52d8448a4df67d098da6da8a0ce0d9748015bd7

Observation 01277ddf-1995-4ea3-b713-52a20ec98685 · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Temos: Generating diverse human motions from textual descriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.171990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.404718Z digest=sha256:d1c0b6f69c5ced3c3e8d650bdac604eb81ca2d3b600f48561889bdddd2623cf4

Observation c7e44568-127d-4782-b635-1021a5a3175c · outbound

This paper cites Language2pose: Natural language grounded pose forecasting.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language2pose: Natural language grounded pose forecasting

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.830066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.488840Z digest=sha256:2b6070cf7efd0396fe8b91eb6284cfd594cf5c3e136cacdf4f7bffb876c47ab6

Observation 41af36a7-045a-4bf5-b95d-ddffd17f0d17 · outbound

This paper cites Motiongpt: Finetuned llms are general-purpose motion generators.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Finetuned llms are general-purpose motion generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.518557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.572722Z digest=sha256:875208ef4990f1f21c41c45ac9568006c85d41f5ced4188c367fd382c31ea1c9

Observation 6f82544c-813a-4c9c-b787-9e4c43a67498 · outbound

This paper cites MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:58.674215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:58.674215Z digest=sha256:83ebdca2e45c4250e4a6ff9cdb5960e82958d39d1affa3b0e4a5b8fff482da14

Observation b5109d80-5552-4ce4-8338-eef799e9e219 · outbound

This paper cites Recurrent network models for human dynamics.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recurrent network models for human dynamics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.209540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.788216Z digest=sha256:760eee2ca13ec948fee3c95c00433b535c84a26b14d1db7413dff877d36fd34f

Observation 34812253-a8e0-43bf-ad8c-09a39ec1b6b2 · outbound

This paper cites A neural temporal model for human motion prediction.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A neural temporal model for human motion prediction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.899001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.901849Z digest=sha256:74b5f8cd200800af9803f0dae58de4370e231c2dc145ac824cfa85fbfe72ee84

Observation add6134e-aa7e-4f2e-996d-651df9ac534f · outbound

This paper cites A stochastic conditioning scheme for diverse human motion prediction.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A stochastic conditioning scheme for diverse human motion prediction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.524160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:58.988702Z digest=sha256:d95a15efd866bcfa29268aa0d34e5157e6790db4428a0066e40fa481203f1283

Observation f2a32cee-4d38-4827-9e1b-3a6acd8c1de7 · outbound

This paper cites Learning diverse stochastic human-action generators by learning smooth latent transitions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Learning diverse stochastic human-action generators by learning smooth latent transitions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.190878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:59.102481Z digest=sha256:014c01bb6280f32a70198e203896360c0f22b21c7bd4139d7352643248d2bc22

Observation 3a35875b-133f-408f-856e-523ea6a1330b · outbound

This paper cites MotionChain: Conversational Motion Controllers via Multimodal Prompts.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionChain: Conversational Motion Controllers via Multimodal Prompts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.395699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:59.216321Z digest=sha256:5b5d1d5094e218d6de7d96788e031b67850f5c8c9404bd2074fe68cd3362fc42

Observation c0d45bf0-6888-4ddc-82c2-125238f37983 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.349922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.349922Z digest=sha256:487a0225c2aa375539a2b4a22d833afcc690d204806332c5f98ae7ffe1755550

Observation e8293c10-4cb8-43e5-b917-088f5d8a3fb0 · outbound

This paper cites Large motion model for unified multi-modal motion generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Large motion model for unified multi-modal motion generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.920463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:54:59.472583Z digest=sha256:b836e2ba73e60124c481e817e9fdbf9370f851735b6de4da2db5e2f7907121cc

Observation 1dda8c6c-d1f5-4128-9e8f-1cc6bfa07b58 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Flamingo: a visual language model for few-shot learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.555702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.555702Z digest=sha256:225302b500bf8283ed44c706fdef237e994877450b6b333f23e8e9313d79ad60

Observation 341ddcb9-d858-4868-aabf-0b7c45401c19 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.654282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.654282Z digest=sha256:067cac0bfa7d1b590654ae0d988b580b3ac56b37a59db0bc6999f932ff1c348f

Observation 1b49001d-75f1-408a-997f-aed91bd5b1ed · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.795393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.795393Z digest=sha256:b42b9af7c74ff7ec83e29bbd89c4d486d64ac671fadb4c555f7ecc1fb513f160

Observation 20663abe-3cae-4f82-a4be-36d4943a4bcd · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.894391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.894391Z digest=sha256:75a51d61ae360db58505379fae2e6a4156c662934234bfe11b3cc9767c22fa10

Observation e480fb9c-08b8-4a63-ac25-0bbcf0bddad0 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.990725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.990725Z digest=sha256:9560b2023954b8fe8be3e511a13a708ffb77cbad5b856224d0e783bc0ce2f334

Observation efa3ce64-3ad0-41e3-b8c0-f154ae49a476 · outbound

This paper cites Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.076186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.076186Z digest=sha256:14776ee786be204ad35ea3646c2ccab061cdc417e57f58d36296e709b4623d28

Observation ce96e7aa-98de-4a25-b0fe-a707dd9a63cf · outbound

This paper cites Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.655546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:00.164563Z digest=sha256:87c8e6409ddead540d6ac758b98b527846190eb6b495f0026b7211677e4ef3b0

Observation 085c8fec-7879-4d83-9ba4-3c5a4fa9332b · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.299651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.299651Z digest=sha256:6502781fe278cdff6fbb43e4ca705d89eeba324fb9a6fbb3c074594393a3a470

Observation a7c21b9c-1ed6-46bd-a882-91d9cc60f152 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Finite Scalar Quantization: VQ-VAE Made Simple

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.438317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.438317Z digest=sha256:69900c143651ccea6157ad1f5ca389558491ea7d9ad10e401c095f8d601eb63c

Observation 19116fe6-42eb-4ba8-9981-4744899a9c44 · outbound

This paper cites The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.212186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:00.545792Z digest=sha256:bbde6505fa9457bc892e93bc4576f4043709a31617ccdb178da1a87306f54014

Observation 37024144-556a-4f56-b275-8c5f52378f5e · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.655357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.655357Z digest=sha256:ae039266670221a5d6432509dcf3cbe1ac68d4eb44cfdeeab7db6545bb779efa

Observation 3ed91865-da65-4d6e-a1db-a804263ec9bb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.759463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.759463Z digest=sha256:2d2b3dfdcbddd1c398fa7f0a5b8fc0d23cf67fbd284287f9120db52c004c7d42

Observation d1bf1a21-a541-42d9-9561-b1a77cfae242 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.851447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.851447Z digest=sha256:76d08f27719975d0b8e613020f0f9f0de4bd2bc9dea4feb3effc6306fc6d13fb

Observation 0dd5e36c-bab4-44e7-aca4-8ca3ae7de1e4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.957517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.957517Z digest=sha256:f5a7b9511d97b2ae7a9b9a992daf0508ab185b1db24cbfda00126f078546cf49

Observation 07474f84-7df4-418b-832d-84842521fee3 · outbound

This paper cites Sigmoid loss for language image pre-training.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Sigmoid loss for language image pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.067511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.067511Z digest=sha256:12e277b7b59208dafbee1a56b367e6644bb12a5fd76cbbbff3d427c95e61750a

Observation d0cc928d-8f53-4146-9b96-1e97f4d9a744 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.153815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.153815Z digest=sha256:965156665f80b8d1f6af9711adb8a1261fc9ae800e33965d3aba1ad4929b9e79

Observation 831fad93-aa76-4d93-82c8-7cf4f39ea056 · outbound

This paper cites Smpl: A skinned multi-person linear model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Smpl: A skinned multi-person linear model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.365209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:01.246734Z digest=sha256:edae5583b850981e66981bcfe595a66126d7a1e698dbcee01a9fc6fd272b9836

Observation 4ae44264-f948-4cca-937d-02e8bf0a804b · outbound

This paper cites ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.023236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:01.311777Z digest=sha256:d5460113b8d0769db8058fd9d9ec65c98b4dc41973df06fc966d395d934d0646

Observation d9f424e3-a602-4eb6-bb9f-bff04676e6e3 · outbound

This paper cites You only look once: Unified, real-time object detection.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model You only look once: Unified, real-time object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.439383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.439383Z digest=sha256:3cdac57b80273771da0b16d8b53b08f9d03c837748922bbc868990b6259a93cb

Observation 16abd518-9938-4017-989e-04af108538c1 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.545093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.545093Z digest=sha256:1df79b1ae94ead8b45cd4aca530b9a9df7d17ba87cb65400b822e2a29d2c4881

Observation 1ef57ef7-c289-471c-9781-0d3376cbafff · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Perpetual humanoid control for real-time simulated avatars

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.629310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.629310Z digest=sha256:51320b36a841eb6bd9302e455813b338f1a01c7d9b1d2074956400a9a67785f6

Observation a4e663f2-0d61-4a5d-acaf-797a31be13c3 · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.735535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.735535Z digest=sha256:1f376494e5480291fb49caf3b7506e54b7bdb2044e91a0f065184b4fff4b9db5

Observation 519ee99e-1e62-4771-9393-ceb389115477 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.847632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.847632Z digest=sha256:d9cc0ec0a3f2bf0c489727864ccb1d6879c7ae85953712ae693030c737307404

Observation 8e52c3ca-4481-48e3-b569-1d8350ff632b · outbound

This paper cites Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.069823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:01.922185Z digest=sha256:18a10d3f94d4bb0bae8269c956436a9c1085d009bcabf6c469c0ea29799c1414

Observation f031606a-9685-4d35-967b-9da7383f2d0e · outbound

This paper cites UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:03.873453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.001344Z digest=sha256:d8e30dc510b36154daa7a7b813a0c3b0885c7d2b0fea6d383dd3144b4d906002

Observation 26d89571-6974-453b-96e8-826bd48f83a2 · outbound

This paper cites The kit motion-language dataset.Big data, 4(4):236– 252, 2016.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The kit motion-language dataset.Big data, 4(4):236– 252, 2016

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.837633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.083256Z digest=sha256:3cf4cc9592907cb0465bd142884a6af12e0ac89127c69987f0d743e9b86fd490

Observation c88c3d4d-616c-42ec-b458-2c6cfc015bd6 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Amass: Archive of motion capture as surface shapes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.186443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.186443Z digest=sha256:425faaac739241fc3aab7f0d1f16d4ad395b97d8543adb6895bd5b9c94df370a

Observation e600edf6-3a56-4126-bdd1-9dcb55995bc4 · outbound

This paper cites Recovering accurate 3d human pose in the wild using imus and a moving camera.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recovering accurate 3d human pose in the wild using imus and a moving camera

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.713662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.270194Z digest=sha256:a9f52edfe49279545be0c915e2126a0f50bd66c383c839fc59b3771fe5d69f25

Observation 95a54321-6915-4f40-b11e-4f536679466a · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.590197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.345136Z digest=sha256:d2463a80a29a04dc66831675dadfcdba822e27057e034520f871bd1b807a29b7

Observation ac2b0286-e91c-4598-8878-12aa803647ed · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.419755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.419755Z digest=sha256:3a1eaf6447bba5e940370d5901d642aba3a7195ff547e7e624b80d069da19c01

Observation 4b1eb614-d23a-4f42-9002-67d3e549c225 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Executing your commands via motion diffusion in latent space

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.463606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.501633Z digest=sha256:6d6a353091ede64dff7b38b3c044cf2c92d8dfebb2a2c697bee0dc76e04458c5

Observation 2938cc62-ac40-4c9e-b2d1-53d03875db61 · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.580215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.580215Z digest=sha256:2fbb9431f418966cc21682b1f7ffb673ebadbbf5b6861e00bfa6d3076deb86c4

Observation 159b2500-41a7-49cc-ac83-ec77fdb8669a · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating human motion from textual descriptions with discrete representations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.313231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.747962Z digest=sha256:77ca1226233fd9e8f882d7be04452ce92159d8a2b3c85b9b4d5f6e98ba77ac6c

Observation 63b8103f-e63b-4970-bbd4-3bcc41f92c0e · outbound

This paper cites Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.826806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.826806Z digest=sha256:444c31b9069110ea97bc47a9c113c9f6f197ac7f713b2ff1a59f011236216ae3

Observation 5c5c879b-afdd-482a-875b-d6ae4e2c782f · outbound

This paper cites Avatargpt: All-in-one framework for motion understanding planning generation and beyond.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Avatargpt: All-in-one framework for motion understanding planning generation and beyond

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.135640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.910283Z digest=sha256:18d4289f3d0e09fb0cecea0fb722dacdfdf859f0e3f94e4d42244881921a8144

Observation 63da06a9-a3db-4858-89f5-3acdbde51454 · outbound

This paper cites llamacpp.https://github.com/ggml-org/llama.cpp, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model llamacpp.https://github.com/ggml-org/llama.cpp, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.927027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:02.982817Z digest=sha256:01f7200fbf9776a27e9d7d08c1a5e506b01b41d6695ef0248c3cfdfdf552a4ae

Observation 60cd66d7-ba14-42f0-8102-68c5edc2b3df · outbound

This paper cites Parco: Part- coordinating text-to-motion synthesis.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Parco: Part- coordinating text-to-motion synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.745015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.075835Z digest=sha256:93a29c750636d8947e76a8233498a2c19fdfa69868d02f2583814aebfdfbeaf7

Observation b5ff2380-7174-4532-9f71-1442afa78375 · outbound

This paper cites Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.563158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.191406Z digest=sha256:e8cfeb5ba84c390647af864b3ab4cc2344a74b5b6ee74fc5501564dcc2e0e33a

Observation f4fe6324-23d1-49cd-8c34-8ab627fa4359 · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.397712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.264774Z digest=sha256:7af408228e61cec0f7623cfb0da75f5c87d7aa3f1a9fb258047b051d58b7dd71

Observation e8dd1cb8-d25d-4708-83cc-06817a31f57f · outbound

This paper cites Microsoft coco: Common objects in context.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Microsoft coco: Common objects in context

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:03.339648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:03.339648Z digest=sha256:de329e77cffce7256d47fefcec9653d955922ed332f9259261c9ed2772924456

Observation 359e48f1-c9e6-47d1-8b3e-316920fdb2b0 · outbound

This paper cites Posetrack: A benchmark for human pose estimation and tracking.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posetrack: A benchmark for human pose estimation and tracking

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.229691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.411500Z digest=sha256:83f018475102fd727748ebd6672060228ff21fda83639d897c91f3add4790b96

Observation 4d4680ae-0b05-493f-b552-7d70fa057295 · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Resolving 3d human pose ambiguities with 3d scene constraints

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.051157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.476511Z digest=sha256:7b86b934256c7b9695f3ca44b7570c13a0a29de1a56274db21f322e2e9e98218

Observation a18f99fb-b686-4ff0-9033-a97193bf97dd · outbound

This paper cites Behave: Dataset and method for tracking human object interactions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Behave: Dataset and method for tracking human object interactions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:04.881522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.555056Z digest=sha256:bac20aecab9dfa9f09eae15bf4a90c33ec4e62689e64f72f9eb9a8ffa8960a30

Observation a0088ee7-8b45-4f9a-9ff4-ac15d6d8754c · outbound

This paper cites the left hand is positioned below the right hand.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model the left hand is positioned below the right hand

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:04.677201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:55:03.686435Z digest=sha256:d3c3da7a17f0a5cbf8b8f7284f04f80b70db87888044099812023dccaae1465f

Pith citing papers

Observation f3f0f2b1-ba20-4233-91d7-cb8d54d2cfd2 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:54.531502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:54.531502Z digest=sha256:a0a78197271a8ef2c684dbdca9db1efd0d1dd42f018c01200435f05355fc4ed2

Observation 76dac94f-3c6c-4fa7-b748-124078c28352 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.514056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:1bfde6ad5af86cbfd020f71fad27c4a48f699dfebda895eb928f006f2894f92b

Observation 5fa7384a-329a-43d7-9e3c-f6db91017093 · inbound

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cites this paper.

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:25.303896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:25.303896Z digest=sha256:2241b6bdacc5660d572c98c77e974b9cf4aebe824145c64ed7d54b82579e2b61