Pith. sign in

Paper Citation Record · LEDGER

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2508.07863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07863 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:55:03.686435Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:25.303896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:47:57.512626Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy28
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dece0443-f9d0-4aed-8a4b-fb9ee5b6a50c · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Momask: Generative masked modeling of 3d human motions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.061147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.061147Z digest=sha256:ee2803ce6757359118be82a0e371ffcfd7605993172182af3c23606131ee7217

Observation e7f6a5fe-b310-4021-88d5-e9b0c9f65b11 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Human motion as a foreign language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.194279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.194279Z digest=sha256:976dc8243f5dec33429a027ec8cc0e531d5ee511914a33a8dfebefbec7095e24

Observation dff48fa4-f852-4f2c-843c-a6a25f92051e · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating diverse and natural 3d human motions from text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.356942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.356942Z digest=sha256:71af5bc1784200d09fac7fc3bf66ba85c2b1fbe8ebf1af93d23dfe9aaa7d6287

Observation fa0611e9-d169-49ac-8e34-ebb202c2a7a7 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:11.071203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:57.480981Z digest=sha256:3c7212c241bc1cda40a8c3cee3b3ec7907e187a3b0e1608974898da7270e20dd

Observation 8b3db17c-7c61-421d-81f4-80040385add2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.627224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.627224Z digest=sha256:e0ff5bc317ed3d5a6a8f1d43c4e677ebfd4a19c77dd1459f6b7a51cda7242dfe

Observation 1c21bda0-ecc1-410b-b15d-9684fde545d4 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Improved baselines with visual instruction tuning, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.739675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.739675Z digest=sha256:ca6849ed77a6d8786c991a93e7eb7d3d069ed79c6b5d3de10a79791be92ce57f

Observation be59e44c-9ed1-4551-8228-2bee0d92fc49 · outbound

This paper cites A large-scale rgb-d database for arbitrary-view human action recognition.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A large-scale rgb-d database for arbitrary-view human action recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.900308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:57.899801Z digest=sha256:05533d11fcd7b8d0bf96a96280128ebcbe954412d6a7d7354688c221f8b37c73

Observation 5113d07b-6e8d-48e6-b842-1a78fd74b251 · outbound

This paper cites M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.531949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.024126Z digest=sha256:ceccaccf5dde9498be66f76c473048e49496eb7cf0e47ddb6470d749c0227d11

Observation 74225f25-b5f1-48d9-a88a-24eecdcaf43f · outbound

This paper cites Scaling large motion models with million-level human motions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Scaling large motion models with million-level human motions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.589979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.170476Z digest=sha256:23e3c0040b90bbded9aa192a26f6e66cae35b0eede85cb7a7a756ce13d964968

Observation 2fe7aa93-f7bd-4160-a6be-71415b12f3f5 · outbound

This paper cites Autoregressive image generation using residual quantization.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Autoregressive image generation using residual quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:58.286611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:58.286611Z digest=sha256:dc00f02b8253a04e14647b969d7cd6da0db311ce020123f6ebdc1d68c2d9c6fa

Observation 01277ddf-1995-4ea3-b713-52a20ec98685 · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Temos: Generating diverse human motions from textual descriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.171990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.404718Z digest=sha256:f09635492f4478337467cac04e65ade06eef775e553cacad92b79129fd0a5cf8

Observation c7e44568-127d-4782-b635-1021a5a3175c · outbound

This paper cites Language2pose: Natural language grounded pose forecasting.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language2pose: Natural language grounded pose forecasting

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.830066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.488840Z digest=sha256:a1494c918568e9ff7faa5c0f79f4492bf629c7c1b78294318039323c295e53cc

Observation 41af36a7-045a-4bf5-b95d-ddffd17f0d17 · outbound

This paper cites Motiongpt: Finetuned llms are general-purpose motion generators.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Finetuned llms are general-purpose motion generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.518557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.572722Z digest=sha256:eea71e1ca63be7318ce1734d7020f39efde295d34140f5a54f709c599be8a878

Observation 6f82544c-813a-4c9c-b787-9e4c43a67498 · outbound

This paper cites MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:58.674215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:58.674215Z digest=sha256:60ead2fe1a68a540004ec7590f62ca2470a6f7ea178789cbaf97df3bad359391

Observation b5109d80-5552-4ce4-8338-eef799e9e219 · outbound

This paper cites Recurrent network models for human dynamics.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recurrent network models for human dynamics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.209540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.788216Z digest=sha256:c784d21f8ed8f45f35ce0a6fc5566019e97bb7b64bb1f91ba024e888e2152ae0

Observation 34812253-a8e0-43bf-ad8c-09a39ec1b6b2 · outbound

This paper cites A neural temporal model for human motion prediction.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A neural temporal model for human motion prediction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.899001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.901849Z digest=sha256:0b17c899a78997056506cd9165d34b3fae020eb715d0f4af3193fbd67e7e9026

Observation add6134e-aa7e-4f2e-996d-651df9ac534f · outbound

This paper cites A stochastic conditioning scheme for diverse human motion prediction.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A stochastic conditioning scheme for diverse human motion prediction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.524160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:58.988702Z digest=sha256:8a0d7e91cb4c5ca348792256eb6f348fdc708d05969e4bfae71d798c4af4def0

Observation f2a32cee-4d38-4827-9e1b-3a6acd8c1de7 · outbound

This paper cites Learning diverse stochastic human-action generators by learning smooth latent transitions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Learning diverse stochastic human-action generators by learning smooth latent transitions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.190878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:59.102481Z digest=sha256:9a7ca3bd727b4266fb6877a648fa2910fc640f2b4742ef88ca1bdea2071ad451

Observation 3a35875b-133f-408f-856e-523ea6a1330b · outbound

This paper cites MotionChain: Conversational Motion Controllers via Multimodal Prompts.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionChain: Conversational Motion Controllers via Multimodal Prompts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.395699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:59.216321Z digest=sha256:3a02ae327510d19acaca2b123b2405dd6be252a20dae0069cd3af3dd567d511d

Observation c0d45bf0-6888-4ddc-82c2-125238f37983 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.349922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.349922Z digest=sha256:88df051162c63b661e95f7c7a29bbdd4f5abeadf50708ddf25af13c0dac15ad5

Observation e8293c10-4cb8-43e5-b917-088f5d8a3fb0 · outbound

This paper cites Large motion model for unified multi-modal motion generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Large motion model for unified multi-modal motion generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.920463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:54:59.472583Z digest=sha256:00a22d3bdb741c3fe50151c60b71317c10cc0f25234ddf93bbab2f6191f2edff

Observation 1dda8c6c-d1f5-4128-9e8f-1cc6bfa07b58 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Flamingo: a visual language model for few-shot learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.555702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.555702Z digest=sha256:c03d244f75db86e0ffa3d65c598575f0f84390184c629c7149db5c259421459d

Observation 341ddcb9-d858-4868-aabf-0b7c45401c19 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.654282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.654282Z digest=sha256:22e0d646b66b6680894916fe9087127faf7862b76b447689f0fc69cb8058c905

Observation 1b49001d-75f1-408a-997f-aed91bd5b1ed · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.795393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.795393Z digest=sha256:9f7ba40e13d692abf9b130df354d7389bc08a401cc0241393cf2537dd9b23dfb

Observation 20663abe-3cae-4f82-a4be-36d4943a4bcd · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.894391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.894391Z digest=sha256:e949c9bf74ed10fbfa609c1241f34cd27d5b972e95fa71b33c569c0ee84378cc

Observation e480fb9c-08b8-4a63-ac25-0bbcf0bddad0 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.990725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.990725Z digest=sha256:b253a4711972529f09ad1337be2d8b09683c728661fbab1eeb1531dc2ec1b34d

Observation efa3ce64-3ad0-41e3-b8c0-f154ae49a476 · outbound

This paper cites Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.076186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.076186Z digest=sha256:8978eba18cde9fffabf451819083701da5cc9bded9f1d39d88fe91386758d59b

Observation ce96e7aa-98de-4a25-b0fe-a707dd9a63cf · outbound

This paper cites Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.655546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:00.164563Z digest=sha256:63b673dea436859d7763cbfb1e9a9038742a8a1dd202da961d811cfd9afc4fd9

Observation 085c8fec-7879-4d83-9ba4-3c5a4fa9332b · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.299651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.299651Z digest=sha256:29ea0763d9ae601fefc8a2c69af96af758b9c3ec0517b791c3a6fb029b4fb396

Observation a7c21b9c-1ed6-46bd-a882-91d9cc60f152 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Finite Scalar Quantization: VQ-VAE Made Simple

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.438317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.438317Z digest=sha256:3216e65bf109e6e36df551e1e1a30aac0550ded77481df64741d27513a2442d2

Observation 19116fe6-42eb-4ba8-9981-4744899a9c44 · outbound

This paper cites The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.212186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:00.545792Z digest=sha256:b4f109abda04e66f9c356c7ddb5c80f9165445369942c7c85475d1d46acdd099

Observation 37024144-556a-4f56-b275-8c5f52378f5e · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.655357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.655357Z digest=sha256:f2f759829414576db153ce1a6bf76f2ab8afe77bbd342599012cd6f43b5c3275

Observation 3ed91865-da65-4d6e-a1db-a804263ec9bb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.759463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.759463Z digest=sha256:fef1bea90ac40f38957cdbfa0abcb23e211fdbf80bbf71f433cd6de026f14b08

Observation d1bf1a21-a541-42d9-9561-b1a77cfae242 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.851447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.851447Z digest=sha256:3605cae37a54c88a23bf013697d2c76b471bcc0a3c0748899c3845c4a9d18d76

Observation 0dd5e36c-bab4-44e7-aca4-8ca3ae7de1e4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.957517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.957517Z digest=sha256:76e71e807fc48735e00ba7b1972458e9ce05680aabed2f4a580693cc4442704e

Observation 07474f84-7df4-418b-832d-84842521fee3 · outbound

This paper cites Sigmoid loss for language image pre-training.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Sigmoid loss for language image pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.067511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.067511Z digest=sha256:8e65459cec8a95ec4557ed0d578a4cb8e6b4235103640a2154300c95c38866d4

Observation d0cc928d-8f53-4146-9b96-1e97f4d9a744 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.153815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.153815Z digest=sha256:f4a65bab0970e89b07acecc89af5c44f1a9ef1ad35a8d260c23db879c1355c84

Observation 831fad93-aa76-4d93-82c8-7cf4f39ea056 · outbound

This paper cites Smpl: A skinned multi-person linear model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Smpl: A skinned multi-person linear model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.365209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:01.246734Z digest=sha256:40684b5c449969d68c9970288e9281de48eb4616545357b06499600a05f99853

Observation 4ae44264-f948-4cca-937d-02e8bf0a804b · outbound

This paper cites ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.023236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:01.311777Z digest=sha256:102f7d63f491f9be73ccc74c34a5ad195f4ee31542cff3587697d72b49f88ba3

Observation d9f424e3-a602-4eb6-bb9f-bff04676e6e3 · outbound

This paper cites You only look once: Unified, real-time object detection.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model You only look once: Unified, real-time object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.439383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.439383Z digest=sha256:f87b14d95f3bddb1611ff196fe056cf62206f8a659cd2361c121410f699908ec

Observation 16abd518-9938-4017-989e-04af108538c1 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.545093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.545093Z digest=sha256:d928f64e4cf09729a1bfe221752f188e86ea3b6a31f9f0b06b12abc7535d6c7d

Observation 1ef57ef7-c289-471c-9781-0d3376cbafff · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Perpetual humanoid control for real-time simulated avatars

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.629310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.629310Z digest=sha256:19ca29b6615b9893b99eb57ecc4a1d08e51f20139011930c0b554ad794a4c46c

Observation a4e663f2-0d61-4a5d-acaf-797a31be13c3 · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.735535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.735535Z digest=sha256:706ed66978421cfa37127ea1fae841f3efb456874ae7394237c558925fa14d22

Observation 519ee99e-1e62-4771-9393-ceb389115477 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.847632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.847632Z digest=sha256:c8343d4ab03d6a261edc92b9b5be2d3048df094bb0f2a47a61c338f0e3f2bc69

Observation 8e52c3ca-4481-48e3-b569-1d8350ff632b · outbound

This paper cites Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.069823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:01.922185Z digest=sha256:086572ff5525d63d2b2c58105686417abca93b0646dacab54402cd01443a56e4

Observation f031606a-9685-4d35-967b-9da7383f2d0e · outbound

This paper cites UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:03.873453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.001344Z digest=sha256:e479410f6e7998f02a61b4a603b3214eb3ba47c67f313aacc0f4969e6fcf4c61

Observation 26d89571-6974-453b-96e8-826bd48f83a2 · outbound

This paper cites The kit motion-language dataset.Big data, 4(4):236– 252, 2016.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The kit motion-language dataset.Big data, 4(4):236– 252, 2016

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.837633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.083256Z digest=sha256:737d37951b9865f7691ea41ba91177c3647c4a9eaa375b762a22cc2b431e2007

Observation c88c3d4d-616c-42ec-b458-2c6cfc015bd6 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Amass: Archive of motion capture as surface shapes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.186443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.186443Z digest=sha256:dafd0f819d0c8be1c63fe8571b6674ad1bf6be6768caceaec4fdb4944dd57b2a

Observation e600edf6-3a56-4126-bdd1-9dcb55995bc4 · outbound

This paper cites Recovering accurate 3d human pose in the wild using imus and a moving camera.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recovering accurate 3d human pose in the wild using imus and a moving camera

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.713662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.270194Z digest=sha256:249e9df53a70f07c9cac0f134227adb46dbebff12139ae0571d73584fc4f3d71

Observation 95a54321-6915-4f40-b11e-4f536679466a · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.590197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.345136Z digest=sha256:73696c48fdfdc9ffde34c08c487e0baeb4fcec07c38d949e4f25c595361609d8

Observation ac2b0286-e91c-4598-8878-12aa803647ed · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.419755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.419755Z digest=sha256:558fd5798d87d586ca46ad35f372b034555b2e718241566c094b997ce3a66a18

Observation 4b1eb614-d23a-4f42-9002-67d3e549c225 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Executing your commands via motion diffusion in latent space

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.463606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.501633Z digest=sha256:c2586bbac583bbebe32af680a478fe0ccb351c04b99aff59d267b82da5cf136b

Observation 2938cc62-ac40-4c9e-b2d1-53d03875db61 · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.580215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.580215Z digest=sha256:04673625744bfbcf466e944d2b6e0aa6dd7cdebc4c82cd3c533b28e271d9891b

Observation 159b2500-41a7-49cc-ac83-ec77fdb8669a · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating human motion from textual descriptions with discrete representations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.313231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.747962Z digest=sha256:5c42b2813fa42c86b138c1ad034680a4af43d9df08aae6f2e70aee8664bbee2d

Observation 63b8103f-e63b-4970-bbd4-3bcc41f92c0e · outbound

This paper cites Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.826806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.826806Z digest=sha256:e706c59ba10c27814ceb1595408fff4f17872c43f9152d68d8e94ed37d15fafc

Observation 5c5c879b-afdd-482a-875b-d6ae4e2c782f · outbound

This paper cites Avatargpt: All-in-one framework for motion understanding planning generation and beyond.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Avatargpt: All-in-one framework for motion understanding planning generation and beyond

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.135640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.910283Z digest=sha256:7270408577ed6fa40203ed5992e87339887284bbd1794d27c6ab6545f6524602

Observation 63da06a9-a3db-4858-89f5-3acdbde51454 · outbound

This paper cites llamacpp.https://github.com/ggml-org/llama.cpp, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model llamacpp.https://github.com/ggml-org/llama.cpp, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.927027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:02.982817Z digest=sha256:9c0a871374d6aba681dfa27a1f689b97b29f9c231453eec12cf69e9fd08fea40

Observation 60cd66d7-ba14-42f0-8102-68c5edc2b3df · outbound

This paper cites Parco: Part- coordinating text-to-motion synthesis.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Parco: Part- coordinating text-to-motion synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.745015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.075835Z digest=sha256:1ba2e07f17e65b52bd31cddb365ccd517db7e2437830ae5c2ceaedcf5e01209d

Observation b5ff2380-7174-4532-9f71-1442afa78375 · outbound

This paper cites Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.563158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.191406Z digest=sha256:073e80bf8d72160a55c56d8b5d471f58cdb2db14cd054b49777a224a081fcdf6

Observation f4fe6324-23d1-49cd-8c34-8ab627fa4359 · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.397712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.264774Z digest=sha256:cc8e74a15fe6d70787ef883b577a6d0d6b4fe8ebcc4c814969b4dac60525712c

Observation e8dd1cb8-d25d-4708-83cc-06817a31f57f · outbound

This paper cites Microsoft coco: Common objects in context.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Microsoft coco: Common objects in context

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:03.339648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:03.339648Z digest=sha256:54a30f33b4257cbc5e09faf8773c4d0710c4298e976cf05c9ba989e7e972e90b

Observation 359e48f1-c9e6-47d1-8b3e-316920fdb2b0 · outbound

This paper cites Posetrack: A benchmark for human pose estimation and tracking.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posetrack: A benchmark for human pose estimation and tracking

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.229691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.411500Z digest=sha256:8009c13cbdfb2029ec4d86fd53b281c443bf172c076cea37255e8efabe1c1d48

Observation 4d4680ae-0b05-493f-b552-7d70fa057295 · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Resolving 3d human pose ambiguities with 3d scene constraints

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.051157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.476511Z digest=sha256:02453f6bfae03fb54dd8dd10c9ffdad9c4abf619262ec1f46a068a7b79ff2215

Observation a18f99fb-b686-4ff0-9033-a97193bf97dd · outbound

This paper cites Behave: Dataset and method for tracking human object interactions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Behave: Dataset and method for tracking human object interactions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:04.881522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.555056Z digest=sha256:7e7533175d3006b52e7e2acf42140549815bb47befc0f122e5dac58ee049f1fa

Observation a0088ee7-8b45-4f9a-9ff4-ac15d6d8754c · outbound

This paper cites the left hand is positioned below the right hand.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model the left hand is positioned below the right hand

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:04.677201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:03.686435Z digest=sha256:c463b8fc9d60ec8f7d80d855fc06ee2274aa402caced6fd8ac380bf4336d29f9

Pith citing papers

Observation f3f0f2b1-ba20-4233-91d7-cb8d54d2cfd2 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:54.531502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:54.531502Z digest=sha256:52c58970caa3afb3869311c353fc5ff3184cf7524901446b4b045c3bca0566f9

Observation 76dac94f-3c6c-4fa7-b748-124078c28352 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.514056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:d11bf8e3c3600947c228c05d17ed3417435fa0f958dc619f0c3b72a920c2e140

Observation 5fa7384a-329a-43d7-9e3c-f6db91017093 · inbound

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cites this paper.

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:25.303896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:25.303896Z digest=sha256:66bbfa945982dd531d1bb0e503ea90833718710a65a1201b3b87b26418d2fec5