Pith. sign in

Paper Citation Record · LEDGER

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 25 inbound Pith citation observations for arXiv:2509.08519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08519 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:35:30.373373Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:04:14.380575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.676485Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f6e52b1-42c4-4c68-8d10-025ce616f829 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Qwen2.5-vl technical report, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.233495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.233495Z digest=sha256:e1b31ad9d3467def5898442a5f8f7156f985d6c895ff157a12142f3737340c43

Observation 6cff7969-361a-4ced-bf5a-e069987fa478 · outbound

This paper cites Goku: Flow Based Video Generative Foundation Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Goku: Flow Based Video Generative Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.238979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.238979Z digest=sha256:77862ecc7cf93ac15c06ba3520d333aca505d24e43e8414a7aef28f24af6b99c

Observation c7a744eb-4b69-4c52-a723-e40767a1831a · outbound

This paper cites Phantom-data : Towards a general subject-consistent video generation dataset, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Phantom-data : Towards a general subject-consistent video generation dataset, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.244265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.244265Z digest=sha256:4e63d23029affdd0dd488ad809bf0c89e2b664f73f24fce2e1a4d65d5499a128

Observation 01de29e7-9c07-4e5c-9b3b-f2d0a4f64690 · outbound

This paper cites Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.248586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.248586Z digest=sha256:1673eb04bd2bd45d3e46ee51300262950db77ac788b84901d0c2868627037989

Observation 73526c4d-fef5-4337-a088-f151684a4563 · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.252509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.252509Z digest=sha256:bce050e9e2a4076fb6106537d475f7bcb1ab629755f046f88abf5297521bd1e6

Observation 1982d1ed-fb0e-4aa2-a0c1-7a6d565eae51 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–1, 2021.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Arcface: Additive angular margin loss for deep face recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–1, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.256166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.256166Z digest=sha256:a8161fb1ad70adadcef9094747b38ba5ec1a0f22c5ff31d023bc22346a42ee82

Observation e04f3538-3bc5-484b-ac68-160620b20b12 · outbound

This paper cites Magref: Masked guidance for any-reference video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Magref: Masked guidance for any-reference video generation, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.259941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.259941Z digest=sha256:18fbc4b4f548649009c9f4313c342e1118f573f4a386dcd4cef75838cff5f28e

Observation 44100d2c-905c-4bed-8b66-757f6c16f9e7 · outbound

This paper cites Seedream 3.0 technical report, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seedream 3.0 technical report, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.264055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.264055Z digest=sha256:e3174a0ebc9e0ea35674c3c57beac2f34106bd4391c9447e3db3a2eb7514305a

Observation 74c364d4-1825-4b38-87bc-97a4f1befbf4 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.268006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.268006Z digest=sha256:2d62535d80ee3c08218cee478425e088c9e554d82467aecce38d776f6cb5465a

Observation 5c884a46-d3b8-45aa-acca-7cf61f7e5515 · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.271855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.271855Z digest=sha256:ad21046a34db394864fb48666d111e2c5d14e957d54f4ad21d37cb446d60a316

Observation 3b99b558-d953-4af1-bf04-df27b6da5e97 · outbound

This paper cites Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.275819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.275819Z digest=sha256:49adad3e402a76e23552172f8eed00dbad085d22d70c9364b3305827f8b600d5

Observation 863638f1-8ad9-4be2-a9a8-6b58c0f34f5c · outbound

This paper cites Curricularface: Adaptive curriculum learning loss for deep face recognition.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Curricularface: Adaptive curriculum learning loss for deep face recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.279526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.279526Z digest=sha256:013e2f3f74b7f72f139c39207b610cee11ed7357620d4de00def45dc5192d149

Observation c5c714b8-89c6-4e78-8a3d-a5fdac551aec · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Vbench: Comprehensive benchmark suite for video generative models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.283029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.283029Z digest=sha256:035a8f7f2485d7735d84c705cec0d595a299c6fcb44ac508f16fec9944151f06

Observation fdbdf8e8-125d-48c5-959f-175d01f7c997 · outbound

This paper cites Multi-reference images to video generation feature.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Multi-reference images to video generation feature

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.286812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.286812Z digest=sha256:6c4c0beacddffd09c7f4760d468920dee551089c22d1f8af245e4814a66c1ffb

Observation e02591c5-48db-4050-9bf1-7b32005ecb64 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.290001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.290001Z digest=sha256:aea948626925595140c4a6bd300da08578f28bb961a4f9d3350d984d2f044604

Observation e85d97a2-7601-46d2-a88a-5a5e715ee43d · outbound

This paper cites Let them talk: Audio-driven multi-person conversational video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Let them talk: Audio-driven multi-person conversational video generation, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.293710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.293710Z digest=sha256:b6260996b52b42072cbeaf4f07e9bc208ae36704050012c0668952ad7cc3bb29

Observation e8056d4d-a20d-4e3a-82c0-26d76dd245ec · outbound

This paper cites Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.297026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.297026Z digest=sha256:8823b7f52c0eec1cab1768c44ce1d6fc73278365ba484e473184c11cb2346f5c

Observation 0fe98423-05f1-4ca0-ad34-4ad69779c0d0 · outbound

This paper cites Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.304093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.304093Z digest=sha256:97a123ea653e15d1811903f1ec533b60d5023662798fab5b957f45897726be17

Observation 6d9df52a-35b1-4f13-a357-79421bbbcf0b · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.307584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.307584Z digest=sha256:4abb4c75957d3db35dcb3327f7c9fbbc754be81fa9cea8c52ac9dbd174ed01a9

Observation 54b180c7-8a38-4f89-8ad7-61606f1148ea · outbound

This paper cites an unresolved cited work.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.311182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.311182Z digest=sha256:7809947bb986f01734d73857347d3486d2b2d4b0450bb2b8d8401a4b181e25aa

Observation 19d52d6e-60b5-4514-95eb-4cfa52ede0b2 · outbound

This paper cites Improving video generation with human feedback, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Improving video generation with human feedback, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.314780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.314780Z digest=sha256:e2ba5e16acc087356a47f5f58a281271c89beab264097a97c51a927aa6f98e9d

Observation 89dac79d-2b20-4695-8017-7d86e6f3cecd · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Phantom: Subject-consistent video generation via cross-modal alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.318096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.318096Z digest=sha256:5b8cb7ea49b8586e2afee28cef6d4b8c0186a368bee11f139b80e7e83b5f57e8

Observation 8edc8a2e-5f5e-405f-83ba-f097b569c2e0 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.321358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.321358Z digest=sha256:24f239c1762d05af5d4343897aa9ddece001fb0158aff68a8976e9fce790735f

Observation dcca47f5-00a3-4da8-9eea-bfe8a32261de · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Movie Gen: A Cast of Media Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.324506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.324506Z digest=sha256:1cb37da4f4c2baceef7b8d2ebf5398a0dce0f053a2b3e12d032da3104621f06c

Observation dd519e46-b700-4561-93c2-7c6a130b742e · outbound

This paper cites Learning transferable visual models from natural language supervision.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.328222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.328222Z digest=sha256:0681045c371a6f46d547e168270c957ae9631b0e1c96d5a27e66e34376e18f6f

Observation e35adaba-2d5d-45de-a961-6add66bae3d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Robust speech recognition via large-scale weak supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.331572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.331572Z digest=sha256:b0d356b4c1fd2f0a3c0665af558cff1a1e12832edd413966e0a606e46e9b39c9

Observation 1342596d-7656-46e5-9770-fa20a0badf5b · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.335032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.335032Z digest=sha256:4e4ef09d0ce30f7528e7a75c62eeb240e0d5530b453dd709aeb476290cd608dc

Observation 404e1143-193f-435e-a946-c7821ae3966b · outbound

This paper cites Seitz, and Ira Kemelmacher-Shlizerman.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seitz, and Ira Kemelmacher-Shlizerman

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.338475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.338475Z digest=sha256:7499a42b5a9903b0d4da821b47b53a261b323d6bb8d6438f451fbe9276049df6

Observation 66a12392-8187-48c7-99b0-49621ece1022 · outbound

This paper cites Gemini 2.5 flash image.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Gemini 2.5 flash image

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.341527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.341527Z digest=sha256:6c2e3e216ddeb4a4481dde422b41b849f75a9dea3aa92a7a5992bd6789baeb37

Observation 5015b6f4-0fdf-4cc8-99f7-331a36e5d40e · outbound

This paper cites Gemini 2.5: Our most intelligent ai model.https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Gemini 2.5: Our most intelligent ai model.https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.344720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.344720Z digest=sha256:090c7b5d2d5643e1a3fd8557cf5d82ef96122a45bd2d82b3772cf7f13c3ead6c

Observation 6761de8e-cee4-4ef1-8e5c-626d1aece653 · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Wan: Open and advanced large-scale video generative models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.348072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.348072Z digest=sha256:00cb77d910b843719ac01c3c6d0c5e96ddecf32fe7a5e8a096b7cd67935ffdf9

Observation 535abdab-5554-4da2-b606-24ca809d2cbc · outbound

This paper cites Fantasytalking: Realistic talking portrait generation via coherent motion synthesis.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Fantasytalking: Realistic talking portrait generation via coherent motion synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.352280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.352280Z digest=sha256:f78873a0c87bf46979face1a0a900cfb8ad7fcf85bcd18c698bebeda4c5d0d9f

Observation e298fb75-861e-4e5f-b1d5-32b3bd102811 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.355742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.355742Z digest=sha256:06c65fb23ba73c92eaabc78a64f5845270e57fabd34a53e0314ff1cdd51bb7ba

Observation b517a61d-d8bc-4ca9-ae08-daba4705f3d7 · outbound

This paper cites Interacthuman: Multi-concept human animation with layout-aligned audio conditions, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Interacthuman: Multi-concept human animation with layout-aligned audio conditions, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.359072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.359072Z digest=sha256:326b95d7df43c0512123a575612dad4b2c217f86e9fb70b39dad132f3c7ef5c5

Observation ca611332-8519-4922-aa86-96592f562257 · outbound

This paper cites Mocha: Towards movie-grade talking character synthesis, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Mocha: Towards movie-grade talking character synthesis, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.362761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.362761Z digest=sha256:9c0dc92df058512d19be684c0a42c41617fa391118ab079c63915c5d6b90ba95

Observation 5d0c4ebe-897b-41e1-9b3f-afff7b7f7a52 · outbound

This paper cites Magicinfinite: Generating infinite talking videos with your words and voice, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Magicinfinite: Generating infinite talking videos with your words and voice, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.366342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.366342Z digest=sha256:69e20bdcdeb62cb0f998478be6e22d4df3dea778ed308ea597e68e49e593e9f0

Observation 200ba8ae-e51a-4cd2-b076-cb23f14777c6 · outbound

This paper cites Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.369765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.369765Z digest=sha256:73dbdc293014c09c76a6a524a4b137a7f28bd4588ba29ba26613908a09417f63

Observation 437606c5-bc8c-40c1-ac7f-c34dc4d3a8af · outbound

This paper cites Identity- preserving text-to-video generation by frequency decomposition.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Identity- preserving text-to-video generation by frequency decomposition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.373373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.373373Z digest=sha256:cd9241a5a94d91e7fa89b9be5efa4789b0b0f8d53372d0c4bf24f780ca60cedf

Observation d05a74d5-491e-44ba-a394-a08056f9ec64 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.300454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.300454Z digest=sha256:6d618acb31e3bd108068b50cd6a4caf9d2d80be4bf822cd6934a7d4575dbcf88

Pith citing papers

Observation cc157498-875d-48b4-a842-beae10033b8d · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.211260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:c2a58ced38c4191a54ef3188fa5220d3713e8bfe995be7b0ce657eb8afb38a38

Observation b2177093-0791-4c79-9f92-221d3a59a587 · inbound

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection cites this paper.

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:48:58.476215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-17T03:48:26.807495Z digest=sha256:5f5eb3a8008cf751d4945f10d68f9a6eea786a8a772b8ac5b8f57f22750fd922

Observation d19122a2-98f1-43bd-a2c9-84765f842cc7 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.907924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:3e3c924d06c666279cf16200f6611c5bd386c8aa834b79bc979b344b1b90489d

Observation f09ff69a-813c-46be-8db5-611a115730d0 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.230425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:81fde78adc6bd20eb2ef1179e33637f02b52a0b403780b69bb6f9ad2daac571e

Observation c43748d0-94d2-4373-8142-ea99c47050ad · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.511227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:0e6a5e1d23d1c6172258886efeb403fbd717169f4b39d15b2036c4a60a3ee6af

Observation 7617bc1d-d26a-4f65-90f6-a0409ef9d3ad · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:10.037400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:10.037400Z digest=sha256:1eee5e287c88c34eeb891f15579e3660b9ce12fcef3ed66d271762dcc8c74678

Observation 64187c61-3b4e-499a-ae99-90f38bc8c8d4 · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:1f66d710b25a142db911c6276e604e94f4e5962e42172b58c0ac2380f90aead4

Observation e55e8956-b8f0-4f28-8a80-3d82ee835944 · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.110781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.110781Z digest=sha256:6187c8b075a38c181e2d30f674db8c53b8237d64bd46ab1b3a41890d721c09e1

Observation c0aa2cf3-d9d9-4912-b478-6bb81d717bd9 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.846071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:0849647b03b161c99b127c7deb24f80343df299c9030a1b96b2e3c91373e3960

Observation 3fbbb5f7-a58b-429f-a599-051e32a9f910 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.305402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:0c27f781cec265c496a6727fc96bfc59ceb6871acfdf222e20eeaa6f12d95bd4

Observation 6a0795f9-e1a4-4d76-8c84-64a7468ddddd · inbound

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis cites this paper.

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:02.046272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T02:24:57.882447Z digest=sha256:e2ee8b7e77eb5c4b8589b1bbfd66ed28d459c5068e9c1c70285031a5c7143e95

Observation 4c170827-95c6-4e97-848b-14e456023853 · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.856032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:007ffe0d6fda097883afbf471ca5f056ee802be367f86e648a0fdd30f988044f

Observation 43191b27-afa2-4dd0-ab05-1b3066bfbcb7 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.416194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:e7d04b9b3d5897e7245b205f8b0644f2f141f08ba6a3f53260d78b65bbd877a4

Observation c6630e7a-2ec7-401d-a29e-2a57d085da97 · inbound

Aurora: Unified Video Editing with a Tool-Using Agent cites this paper.

Aurora: Unified Video Editing with a Tool-Using Agent HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:48:12.712658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:47:05.038308Z digest=sha256:7b35ed1d52e2c0062d7271441ca55ffb57ef320117843b8ca7ca2d1a860c9dc5

Observation ecb2003d-3b8c-4c9f-bf19-2a13a20370de · inbound

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models cites this paper.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:41.246048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:39b61aebbc7e9bddb590e83132250d65ee15adb41b751ee9985317b8583e8089

Observation b3cb5746-2fbf-4d8a-9bfd-8afa6b38967f · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.636837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:f9e05caf1d71341cda2b636a25575d1c064a3fe06648f27adc332015486ccb71

Observation b2f10b8f-5a9c-4371-9d50-d5d54890b41c · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.473972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:f514ae8a46934b7c00bc8cb4a04590b19c7612c6630ed930cac0399f255c8206

Observation 23737e46-c8ca-483b-a636-1a1101528f32 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.651576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:786a7c0d4e8ea78863d0cdf90fc4747c6a42c0f9a909b4bb65d7262cfb1f00fd

Observation f6df4e3b-1a50-4e2e-a38f-4b7da48f6258 · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.678040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:9e85c818a34689f8c90c64c0c3612ccd54b442057b23930d1f6508f20d5c4d49

Observation 8f0d2e04-551f-41c2-9952-bb51c793bb71 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:e3362325156a05a0453b1e5b2e02bdd8f28e27d8e9be8136088e1a5022d7b60b

Observation a7effad9-82da-4ba1-ae87-441f648ab486 · inbound

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation cites this paper.

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T10:37:57.061257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:37:57.061257Z digest=sha256:e4e2b8ff2c3affad684643273c47d498685badb3f439817416b121862d8717f3

Observation 65597ed5-91f9-48d8-8830-5a7d2d18c6a9 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:15.266478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:15.266478Z digest=sha256:bd30ce085aab8816d9fe7f7b42f91d1f63993df89edaa295460c361472de2a80

Observation 8917a336-66f4-4a14-b688-44e063684b20 · inbound

ID-V2V: Identity-Preserving Video Restylization cites this paper.

ID-V2V: Identity-Preserving Video Restylization HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:31.607132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:27:31.607132Z digest=sha256:7b80da79ab98e5db426ccbbd32aebf80869af6992e8b13e1095ecc2b7c1edbf4

Observation 99433b28-261f-4369-903d-a30387fef95f · inbound

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction cites this paper.

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T20:31:38.681716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T20:31:38.681716Z digest=sha256:6072899e1f40804323bb041f33ca8a04f1ac668310e594cae1555bd3613ca4c4

Observation ff8d26ba-0435-4682-9be0-1eaf459188d7 · inbound

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing cites this paper.

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T13:04:14.380575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:04:14.380575Z digest=sha256:87d4fc2786d89f8e7185643972cb988b55164cf6f58d1bec342c6bb79699721b