Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:35:30.373373Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 25 inbound Pith citation observations for arXiv:2509.08519.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:35:30.373373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:04:14.380575Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T21:10:09.676485Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f6e52b1-42c4-4c68-8d10-025ce616f829 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Qwen2.5-vl technical report, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cff7969-361a-4ced-bf5a-e069987fa478 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Goku: Flow Based Video Generative Foundation Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a744eb-4b69-4c52-a723-e40767a1831a · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Phantom-data : Towards a general subject-consistent video generation dataset, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01de29e7-9c07-4e5c-9b3b-f2d0a4f64690 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73526c4d-fef5-4337-a088-f151684a4563 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1982d1ed-fb0e-4aa2-a0c1-7a6d565eae51 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Arcface: Additive angular margin loss for deep face recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–1, 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04f3538-3bc5-484b-ac68-160620b20b12 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Magref: Masked guidance for any-reference video generation, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44100d2c-905c-4bed-8b66-757f6c16f9e7 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seedream 3.0 technical report, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c364d4-1825-4b38-87bc-97a4f1befbf4 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seedance 1.0: Exploring the Boundaries of Video Generation Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c884a46-d3b8-45aa-acca-7cf61f7e5515 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b99b558-d953-4af1-bf04-df27b6da5e97 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 863638f1-8ad9-4be2-a9a8-6b58c0f34f5c · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Curricularface: Adaptive curriculum learning loss for deep face recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5c714b8-89c6-4e78-8a3d-a5fdac551aec · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Vbench: Comprehensive benchmark suite for video generative models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdbdf8e8-125d-48c5-959f-175d01f7c997 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Multi-reference images to video generation feature
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e02591c5-48db-4050-9bf1-7b32005ecb64 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85d97a2-7601-46d2-a88a-5a5e715ee43d · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Let them talk: Audio-driven multi-person conversational video generation, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8056d4d-a20d-4e3a-82c0-26d76dd245ec · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe98423-05f1-4ca0-ad34-4ad69779c0d0 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d9df52a-35b1-4f13-a357-79421bbbcf0b · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b180c7-8a38-4f89-8ad7-61606f1148ea · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d52d6e-60b5-4514-95eb-4cfa52ede0b2 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Improving video generation with human feedback, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dac79d-2b20-4695-8017-7d86e6f3cecd · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Phantom: Subject-consistent video generation via cross-modal alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8edc8a2e-5f5e-405f-83ba-f097b569c2e0 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcca47f5-00a3-4da8-9eea-bfe8a32261de · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Movie Gen: A Cast of Media Foundation Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd519e46-b700-4561-93c2-7c6a130b742e · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Learning transferable visual models from natural language supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e35adaba-2d5d-45de-a961-6add66bae3d7 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Robust speech recognition via large-scale weak supervision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1342596d-7656-46e5-9770-fa20a0badf5b · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404e1143-193f-435e-a946-c7821ae3966b · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seitz, and Ira Kemelmacher-Shlizerman
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a12392-8187-48c7-99b0-49621ece1022 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Gemini 2.5 flash image
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5015b6f4-0fdf-4cc8-99f7-331a36e5d40e · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Gemini 2.5: Our most intelligent ai model.https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6761de8e-cee4-4ef1-8e5c-626d1aece653 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Wan: Open and advanced large-scale video generative models, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535abdab-5554-4da2-b606-24ca809d2cbc · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Fantasytalking: Realistic talking portrait generation via coherent motion synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e298fb75-861e-4e5f-b1d5-32b3bd102811 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b517a61d-d8bc-4ca9-ae08-daba4705f3d7 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Interacthuman: Multi-concept human animation with layout-aligned audio conditions, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca611332-8519-4922-aa86-96592f562257 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Mocha: Towards movie-grade talking character synthesis, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0c4ebe-897b-41e1-9b3f-afff7b7f7a52 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Magicinfinite: Generating infinite talking videos with your words and voice, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 200ba8ae-e51a-4cd2-b076-cb23f14777c6 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437606c5-bc8c-40c1-ac7f-c34dc4d3a8af · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Identity- preserving text-to-video generation by frequency decomposition
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05a74d5-491e-44ba-a394-a08056f9ec64 · outbound
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc157498-875d-48b4-a842-beae10033b8d · inbound
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2177093-0791-4c79-9f92-221d3a59a587 · inbound
MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d19122a2-98f1-43bd-a2c9-84765f842cc7 · inbound
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f09ff69a-813c-46be-8db5-611a115730d0 · inbound
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c43748d0-94d2-4373-8142-ea99c47050ad · inbound
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7617bc1d-d26a-4f65-90f6-a0409ef9d3ad · inbound
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64187c61-3b4e-499a-ae99-90f38bc8c8d4 · inbound
What if? Emulative Simulation with World Models for Situated Reasoning HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55e8956-b8f0-4f28-8a80-3d82ee835944 · inbound
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0aa2cf3-d9d9-4912-b478-6bb81d717bd9 · inbound
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fbbb5f7-a58b-429f-a599-051e32a9f910 · inbound
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a0795f9-e1a4-4d76-8c84-64a7468ddddd · inbound
ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4c170827-95c6-4e97-848b-14e456023853 · inbound
Generate Your Talking Avatar from Video Reference HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 43191b27-afa2-4dd0-ab05-1b3066bfbcb7 · inbound
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6630e7a-2ec7-401d-a29e-2a57d085da97 · inbound
Aurora: Unified Video Editing with a Tool-Using Agent HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ecb2003d-3b8c-4c9f-bf19-2a13a20370de · inbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3cb5746-2fbf-4d8a-9bfd-8afa6b38967f · inbound
LongCat-Video-Avatar 1.5 Technical Report HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2f10b8f-5a9c-4371-9d50-d5d54890b41c · inbound
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23737e46-c8ca-483b-a636-1a1101528f32 · inbound
HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6df4e3b-1a50-4e2e-a38f-4b7da48f6258 · inbound
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f0d2e04-551f-41c2-9952-bb51c793bb71 · inbound
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7effad9-82da-4ba1-ae87-441f648ab486 · inbound
StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65597ed5-91f9-48d8-8830-5a7d2d18c6a9 · inbound
AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8917a336-66f4-4a14-b688-44e063684b20 · inbound
ID-V2V: Identity-Preserving Video Restylization HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 204
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99433b28-261f-4369-903d-a30387fef95f · inbound
Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8d26ba-0435-4682-9be0-1eaf459188d7 · inbound
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.