Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:01:30.793099Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 14 inbound Pith citation observations for arXiv:2508.19320.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:01:30.793099Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:45:46.122397Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:49:57.706870Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 17051457-5094-405d-b920-58048bd69c7d · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation LRS3-TED: a large-scale dataset for visual speech recognition
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9079251d-3480-4c4e-a654-a288346b7d44 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58175350-9275-4437-beee-9c33f9f6a819 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e2ecc3-0a5f-4fa9-ba94-ac0d8c4a75e5 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7dc1f40-57f5-44e4-a77f-f664e284d62b · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b49db34c-bbd5-4f14-8308-656af95ccf76 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe4cfe71-a521-4a47-b5ef-56faa10a8e68 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9928428-2bc5-4a59-b748-6472665c45af · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation V oxceleb: A large-scale speaker iden- tification dataset
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a1348e7-2ac6-4672-8192-6cbeaeab58f0 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Robust Speech Recognition via Large-Scale Weak Supervision
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be07b53f-fc56-413b-b096-767400a84c2a · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MAGI-1: Autoregressive Video Generation at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3902634-2cbb-4768-91eb-efa0626f97cf · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5abb561e-3940-4e21-9c09-452074a8ff51 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Diffusion Models Are Real-Time Game Engines
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3361aa-f299-458b-9e9d-02fd7de75b16 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9078eba-f632-477d-9580-59079727a280 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Body of Her: A Preliminary Study on End-to-End Humanoid Agent
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3481134-6850-4cd3-8ffa-015982987b44 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3dc2ccd-0c20-40ec-9c59-4cc98d8737e8 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Long-Context Autoregressive Video Modeling with Next-Frame Prediction
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b852dd8-9b9b-4d85-8dee-48b1a43d6c9b · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Qwen2.5 Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637da00f-17f4-4838-b577-1697b3dc2884 · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e51af2a-bd1f-4de3-bd17-ae5ac711161a · outbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation V oxceleb2: Deep speaker recognition
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53c7a565-32d3-4fa6-b7a2-cba71f4677f3 · inbound
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cb89d86-de96-464c-a2f5-08274b1f2d33 · inbound
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108d7ebf-e767-4e25-bd24-cec23c714294 · inbound
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 665fde9b-a8a6-4d66-a681-4e7eac0bc3bb · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5db50733-5e4c-431f-b55a-f56a089308b9 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8026f9-42e5-4e45-b571-5ba2bc0952fb · inbound
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2f4e523-8418-4ece-a17a-6951507b5999 · inbound
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation add3ca7b-f2a2-4f9d-a150-b6b9f950d996 · inbound
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49a4d07d-3d86-40f9-95b8-099f10a15b9a · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7e421ae-0ffe-48aa-b59e-219dd8b4607d · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 482642b1-8e99-4271-acf3-86030bcb6b72 · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e7788fd-13a0-44e1-ab20-7915a5a6513e · inbound
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55e9fcb2-b64a-476c-802f-ffb5b005e91e · inbound
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d958e652-9421-4996-b66f-9ef6324edd8e · inbound
EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.