Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:13:58.243293Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 19 inbound Pith citation observations for arXiv:2506.03099.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:13:58.243293Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:45.075199Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:49:57.653031Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be2a126a-4132-448e-a691-6638350343c6 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Paddleocr, awesome multilingual ocr toolkits based on paddlepaddle
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f2650c3-b200-4c0a-9c7d-954cb1055510 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Video generation models as world simulators, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2ebadf33-3ada-42e2-b5d9-3759e9261aa0 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Pyscenedetect: Python and opencv-based scene cut/transition detection program & library
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7d2e49b-57fc-45f2-969d-6c9133ef8812 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Diffusion forcing: Next-token prediction meets full-sequence diffusion
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44305ca2-d0ae-454e-9325-cdafbc7213f1 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Out of time: automated lip sync in the wild
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 788de9a0-d16f-45c6-8538-c159f581c7c2 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa3d8d8-b633-412b-a673-4d584f534ed0 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26dbfc62-4f82-40aa-8f92-a91c286423cc · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ee1ae07c-8a14-4eae-a7c3-456db9e08f60 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Megaportraits: One-shot megapixel neural head avatars, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b0ed24da-a047-4360-8a37-57dee11d2c5d · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models $R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a333e764-c15a-48ab-9d4c-c2793ff3c004 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3617d214-d00d-4a79-991e-e21383fe4cdb · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e577f335-9c5d-4610-a9bd-fbc087ae0c2d · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models An introduction to variational autoencoders
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9935acfc-626f-48b4-9bdb-8e0340511a27 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107e56ae-08ef-443f-8670-4dba779c53d8 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Livekit: Open-source webrtc infrastructure for real-time audio and video, 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1cae25c6-955c-4203-9488-c8445c07b4db · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for? Queue, 6(2):40–53, 2008
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5236f2fc-6bd4-4bbe-b170-0bc0b21ebc4f · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Zero: Memory optimiza- tions toward training trillion parameter models, 2020
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfbccb1d-a200-4167-a993-4932198e9b82 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models High-resolution image synthesis with latent diffusion models, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4339643c-837b-4050-ba16-5cbde8336b16 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Laion-aesthetics: Predicting the aesthetic quality of images, 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2c48512b-077c-4442-8d15-12d8e7565650 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92c60cb-edd0-47e8-b484-7aebe67e433f · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Raft: Recurrent all-pairs field transforms for optical flow
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a768ea-2ffd-4e46-9cf5-e023d00b4bcc · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a054e9ea-4728-44a9-b187-273d4de7624a · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741262d4-f416-4402-b419-9441062c00fe · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Attention is all you need
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f1d52f-ba1e-485d-adcc-71b92e63aacd · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Wan: Open and Advanced Large-Scale Video Generative Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd8088e-857c-4e2b-a946-8d573e7cf149 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Vasa-1: Lifelike audio-driven talking faces generated in real time, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58efbcb1-e9d4-47f8-ad9f-71c77983bf46 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67181f11-e43f-412b-b4fd-350470c0bc1e · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Improved distribution matching distillation for fast image synthesis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72504a90-1dbd-436a-9501-d36414f2d534 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models One-step diffusion with distribution matching distillation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464facd5-8605-4f95-bc3b-2eabcd8ff008 · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models From slow bidirectional to fast causal video generators
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff555d9c-c9fe-401f-8fd9-96615bb22c2d · outbound
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a0ede31-7c2c-4aaf-ade2-f04637144bc7 · inbound
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe4cfe71-a521-4a47-b5ef-56faa10a8e68 · inbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ed816c-b2b5-481e-9245-bf9d2eeba0e2 · inbound
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 01cebfeb-107d-4571-a151-bcf6615959b0 · inbound
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a3213b5-584f-4e40-8b81-1465a29f7d90 · inbound
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de620223-5563-46f0-a22e-37ff66af5d5f · inbound
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27bc7c80-632f-4dde-a893-ec739075638f · inbound
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c2b8408-c1d2-412c-a640-b66edd062f47 · inbound
LPM 1.0: Video-based Character Performance Model TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 89931edb-4006-4421-83fe-abdf5621cfe1 · inbound
Efficient Video Diffusion Models: Advancements and Challenges TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d46b214c-f9ee-4c93-ba81-3dab2d71ab11 · inbound
PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 236a9de2-7f14-49b4-af42-24272acc0e49 · inbound
Image-to-Video Diffusion: From Foundations to Open Frontiers TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8824b76c-9096-4e16-acbf-911c9d25e250 · inbound
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ffbd3f46-e128-4644-a43e-155a3cff8aa1 · inbound
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 20453d1d-5921-43a2-8ac6-dbd35725fa69 · inbound
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43dfb4ac-d655-4536-92da-3ed0321accc6 · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9beee4c7-9877-4738-9d74-9ecb7464a9b0 · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c501338-3a02-431f-9558-da41b865ca4f · inbound
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f04e79c-db6c-4a78-a9e3-beeb532e6853 · inbound
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8ab426-3d39-468f-8d82-a2b5be377eeb · inbound
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.