Pith. sign in

Paper Citation Record · LEDGER

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 19 inbound Pith citation observations for arXiv:2506.03099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03099 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:13:58.243293Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:45.075199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.653031Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be2a126a-4132-448e-a691-6638350343c6 · outbound

This paper cites Paddleocr, awesome multilingual ocr toolkits based on paddlepaddle.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Paddleocr, awesome multilingual ocr toolkits based on paddlepaddle

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:55.247754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:55.247754Z digest=sha256:e02beb028b783c4e24e0216a9ec753386a73ae801a7c4b5faf4e9f2810640319

Observation 0f2650c3-b200-4c0a-9c7d-954cb1055510 · outbound

This paper cites Video generation models as world simulators, 2024.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Video generation models as world simulators, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:01.986958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:55.363593Z digest=sha256:c48030aac6c11722c23ef71a43a55331de028ccc2d0a91aab66cb5dc90d2764d

Observation 2ebadf33-3ada-42e2-b5d9-3759e9261aa0 · outbound

This paper cites Pyscenedetect: Python and opencv-based scene cut/transition detection program & library.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Pyscenedetect: Python and opencv-based scene cut/transition detection program & library

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:01.736592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:55.459398Z digest=sha256:b48a6efe038ec33734ca4bd31452259542a14d5fecb682ea675efb07566208c5

Observation b7d2e49b-57fc-45f2-969d-6c9133ef8812 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Diffusion forcing: Next-token prediction meets full-sequence diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:55.541205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:55.541205Z digest=sha256:f328f96b8b0e44442c0624426f84ba398653589eb148984c7a0e165fa851ba0b

Observation 44305ca2-d0ae-454e-9325-cdafbc7213f1 · outbound

This paper cites Out of time: automated lip sync in the wild.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Out of time: automated lip sync in the wild

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:01.392067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:55.615493Z digest=sha256:f1d3caa412afb6ecb4cc6514ebf0978f167c45070665bfdda37bb13fefc50544

Observation 788de9a0-d16f-45c6-8538-c159f581c7c2 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:55.731959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:55.731959Z digest=sha256:87ccba08ee524a474e12e65cbafe2edb5889ddddc81ddaa1c3a56de12b2f97d9

Observation daa3d8d8-b633-412b-a673-4d584f534ed0 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:55.862202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:55.862202Z digest=sha256:ab746e1deab8bdb7c497d2e45119a4246e1656041a81386ccf04961335c18336

Observation 26dbfc62-4f82-40aa-8f92-a91c286423cc · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:01.052274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:55.961042Z digest=sha256:d639551a49e330773c3a5afa415829b6eb10447426e3d850d66007382602e0cf

Observation ee1ae07c-8a14-4eae-a7c3-456db9e08f60 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars, 2023.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Megaportraits: One-shot megapixel neural head avatars, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:00.775196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:56.077398Z digest=sha256:fbb9583ef71b70daecfa8e861df0b03515d243e2d2d7ee22fb585fc12bddce3a

Observation b0ed24da-a047-4360-8a37-57dee11d2c5d · outbound

This paper cites $R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models $R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:56.220334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:56.220334Z digest=sha256:125a20e1eb59aa7df2911a89c2a93991606e46859cb471c0392334414618a47b

Observation a333e764-c15a-48ab-9d4c-c2793ff3c004 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:56.315933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:56.315933Z digest=sha256:0f5dfb84079cfe4300f722552a837f40db69434560b8e23c060fc8098ae8666e

Observation 3617d214-d00d-4a79-991e-e21383fe4cdb · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:56.372275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:56.372275Z digest=sha256:2aa2bb67ffaf67e31e0de08dea972c7a531fba57eb1b385439d3d8026cbd978e

Observation e577f335-9c5d-4610-a9bd-fbc087ae0c2d · outbound

This paper cites An introduction to variational autoencoders.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models An introduction to variational autoencoders

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:00.516620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:56.439850Z digest=sha256:fb2f0a56f439b29ec3e8327092474e023f6335eef2ea81b6930fa9e1e95ebbff

Observation 9935acfc-626f-48b4-9bdb-8e0340511a27 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:56.524029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:56.524029Z digest=sha256:9f44385d533ade0990a13a1badff51d3661d55d0c5258932ffdcc116f20948f9

Observation 107e56ae-08ef-443f-8670-4dba779c53d8 · outbound

This paper cites Livekit: Open-source webrtc infrastructure for real-time audio and video, 2025.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Livekit: Open-source webrtc infrastructure for real-time audio and video, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:00.185410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:56.608705Z digest=sha256:d61435981c64ee99531ebcf15c882870c5b896dacf12f9301d83db5b5c6f0c35

Observation 1cae25c6-955c-4203-9488-c8445c07b4db · outbound

This paper cites Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for? Queue, 6(2):40–53, 2008.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for? Queue, 6(2):40–53, 2008

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:13:59.908784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:56.709513Z digest=sha256:6a443192f55e8d5fe3170688e851a5cccc4a19693b32ad68f09fe3c7293329a7

Observation 5236f2fc-6bd4-4bbe-b170-0bc0b21ebc4f · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models, 2020.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Zero: Memory optimiza- tions toward training trillion parameter models, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:56.813355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:56.813355Z digest=sha256:6834452c0590dfb6802a3c8ebe3fafb9159afa99f9ae5ff2974bfead49cc34ff

Observation bfbccb1d-a200-4167-a993-4932198e9b82 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models High-resolution image synthesis with latent diffusion models, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:56.919267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:56.919267Z digest=sha256:5323dc31c0f62d1355aae4700f7f78a64e61cef7ae84514549b1697ceff2d43f

Observation 4339643c-837b-4050-ba16-5cbde8336b16 · outbound

This paper cites Laion-aesthetics: Predicting the aesthetic quality of images, 2022.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Laion-aesthetics: Predicting the aesthetic quality of images, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:13:59.631482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:57.036868Z digest=sha256:48b285988d7fad63e13b0ac1718842ef2d7e6679145434c0eea8e56df61c43d2

Observation 2c48512b-077c-4442-8d15-12d8e7565650 · outbound

This paper cites an unresolved cited work.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.139344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.139344Z digest=sha256:ee33be23a97063b11b625d212b01011813ad282affe99c32bd06ad38444d897c

Observation b92c60cb-edd0-47e8-b484-7aebe67e433f · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Raft: Recurrent all-pairs field transforms for optical flow

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.276095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.276095Z digest=sha256:4dc6a2f2e7b2baa30ef9f50ae05a539d5999dbd7d0359a10030fcb058577c012

Observation 82a768ea-2ffd-4e46-9cf5-e023d00b4bcc · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.339398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.339398Z digest=sha256:bf964941c2c75636c42ab5f527d0aa61212707c630bd1e00ae5c95de1dc679e7

Observation a054e9ea-4728-44a9-b187-273d4de7624a · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.442069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.442069Z digest=sha256:940f98292acb89d71f1ef8ae50dfad99e54a25aca67f1e6329a21aab377b141b

Observation 741262d4-f416-4402-b419-9441062c00fe · outbound

This paper cites Attention is all you need.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Attention is all you need

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.556565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.556565Z digest=sha256:e17cdcacb8d10eff5b673e51285f0941d24a5c252a1eaf49afacc4b5af442a9d

Observation d6f1d52f-ba1e-485d-adcc-71b92e63aacd · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.658994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.658994Z digest=sha256:5d058a9a2b6d6e8e9e8dc4674f88df1f4b9967a21e96e6da1ce2c6ab9757ea1b

Observation efd8088e-857c-4e2b-a946-8d573e7cf149 · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time, 2024.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Vasa-1: Lifelike audio-driven talking faces generated in real time, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:13:59.367704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:57.765627Z digest=sha256:c1ebb5a98acb76b98976c0fdef533880dc209959e28628ca717270943a8c7715

Observation 58efbcb1-e9d4-47f8-ad9f-71c77983bf46 · outbound

This paper cites MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:57.860540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:57.860540Z digest=sha256:856ebeb42e1b6554e6a69f192be320c84ecdcfc0cbcbedd440abed0b45e89a6d

Observation 67181f11-e43f-412b-b4fd-350470c0bc1e · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Improved distribution matching distillation for fast image synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:13:59.079749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:57.946790Z digest=sha256:41aee73a3e8873f93618ccce6c9228e280d1f3e111acefc7e0f30a8fbf9be545

Observation 72504a90-1dbd-436a-9501-d36414f2d534 · outbound

This paper cites One-step diffusion with distribution matching distillation.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models One-step diffusion with distribution matching distillation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:58.034163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:58.034163Z digest=sha256:e2ef8aeb16eb248bc5b7724ed0a3e05b2f6ea6596d41bae759330639cec185fd

Observation 464facd5-8605-4f95-bc3b-2eabcd8ff008 · outbound

This paper cites From slow bidirectional to fast causal video generators.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models From slow bidirectional to fast causal video generators

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:58.127451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:58.127451Z digest=sha256:820596eb08d454cbc4b767a102eca07faddbcdb95704553162a4ba115fd4207a

Observation ff555d9c-c9fe-401f-8fd9-96615bb22c2d · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023.

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:13:58.703078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:13:58.243293Z digest=sha256:05c5f9590a0fe328917871697fba91ebfafb7ff13090a4b91c67ff481f1ac0a6

Pith citing papers

Observation 9a0ede31-7c2c-4aaf-ade2-f04637144bc7 · inbound

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation cites this paper.

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:45.075199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:45.075199Z digest=sha256:ec9fa1b4665b64dce989b5af4f9de30d15889a2b69b9cfd6fbfce169ed0937e3

Observation fe4cfe71-a521-4a47-b5ef-56faa10a8e68 · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.768612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.768612Z digest=sha256:98052518e8202b2a015b8e66a721defd8c161eb1c12ac61b7507c0c1d24175e1

Observation f1ed816c-b2b5-481e-9245-bf9d2eeba0e2 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.270808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:b6682e1b3dd457ed9342acb226efd2e09fb30e5fd217670269ade0d3846085fd

Observation 01cebfeb-107d-4571-a151-bcf6615959b0 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.445487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:9c0ce52fffb45fd4e4938dd7432431570a229adc879b7f41261cb139ed53ab76

Observation 4a3213b5-584f-4e40-8b81-1465a29f7d90 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.890004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.890004Z digest=sha256:a6c24f847364232d1900c712028ab2acc4af1cd493432f6141d554b63d52c6c2

Observation de620223-5563-46f0-a22e-37ff66af5d5f · inbound

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation cites this paper.

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T13:04:40.773631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:04:40.773631Z digest=sha256:3224a2fd81b4c58e65b3dfa4c1524042d02d5e2ecfc586e41ae0aaf5f2344915

Observation 27bc7c80-632f-4dde-a893-ec739075638f · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:07:29.811065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:29019619b667ae59f034c08dbc2c7064aa86a238d19f21a0aa770ab2a29c828b

Observation 8c2b8408-c1d2-412c-a640-b66edd062f47 · inbound

LPM 1.0: Video-based Character Performance Model cites this paper.

LPM 1.0: Video-based Character Performance Model TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:59.749809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:09:17.360697Z digest=sha256:3441ab04d7c5576c2a372bee0076b4e2db9c72e1b34fc1b21ca69bd4a17af040

Observation 89931edb-4006-4421-83fe-abdf5621cfe1 · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:25.925005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:052747df914406c387f965ccfd843022fcb1fe6c6dd6fcbdecc7c977a751614e

Observation d46b214c-f9ee-4c93-ba81-3dab2d71ab11 · inbound

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment cites this paper.

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:04.247728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:30:18.645845Z digest=sha256:078b3d2d5069fea2b3d78db89504e6936eedf83a7758ecdb0128de900fcc212b

Observation 236a9de2-7f14-49b4-af42-24272acc0e49 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.053134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:43eec23a758e326177ca6ab815e505c99c0706ac735274614ea067e413b8b09d

Observation 8824b76c-9096-4e16-acbf-911c9d25e250 · inbound

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs cites this paper.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.028231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:4dbf972fd3becd80ae077e25286681013fb6aabfc3f02808ecbe921074139d8a

Observation ffbd3f46-e128-4644-a43e-155a3cff8aa1 · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.588704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:09:06.925645Z digest=sha256:e21b4050b59b7efa2f724dc638dc287d3d9530d34262dec13870437a61273495

Observation 20453d1d-5921-43a2-8ac6-dbd35725fa69 · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:29.162859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:00:53.496569Z digest=sha256:06181a53a780f655035268a317f5e1c533e3c62b1491ef8df2784c002afcb799

Observation 43dfb4ac-d655-4536-92da-3ed0321accc6 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.654524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:13:38.523397Z digest=sha256:43be915e4a1e4a1d748b64d530a16c883f15e397f1ff78da22d1d515253e6c5d

Observation 9beee4c7-9877-4738-9d74-9ecb7464a9b0 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.984847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:27:24.078053Z digest=sha256:1f35bfa0dc719948f6018344bf1965cb9fcdde3cf02574bfe50eb2f575cede18

Observation 7c501338-3a02-431f-9558-da41b865ca4f · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.396418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:43:09.347118Z digest=sha256:7076676cf591a0fb0df3f87052e1ccec67b52c21aca050244e1697347e07851e

Observation 7f04e79c-db6c-4a78-a9e3-beeb532e6853 · inbound

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars cites this paper.

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:55.381722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:52:55.381722Z digest=sha256:ffbc63f799bb0dcb868a1c034380cb0c23a013460003e392a959eba837ce53b0

Observation 0c8ab426-3d39-468f-8d82-a2b5be377eeb · inbound

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos cites this paper.

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:25.351877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:25.351877Z digest=sha256:d2bd1cfe15e5ebb352fd012ed74f367e8be50d2e3a868ef9589101de9aef5ef0