Pith. sign in

Paper Citation Record · LEDGER

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 14 inbound Pith citation observations for arXiv:2508.19320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19320 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:01:30.793099Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:45:46.122397Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.706870Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17051457-5094-405d-b920-58048bd69c7d · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation LRS3-TED: a large-scale dataset for visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.733718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.733718Z digest=sha256:e589d9215457bea10a1cf74d3c96523e9a4b1dc317579e37652d1fed68081f6f

Observation 9079251d-3480-4c4e-a654-a288346b7d44 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.743448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.743448Z digest=sha256:928e912b4dd7cd0675fe4c207622503e79dd4100a69027c426a8971ed67c1042

Observation 58175350-9275-4437-beee-9c33f9f6a819 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.751310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.751310Z digest=sha256:52cd24339940778fe4a27d9c2f2c248af3717e7e44362631fb994ee9a763e7fd

Observation 87e2ecc3-0a5f-4fa9-ba94-ac0d8c4a75e5 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.758960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.758960Z digest=sha256:22558ef3bd411fa65252a7cd4c7b90d00af64b6bf11cbff0a8eeb3b8ba9b3192

Observation b7dc1f40-57f5-44e4-a77f-f664e284d62b · outbound

This paper cites Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.762208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.762208Z digest=sha256:c37034efbe1b51a4e91cb6a38671b2c121eef1300beb4d4f683b22faea5eee68

Observation b49db34c-bbd5-4f14-8308-656af95ccf76 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.764940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.764940Z digest=sha256:d27cc7a171bb3b309b7360e11abc5036ed29128c1397238a4f6c902b4d6c91df

Observation fe4cfe71-a521-4a47-b5ef-56faa10a8e68 · outbound

This paper cites TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.768612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.768612Z digest=sha256:6811b1df6b406434d1ff4c54ebca24d317781c2d3758a6c8975c2d88d8bcc1f2

Observation a9928428-2bc5-4a59-b748-6472665c45af · outbound

This paper cites V oxceleb: A large-scale speaker iden- tification dataset.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation V oxceleb: A large-scale speaker iden- tification dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:30.946894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:01:30.772271Z digest=sha256:631e8f75c1aa5cf72720b2a9443ffb1ab3d4a5135c3de9948062832ba2558874

Observation 2a1348e7-2ac6-4672-8192-6cbeaeab58f0 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.778327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.778327Z digest=sha256:eefdcf046e15f10d51fcf8ae48b5bc4e73febf4e6b0031350cdba319c8967efa

Observation be07b53f-fc56-413b-b096-767400a84c2a · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MAGI-1: Autoregressive Video Generation at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.781337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.781337Z digest=sha256:9d170c663805d9bb4229b8920e69921caf859127b666aac1a455d4be1674654e

Observation c3902634-2cbb-4768-91eb-efa0626f97cf · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.783917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.783917Z digest=sha256:c1651448f2f9988be39053065f6f060d4bc5f9e53cba35ed4e767c7c0910db79

Observation 5abb561e-3940-4e21-9c09-452074a8ff51 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Diffusion Models Are Real-Time Game Engines

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.787197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.787197Z digest=sha256:580f7380a10deeb070e05d0b54ffef2e6eb479a2ba8cf5ff61fc6ba41f42ff6e

Observation fb3361aa-f299-458b-9e9d-02fd7de75b16 · outbound

This paper cites ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.775017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.775017Z digest=sha256:077ecf4b9b33dbdd563e0faa52151b83efc2f7b6e1de556de530def88bacc617

Observation e9078eba-f632-477d-9580-59079727a280 · outbound

This paper cites Body of Her: A Preliminary Study on End-to-End Humanoid Agent.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Body of Her: A Preliminary Study on End-to-End Humanoid Agent

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.737383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.737383Z digest=sha256:e17d0b333f2689b4fee57e3a82d6e81da51586fa939743e7a785ed41d9891a19

Observation e3481134-6850-4cd3-8ffa-015982987b44 · outbound

This paper cites MoCha: Towards Movie-Grade Talking Character Synthesis.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.790252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.790252Z digest=sha256:b1d04764eb70ac82697ecd6a626b5c468393f851da367669a88dbdd310df14bb

Observation c3dc2ccd-0c20-40ec-9c59-4cc98d8737e8 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.747278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.747278Z digest=sha256:c454875eb6b369a4ee30d8c681c21781786824ed13dd32b2f08e2aa763fafff5

Observation 6b852dd8-9b9b-4d85-8dee-48b1a43d6c9b · outbound

This paper cites Qwen2.5 Technical Report.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Qwen2.5 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.793099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.793099Z digest=sha256:2a5e4362bc1c1ef327da65c04eb8520553c1447d1cd25d2bcbb3321ee644c9bd

Observation 637da00f-17f4-4838-b577-1697b3dc2884 · outbound

This paper cites Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.754418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.754418Z digest=sha256:0ec9bd36056fb8a682d0a56963b4160cc62bfbf3096d885d1da0ecb0f7584717

Observation 7e51af2a-bd1f-4de3-bd17-ae5ac711161a · outbound

This paper cites V oxceleb2: Deep speaker recognition.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation V oxceleb2: Deep speaker recognition

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:30.957306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:01:30.740367Z digest=sha256:9cd91e99d5f20db9b0b8509ede8756eb9732f01c1bcf9df1fcab188261246496

Pith citing papers

Observation 53c7a565-32d3-4fa6-b7a2-cba71f4677f3 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.499544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:8cd7e53c2d7b1bd6f31f686a171ed1d2b7391e76c9dd8892f412c5833b4d089f

Observation 0cb89d86-de96-464c-a2f5-08274b1f2d33 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.196491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.196491Z digest=sha256:634945c8e7b9a7f8d664acc24289960020bdd4dcfb325e263da85ebdd7fb08d7

Observation 108d7ebf-e767-4e25-bd24-cec23c714294 · inbound

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner cites this paper.

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:28:40.883660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:26:21.855485Z digest=sha256:ab6f4f1ff863bd3c08624e6c8764a92f1f85df8c8be2c18c620cc4bfdc537348

Observation 665fde9b-a8a6-4d66-a681-4e7eac0bc3bb · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.876833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:d50c9b8aaa3fc37048261f4d963cf972f9747301e163e10d9f1a07a3f4189405

Observation 5db50733-5e4c-431f-b55a-f56a089308b9 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 113

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4f888e11fa45a80161cee46366d934acccc0f94c56bcc2f7f176768c7124c1ee

Observation 4b8026f9-42e5-4e45-b571-5ba2bc0952fb · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.366786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:ca3babad0ddcbff81463655be9779060c562fce0e7e03ae7678c9b08f8839dba

Observation f2f4e523-8418-4ece-a17a-6951507b5999 · inbound

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation cites this paper.

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:23.960771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:28:30.471533Z digest=sha256:ebae43fc9e5d30f9405627393b5c6515bab48d6566d7edcdd53fc8878e4a7ec5

Observation add3ca7b-f2a2-4f9d-a150-b6b9f950d996 · inbound

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing cites this paper.

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.757856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:34:32.558440Z digest=sha256:4c22b32408ef549fcd204c376d722fc6a507975de15a2b3ea2f60706e4f01a55

Observation 49a4d07d-3d86-40f9-95b8-099f10a15b9a · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.708442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:13:38.523397Z digest=sha256:fb9cce9fccd87dd6b00ae15df760d9fb751288e0fe7e579fd3b84ba887c113fd

Observation b7e421ae-0ffe-48aa-b59e-219dd8b4607d · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.999411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:27:24.078053Z digest=sha256:973cf6506a6122989bba772bc08ea96a53b6837bade98e376e3091dd8e80f7bd

Observation 482642b1-8e99-4271-acf3-86030bcb6b72 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.344060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:43:09.347118Z digest=sha256:22f80b8695fc11a89458aa5c7ddc85f93ddda306415ae7f82218b0da964aaf64

Observation 0e7788fd-13a0-44e1-ab20-7915a5a6513e · inbound

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations cites this paper.

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.194113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T05:04:31.366845Z digest=sha256:8f7bc5d3bb5ce1496738b2e0ff71e86d8a135171964cc0d36f6875cdd6364746

Observation 55e9fcb2-b64a-476c-802f-ffb5b005e91e · inbound

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars cites this paper.

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:55.445287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:52:55.445287Z digest=sha256:119ca23a0d8638c7711df971ca4e4318b5a38303d3cc937323b7eefd8b033d23

Observation d958e652-9421-4996-b66f-9ef6324edd8e · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.122397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.122397Z digest=sha256:807dc45df95bf95748fcda1a22bf44aab9b30b8abf63fd74b9aca506a205ffd3