Pith. sign in

Paper Citation Record · LEDGER

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

As of 16 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 14 inbound Pith citation observations for arXiv:2508.19320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19320 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:01:30.793099Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:45:46.122397Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.706870Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17051457-5094-405d-b920-58048bd69c7d · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation LRS3-TED: a large-scale dataset for visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.733718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.733718Z digest=sha256:da70b7cdee916d224b3dd27e22047ba792e7406143649853bd22f7dd9edbc03b

Observation 9079251d-3480-4c4e-a654-a288346b7d44 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.743448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.743448Z digest=sha256:c2f0a1ad451f9a2acac1f3df09d4a78a1ff3dfbbb6c88a3ecbe86016e1fe9152

Observation 58175350-9275-4437-beee-9c33f9f6a819 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.751310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.751310Z digest=sha256:e7fc6ff7b3aab9a070c2467ca07356f354dc57cb97b6184bf94959d44993308b

Observation 87e2ecc3-0a5f-4fa9-ba94-ac0d8c4a75e5 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.758960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.758960Z digest=sha256:cd43bcdd9ecd2a7b85e4131e9f6a805e24b968fadd176c375695cfd2a798cfc3

Observation b7dc1f40-57f5-44e4-a77f-f664e284d62b · outbound

This paper cites Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.762208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.762208Z digest=sha256:66f3b8413169a02194dbebb15574d3490081d9c5eaf09e18e21723b493c18fbf

Observation b49db34c-bbd5-4f14-8308-656af95ccf76 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.764940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.764940Z digest=sha256:d37d74927be452068dfbc7e532639fdfe7d5cbbb16a07a797072790cff850206

Observation fe4cfe71-a521-4a47-b5ef-56faa10a8e68 · outbound

This paper cites TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.768612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.768612Z digest=sha256:0d22ef609ec30526f3cee5bcfdca9f856201e206429496b1107daac9383bb22c

Observation a9928428-2bc5-4a59-b748-6472665c45af · outbound

This paper cites V oxceleb: A large-scale speaker iden- tification dataset.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation V oxceleb: A large-scale speaker iden- tification dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:30.946894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T16:01:30.772271Z digest=sha256:0818cc3bea38a51c83e8a6843558b5733fcc12ea5dd846ddcce4e86ecd471c2d

Observation 2a1348e7-2ac6-4672-8192-6cbeaeab58f0 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.778327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.778327Z digest=sha256:a1324964959339a0424d37ab1017539c80c5ec379867fd0754e326bc8285447f

Observation be07b53f-fc56-413b-b096-767400a84c2a · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MAGI-1: Autoregressive Video Generation at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.781337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.781337Z digest=sha256:60f0ff1cf51acc23a1d217b2ab7645df58908352790e4d7b44a6f20676caf8d4

Observation c3902634-2cbb-4768-91eb-efa0626f97cf · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.783917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.783917Z digest=sha256:5e78e7b8086be3165b911d3340c96af06e9b5324d18b6f7158115c1f710a685e

Observation 5abb561e-3940-4e21-9c09-452074a8ff51 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Diffusion Models Are Real-Time Game Engines

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.787197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.787197Z digest=sha256:e7a22d736e895a9ac722d77dc0671162f45d19c05575152e9247e4ccdba51c4c

Observation fb3361aa-f299-458b-9e9d-02fd7de75b16 · outbound

This paper cites ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.775017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.775017Z digest=sha256:c5c77fb22adbbca0eb6d0b99e0139b90dbf08bc837a2ea4d9b7062da34907f8e

Observation e9078eba-f632-477d-9580-59079727a280 · outbound

This paper cites Body of Her: A Preliminary Study on End-to-End Humanoid Agent.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Body of Her: A Preliminary Study on End-to-End Humanoid Agent

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.737383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.737383Z digest=sha256:bcac4a259a608e1f7aacbe715fbac82ef1f6944844159a6f6eab14688d8b708c

Observation e3481134-6850-4cd3-8ffa-015982987b44 · outbound

This paper cites MoCha: Towards Movie-Grade Talking Character Synthesis.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.790252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.790252Z digest=sha256:ed955a2c39f7a1f199d0949a6ef8025adcf5b6f5681b766e3d3dc4747cb375f8

Observation c3dc2ccd-0c20-40ec-9c59-4cc98d8737e8 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.747278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.747278Z digest=sha256:4f6047e3ba0df3ac7dca53a4a7180c44700e6c07e82914a7ac8c9356dfe74a0d

Observation 6b852dd8-9b9b-4d85-8dee-48b1a43d6c9b · outbound

This paper cites Qwen2.5 Technical Report.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Qwen2.5 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.793099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.793099Z digest=sha256:a3757c6b19986eeede5804ccbe108d7565d07fd0a3822ce9087823df598045dc

Observation 637da00f-17f4-4838-b577-1697b3dc2884 · outbound

This paper cites Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.754418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.754418Z digest=sha256:b5d0d5ff0cc81ce2eb75733d32384ed62c18199cd2ec2345d89134723490130f

Observation 7e51af2a-bd1f-4de3-bd17-ae5ac711161a · outbound

This paper cites V oxceleb2: Deep speaker recognition.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation V oxceleb2: Deep speaker recognition

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:30.957306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T16:01:30.740367Z digest=sha256:216404d90be7bc189684f328a94901e126919326471742237612b01c967ec28c

Pith citing papers

Observation 53c7a565-32d3-4fa6-b7a2-cba71f4677f3 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.499544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:75119f8a88ce38cdb8328e3623d9dc7a410f647fe12928f008800223093e2300

Observation 0cb89d86-de96-464c-a2f5-08274b1f2d33 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.196491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.196491Z digest=sha256:c6b58de695a3b67db6741c4cc4c869897b688f7be42e862e0bcacce3dbb059f7

Observation 108d7ebf-e767-4e25-bd24-cec23c714294 · inbound

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner cites this paper.

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:28:40.883660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T23:26:21.855485Z digest=sha256:441a9fe381e6dfd7ea3d184c706f791b4303387a1d5de1e30448c632f2d6f108

Observation 665fde9b-a8a6-4d66-a681-4e7eac0bc3bb · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.876833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:bda17f352324a4402380468c23662ae449a3f95b36926a35b872fb32e88626bd

Observation 5db50733-5e4c-431f-b55a-f56a089308b9 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 113

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:1aa2e1bf02ca50be133d59fbcd18a032b387dca8c2b16caa44b6151e3329466e

Observation 4b8026f9-42e5-4e45-b571-5ba2bc0952fb · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.366786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:b92bde490c1fab416a18a80955bea293e3772b4a9e2908b400052626c65da2ca

Observation f2f4e523-8418-4ece-a17a-6951507b5999 · inbound

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation cites this paper.

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:23.960771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:28:30.471533Z digest=sha256:057a8f59a5466a3bef88197f670681d65ff49db2b10d98104c0dcccce6cd98ae

Observation add3ca7b-f2a2-4f9d-a150-b6b9f950d996 · inbound

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing cites this paper.

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.757856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T11:34:32.558440Z digest=sha256:8d6d808141605e95f6b79c9fd58e8b6b25330ac1ff4c387ecda606a20deb3af6

Observation 49a4d07d-3d86-40f9-95b8-099f10a15b9a · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.708442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T00:13:38.523397Z digest=sha256:850d9926c1c7623a176933a901a91eea1d70d89dae8216bce638675e03ffb8aa

Observation b7e421ae-0ffe-48aa-b59e-219dd8b4607d · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.999411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T05:27:24.078053Z digest=sha256:afe494a5b9876d97fa456b33687df358770c12e67fd312c481719fc85efcf941

Observation 482642b1-8e99-4271-acf3-86030bcb6b72 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.344060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T09:43:09.347118Z digest=sha256:a6a4339af7ee81892713b1535704390be0881ca443a15d04565fd824abffe3d4

Observation 0e7788fd-13a0-44e1-ab20-7915a5a6513e · inbound

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations cites this paper.

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.194113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T05:04:31.366845Z digest=sha256:505383ae3d14092f96c5255fab4df75c7e003159c5bd434dcf78d4f5f3e21724

Observation 55e9fcb2-b64a-476c-802f-ffb5b005e91e · inbound

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars cites this paper.

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:55.445287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:52:55.445287Z digest=sha256:33d1466ce3a5ac7953d7ffe0687d5edf4a136875cacaf8c4db3f4565fb4a4440

Observation d958e652-9421-4996-b66f-9ef6324edd8e · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.122397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.122397Z digest=sha256:aeda881ce5149121350c4ff315ab44e63ae9caa7c2795c24bd5813a20c6efe45