Pith. sign in

Paper Citation Record · LEDGER

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

As of 23 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 35 inbound Pith citation observations for arXiv:2505.20156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20156 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:30.549667Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:21:12.302668Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T12:41:04.394160Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 588288d3-38ca-4fd1-9cb4-6d05b9331ed2 · outbound

This paper cites Body of Her: A Preliminary Study on End-to-End Humanoid Agent.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Body of Her: A Preliminary Study on End-to-End Humanoid Agent

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:26.629471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:26.629471Z digest=sha256:33a0757adf2f65a19f173fe9ff41435ad05d45b3cf802f39d624e5f6c9f437d9

Observation b6cfadab-7d82-42b5-b151-c18d0e345c4c · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:26.692997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:26.692997Z digest=sha256:f7852954a6b782c2c4ee256c616e550cd536a8dba66624841011f3293cdc8ae8

Observation 6286e85a-4e05-42b9-8c58-064be661215b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:26.769430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:26.769430Z digest=sha256:3ecb0f2ca5e52be3e5156af0031577a962184f035eddad4f44514b22cf73ddcb

Observation 3b72457c-bd3a-4392-922a-6f768920a565 · outbound

This paper cites Blattmann, R.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Blattmann, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:26.852513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:26.852513Z digest=sha256:c2931cae9bd92c6e747a1c0be4e44cadfdfdcd07088509c1de1e68a28ea1e422

Observation 3de5cfbe-a35b-445d-ba88-47bc8fb13b2a · outbound

This paper cites Brooks, J.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Brooks, J

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:26.939627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:26.939627Z digest=sha256:628fdd3c8b8896babae720853105e858f09c1893b304e05215e87e331874aaca

Observation 8af66ef0-013a-47a4-8c0d-b09356ec883e · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:32.378131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:27.034169Z digest=sha256:d7184dbbf247fb8d83c2793ed11b427c6be40237100ac2c779faa77c198c6f84

Observation ed5ff8a9-c3c7-4bfe-9597-dfb59813d059 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.124310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.124310Z digest=sha256:ba36dd9c27b1a90d48b6bd2d592be1131c93f399a2d2f20bb07041c1dd555107

Observation 52116a24-61ed-4d4a-a8a7-a3f2d6c644b5 · outbound

This paper cites Esser, S.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Esser, S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.208294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.208294Z digest=sha256:9ad22fceca9f98873fb388ea989867f3987625e3228b324b73a205521b529280

Observation eb2cd1cd-5658-4ae3-809a-a0bc76c14fcd · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.315202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.315202Z digest=sha256:2eec5f8b5faaf295d6f28c36072badf6a877e58604ecedc98452a44fc966e296

Observation b77c0f73-d6df-4241-b1fe-047f7242ff90 · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Photorealistic Video Generation with Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.393281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.393281Z digest=sha256:b027c3d587bc80f0d1a14f066d1b7027291baf690f46e0d0cbd5d59d8d69ef0a

Observation 09089bcf-4189-4c28-9786-1c0df99829b8 · outbound

This paper cites Heusel, H.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Heusel, H

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.485775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.485775Z digest=sha256:32462fce650220ecd2447b3fd2ed6ba7e1cb456126ae92f4ad46fad897a1d799

Observation a6cadf59-6741-4182-8f17-a123d1f33aef · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:32.218678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:27.558464Z digest=sha256:2084e762bbbe0ed57c255bc68b6d17d01de9f7c0a48d6b7b5fa809398fbe657a

Observation 28e4319b-20ee-4744-886a-73ac8278426f · outbound

This paper cites Hogue, C.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Hogue, C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:32.063655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:27.649201Z digest=sha256:f63e5d95373c0a62d232e14305d6a9890b981a0b7d1d7c86e3fc6ee6a3d16de3

Observation cab230a6-a858-44a9-905d-0067eb4dfcfd · outbound

This paper cites Huang, Y.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Huang, Y

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.757517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.757517Z digest=sha256:685572aeb408da5c5ca810db2af1a4458956f06f8e2eec76d8761d0cb05661fb

Observation 8d5fd413-4126-48a8-b2b7-899499011224 · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.838081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.838081Z digest=sha256:36936b32c2ed8ef5048711ad85e33fad834c03202b27ecd1a2f73369fc225520

Observation 75e72959-7ebf-4db8-90cb-21aa6a743f7f · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.924996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.924996Z digest=sha256:7959d1fd7beaf12f81dcd0286957db146762d2edc3b528801cf9644152e8acaa

Observation c58d42b3-0bdb-4247-b80f-6bf45c7ff7d9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:27.998790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:27.998790Z digest=sha256:8d2c4a0ddfa3b806c79c08e161c286bd6382a60a730c654a468dc88164589e75

Observation 3dc06998-993a-4eb3-a842-9d38f9c56de3 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.063276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.063276Z digest=sha256:baeb877a254b359e544698579b5971fada8bc01a1a64cd63e8c62be3d4ac6002

Observation 73ff64f8-1f63-41f8-8e3d-d2fa1a80bbf2 · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.169736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.169736Z digest=sha256:3b3f5d277073164be7651fd840bc8d244a3b43d54476996ab57fb7dbc813888f

Observation e7535f10-8dfe-4ede-9af2-d9dcacc654a6 · outbound

This paper cites CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.276147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.276147Z digest=sha256:6660ae6d14ef430804e863cb6a74203e4437c585bfc67caac737d79360a4236a

Observation 5693d6cf-d8f2-45e0-96a7-7bf224c2da88 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.347532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.347532Z digest=sha256:4d128b228db15941dc2a5c8e3618a131ec409b1857e709b4898769502186d4e7

Observation 9ea3e67c-bddb-4b7c-bd6d-aaa1480bdf70 · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.454820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.454820Z digest=sha256:abd7a7894b6f6d9997c653dc130c52f09dd7a70badc0f43ff7fd7fd03baa7ac6

Observation 00f4672b-2bf2-42dc-a219-f1e0be9861d7 · outbound

This paper cites Flow Matching for Generative Modeling.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.536605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.536605Z digest=sha256:b2c75d82a1dd3918071e4dd8a6f15846abc8a55f0ec1f924b8c88d9e8b20f4dc

Observation f0d9104a-09bb-43ca-96a8-419218a2bb81 · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.618503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.618503Z digest=sha256:600c44cc21b03ae1b7e97fa8d77c2962d768620dd5e0fcb60707dbbec2160a53

Observation e229cb43-5360-4b04-ab56-fc688c2234fa · outbound

This paper cites Radford, J.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Radford, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:31.907480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:28.712294Z digest=sha256:000c2ce3dfbbc13c744bee692819781b3cfdeafcdda0c7016305c8decd508632

Observation efc2c1e2-cc3e-4aae-80fc-46a7f7f914b0 · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:31.734865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:28.794453Z digest=sha256:3c30a63ecb6355fc476a0f7e72f34011c00507eb754e8fde3ddd11cb0dc52b08

Observation 77c53fce-1cf1-4c31-a5fa-7b30f13d5543 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.903520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.903520Z digest=sha256:c7f59d98c2e2934601ad64a685ae1aa0178014168b7e97c4c85455ed1089ed76

Observation 9d4b4b41-c505-4e2a-8ab3-ba265963146c · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:28.983264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:28.983264Z digest=sha256:2220a94002c542f596cffadced2ded39f80a8b80bd6233050d28f73bcae9041b

Observation 3e9ae884-1849-4e20-908b-ebf2eb80769b · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.076310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.076310Z digest=sha256:904141c84d5a77e0f4aacb1d23a6c297f26f47aa4d691c81e53643868e5a1749

Observation 2e5e23a9-7cc2-41a1-8fa0-39c5b24cb736 · outbound

This paper cites Unterthiner, S.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unterthiner, S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:31.589649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:29.161941Z digest=sha256:6a01abf48ba1d25e0c8de0374dcf272270e5c62336a82d51a0e840730da415e6

Observation 4aa870e9-0003-4f2a-a7c5-f004ed7116f8 · outbound

This paper cites Villegas, M.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Villegas, M

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.248644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.248644Z digest=sha256:f98c84e73f7bc72a2ca65a7307dc65bc754a72850a2ad945342e6e4690ff55e3

Observation 02778bfb-b2fb-44a1-828c-6c2bef5cc689 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Wan: Open and Advanced Large-Scale Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.329734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.329734Z digest=sha256:cb3100f557deb6f9ce9ff4b28c5b80093505c8843c4c3a093671899e98b26f1d

Observation b23639c4-e012-4a26-938e-9e2c4ca63775 · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.417057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.417057Z digest=sha256:3dfe2f62aac1f8b534f5254ad446e3a572791e554603300b8733e5f888fd2e04

Observation abd11fd9-bd30-4c68-978f-86c3d5eafdac · outbound

This paper cites ModelScope Text-to-Video Technical Report.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters ModelScope Text-to-Video Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.525633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.525633Z digest=sha256:82b0a2b49af00b79e32cca5f7337c4f2e9a1143857fe51d3d6e2f0289acbb07c

Observation 23562447-b0cf-4b97-8e8b-76addc290927 · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.621663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.621663Z digest=sha256:0a36bae1db0888b29fe56aca6e0a98966ea33d3782dd8b884ca0ccd5965236cc

Observation 27e88653-d264-4b8e-b612-726e40321e3a · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.718723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.718723Z digest=sha256:f864ac4a8e4f8d0beb2ad127417e94937c51d7465e87fc9d17d77abc559e607c

Observation a26cb2eb-a6fe-430e-b3dc-578da54c9208 · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.798615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.798615Z digest=sha256:1c370779c3f37d55c53c02dba0354df3cdf0a70de01e4efeb5b67a91836469ac

Observation 9566cd20-ce5f-4161-8dfa-e9982b076577 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:29.947862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:29.947862Z digest=sha256:2df682d06df697b111894d612dcdb656864532e60731e29a7c1dafb5b55339c5

Observation 0bc3d107-8172-44b4-a7e9-6efaf9a46f08 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:30.045227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:30.045227Z digest=sha256:f403f9ee37dca1609b0da6c2db8ba3ff234866086cf80dd4fc9f08721309d64b

Observation 5f3d60eb-7299-4daa-b8d4-5cc7dcb8d36f · outbound

This paper cites Zhang, X.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Zhang, X

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:31.428928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:30.138964Z digest=sha256:f044a6535b11c1fe600241a2069b82346d2028a5698424a3454dc616b9e0b7a6

Observation a169bb99-bf0f-45df-846d-0df64eec789f · outbound

This paper cites Zhang, L.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Zhang, L

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:31.278250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:30.232211Z digest=sha256:ed2a90e23f2df672be0245c5711fcbc5db67d9b2ca7cb223a867cb89c7cb9838

Observation 86386a17-2dac-4eb0-bcff-7dc885138102 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:30.356365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:30.356365Z digest=sha256:432c915e490bcd64a6c3d594c9f3158d46ae79bc528acf427381d5bf11db6782

Observation 674d369a-5816-41c1-9ae9-3c02adad90f1 · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:30.456640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:30.456640Z digest=sha256:9db5fdee0e358bea6eb7fc0e6a7004cb350d40383831b4f8e6d3c24ac28d5443

Observation bf5a2aee-6659-466c-b602-cb8ba9d90e6d · outbound

This paper cites an unresolved cited work.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:31.116744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:30.549667Z digest=sha256:ab6264302af4162345c80a9f193d167fef333b6dbaea4c768f798848a14a8882

Pith citing papers

Observation bbce61ed-b268-474d-b389-49af8bed024c · inbound

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router cites this paper.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.897824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.897824Z digest=sha256:4c5589066db9b0bf11e31861206746edad97059060e306bbd0cd5f512057c579

Observation 79609944-7dee-4bd5-aeee-e6dcfefb663d · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.893203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.893203Z digest=sha256:d7b74833e679a837ec071a9c22be8bd13c2d8e350f5e5c7a3fbd408039abcaa1

Observation 0119b901-ac10-44ef-9ef1-f8f5ee3bdc63 · inbound

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation cites this paper.

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:42.835896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:42.835896Z digest=sha256:26d4c884d8e68176a55185b4b851921730428d6eb7754f81b749337adc27cd1a

Observation 0301e8b9-c0db-4444-bf66-3fabdaf38afe · inbound

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation cites this paper.

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:42:41.019438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:42:41.019438Z digest=sha256:e0249bc6fa8b8df690ad7f6a373d3ffe71e64524de2c5c543ee01d33c51c5dfb

Observation c1e1bfbd-c25a-43d2-95b4-104ee8d9c7c3 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:45.351100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:45.351100Z digest=sha256:ebe5e8e65c9f066f471186021647c0db384abf89a8de89b79597da147280632f

Observation 4e38b2e0-a4b8-4d3e-8f2d-23f112e447c1 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.212957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.212957Z digest=sha256:47f9eb6ab792869c921f5e076ed0ee3c05fddef52963c349524e0a881b4ba0e1

Observation 3e091873-51f3-4079-b666-ccbce3d342bf · inbound

Wan-S2V: Audio-Driven Cinematic Video Generation cites this paper.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.625126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.625126Z digest=sha256:9f578a79b5d188cb22476631da58b2693ea337b3112cfb7cabcdb607c562e94d

Observation 69e9fe5b-c627-4661-b3b4-3e72b6611d20 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.088684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.088684Z digest=sha256:cf553c1c899dc0b329f50b2ad0294a18f265dda683578598767a79549021a20c

Observation e5cf9909-4892-42ea-8df7-cc66e9a7d1da · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.280317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.280317Z digest=sha256:2598fe8bbc0be9b5ba34ca18d23aae053d2eafb1b185b5704b108922a9c79be5

Observation 2818ee41-0988-4952-81f6-d2edff52087b · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.145826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.145826Z digest=sha256:51c3a857f87ee3218dc4f191895a090234560185b36a13368dad685f95dab4e0

Observation 15816203-5205-4098-927d-b2ccc98fb5e1 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.405659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:baa06d788e151a35f2b1340a591c2b61e423696607696dc07e733d666cf9a6b9

Observation 48f3ba57-29a5-4ed1-8edf-3887bff17dad · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.617276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:b50f7667bdaa2efd3413475d8593e4db2e61ccda77e87106cbe4fef87d436d59

Observation a70e2c08-2720-43c8-ab53-4ee090b879ef · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.831890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:fe51fcdea9142ae77958704040235a919d167b74d3e1d48e22a5b017a12f7277

Observation ce62eea3-9924-4266-8bff-633d54a0591f · inbound

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation cites this paper.

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:20:54.957167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:19:49.671136Z digest=sha256:d6aecfa790d83a5a6c36dbd94525ac9101319ddec778325efdb34fea8495355b

Observation 16c1e838-08b9-4157-b9ce-21d3e0227c44 · inbound

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation cites this paper.

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:41:04.395475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-05T12:34:41.758238Z digest=sha256:f488f4477c11f27ac8bd49c6b2172b5692b2081e557bf4a64665cf6e357b94fc

Observation 45966fcc-bf53-4162-a8f3-8c566350d129 · inbound

From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation cites this paper.

From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.993284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T05:12:55.062360Z digest=sha256:38d9ebfe12f2a5430afc216ca36597340cb25328d86251eb3f9098df8b331982

Observation 9b87c130-bc3b-4952-be35-5a430522ed65 · inbound

From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation cites this paper.

From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:56:14.940154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:55:13.370339Z digest=sha256:a5083faf98c9913de38e551333fe5c0902b41a7fac203a9087dac79ad70e67bd

Observation cd038643-1325-41d9-97b7-c5b75ad4e314 · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.817723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:cb88ea12ab135b2620e3c979b0f9bfc52cb5c97228927dc039690d11d732bd8c

Observation 7254af00-a8c2-424c-83b9-bb3452c290ad · inbound

Advancing Reliable Synthetic Video Detection: Insights from the SAFE Challenge cites this paper.

Advancing Reliable Synthetic Video Detection: Insights from the SAFE Challenge HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:57.489097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T00:53:18.735506Z digest=sha256:4807978b6629e67d66bffa80d7d8470f9b511a2697ec31d1d5be0517f0ae14d1

Observation 2840d6b3-c54c-4c46-9b33-871a4055e74e · inbound

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation cites this paper.

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:23.980846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:28:30.471533Z digest=sha256:a21f3611c8e0502aef9eeeaad80787d0626cdbc8d88eeaf788741d4f3d89685f

Observation 0f1ac7a1-5ba9-4776-9c18-c41a47b34b47 · inbound

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration cites this paper.

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:25:04.977138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:21:58.123630Z digest=sha256:ef22a96dac0b881d00bef72e7421bba905ef71e0f4a104627d808470fbad2e60

Observation 6993e1d7-5a78-4019-a386-b96d18e52023 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.087991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:bf7e72f2b81923cb302ee7bb2a27fe099b2e8b1ef4a7e11ece7b7461710ad998

Observation 7f083a95-514a-49f4-84fa-b92b464bbf32 · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.666638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:3050fc783da97a4c4a875336f8b42be8cb221f3389a35c7ac4b7f675ce75ed3a

Observation de7144ba-08db-4dbb-80bb-e66e8427185d · inbound

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation cites this paper.

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.133861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T08:30:00.438334Z digest=sha256:d7cabdf2c7f59636c319681369d0d7d52da1edafe3de7958e0fb93cab9ea60d4

Observation f9769dbf-ff53-41ef-9735-3571dab8ec85 · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.578759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T09:09:06.925645Z digest=sha256:bf47b5abdc137deeb10cb4a896428f6853fc5228ffa17ed8531f79dce8c2156f

Observation ba3ef183-f07e-40c1-984b-5a75ef2ad9d3 · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:29.096627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T07:00:53.496569Z digest=sha256:5e039119b8345e46a71fc5c1eb0b5761362bcf155de2b492e87fdcaf56a45cc1

Observation 99bb1fc8-4679-4fe4-b37a-b62d049c71b9 · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.333810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:26e9fcddf056cc4893ae155d5dcd33741cd40629aafcb164c8ce0d2eac2f34b6

Observation 9ddd5501-accf-41aa-b8e0-34c74969c051 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:15.550477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:15.550477Z digest=sha256:e4a5eef6e05e75e2fa15e1168d72a27a74866ce98b4bc0b252a66f3ce7782490

Observation 5debf619-80d7-4f30-b701-caf13a5f782e · inbound

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars cites this paper.

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:55.409252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:52:55.409252Z digest=sha256:fe3058cbb250e814816e627898e3514c136c677d794252edddb65d7ade6b6c6e

Observation bdcf06bc-8e18-4506-977f-6f1d8ac4919a · inbound

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars cites this paper.

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:59.579082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:18:59.579082Z digest=sha256:d3e8409d903183b795d193446900679da758412909dbfcb7790a857740f5cd17

Observation 408e0bc6-58b0-4827-b1be-96613547c5e7 · inbound

CoT-Edit: Let CoT Guide Instruction Video Editing cites this paper.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.982238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.982238Z digest=sha256:91c01eb0a7693396a04c4be63d8e5f8b6d58944349cb48f865d92ddf882498a1

Observation c3a629d6-5bce-4156-9c88-8156fd355e42 · inbound

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos cites this paper.

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:25.258190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:25.258190Z digest=sha256:ac4a4a961aec1cba4040dd04281dad50eef6ec7141171a8dde6328e9653ad585

Observation b1bd9306-469f-4ce9-90b3-eab397e300ab · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.240639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.240639Z digest=sha256:a69fb6d84b628de8bbfa0455afdedcaee375b47a6853b892e9966a7f19ff00bb

Observation f59e2f68-2862-46a1-b819-776d994b7e5e · inbound

Foundation Models are Implicit Deepfake Detectors cites this paper.

Foundation Models are Implicit Deepfake Detectors HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:41:09.275832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:41:09.275832Z digest=sha256:351d619436539f17b1e0cf3c914f54a894d205590b9f3383f8de1b8e04c136e7

Observation 08990fbb-5c71-46f1-bfd7-354b05afb300 · inbound

Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars cites this paper.

Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:21:12.302668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:21:12.302668Z digest=sha256:1213371d5ad3f6487a44b2924cc118c009245b1aa0efac4f835e86f42e39de92