Pith. sign in

Paper Citation Record · LEDGER

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

As of 15 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.22065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22065 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:10.988275Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:39:42.261377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:58:51.378798Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 734f4ac1-2d33-47ab-8fe3-40ea6d5b65de · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:16.063967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:06.783016Z digest=sha256:71e0caf94e85582d840fff20f27663a67c398400cc93be26dc907318f542809f

Observation 74d212de-7d3e-42de-bd6e-ce13f7bf98f7 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:06.817759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:06.817759Z digest=sha256:469cdfa9c4b377feab1ab3541b3911e5b8a0042f284f74c59fa2a6554753dfc5

Observation 2a41ba8e-4dd0-466a-969b-6a826e4cb14e · outbound

This paper cites VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:06.874990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:06.874990Z digest=sha256:70784720d0d440e2449636b01afa77f25d2c8b0b6339ea8aa90312be53a44981

Observation 42158464-7561-4772-8727-c5bc0af53700 · outbound

This paper cites Headgan: One-shot neural head synthesis and editing.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Headgan: One-shot neural head synthesis and editing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.876786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:06.971448Z digest=sha256:8ed1c63cf92a153abd9659ed5d1969fe3a8fab8edb25251d829f4a616c3738ac

Observation cce68330-9e08-4413-b3d1-f85e84fe076d · outbound

This paper cites Emoportraits: Emotion-enhanced multimodal one-shot head avatars.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Emoportraits: Emotion-enhanced multimodal one-shot head avatars

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.695439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:06.987881Z digest=sha256:84c526e57454f9204c6080b139da9e170288c1d1e2233b774697c151b1a95149

Observation affa1dd5-84ed-4997-8bd7-1dfa608acfaf · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Scaling rectified flow trans- formers for high-resolution image synthesis, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.566897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:07.110717Z digest=sha256:95a551cac5d589c3f29a9dced07e13396f96b75ade5d38a022d7bad7bcff618c

Observation 6879313c-bf07-46fe-860a-d286b111d2d4 · outbound

This paper cites High-fidelity and freely controllable talking head video generation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation High-fidelity and freely controllable talking head video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.354198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:07.185335Z digest=sha256:9c7dc68b73e5e8309439b66ca2134fca908681a6fad52825f801ef73f84bb660

Observation cb997797-a4cf-4486-b97a-b997fb18f8ee · outbound

This paper cites Generative adversarial networks.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Generative adversarial networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.250503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.250503Z digest=sha256:54a3bd567bd6cfffc13605db8b1d719ec01c1d0aa746f13feb5f1c6c4a2f4c0f

Observation cc17ce50-34b9-42a5-9e90-afb49ffba882 · outbound

This paper cites Densepose: Dense human pose estimation in the wild.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Densepose: Dense human pose estimation in the wild

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.146492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:07.340755Z digest=sha256:755bbff57eb8dafc71c9c1ae4beebf2bc31ec373fb1a78bf98ef9691d5257ed0

Observation 6aaf8965-6a13-4c4f-8949-6642e08a1285 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.444437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.444437Z digest=sha256:dd2aa4152dddb602e69663b49527da884256c493e5aca04601de51cc35bf064d

Observation 00df9e36-3b7a-4e9c-b392-daa4396af7f0 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.988980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:07.547096Z digest=sha256:e036a701a1ae8a6e77a6b89ed5384ab0e4c19743c07cf1a1918394015ce71e79

Observation b03dd559-c26f-4097-8167-74dedf99b2e1 · outbound

This paper cites Ltx-video: Realtime video latent diffusion, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Ltx-video: Realtime video latent diffusion, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.772701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:07.614477Z digest=sha256:be44f45629d9373a7ba4ef8f1278613aede31cee78c02d7be5ba1c298e7dbe11

Observation 8c18808e-9e06-4a0f-93ec-6ce309bbc1f5 · outbound

This paper cites Image quality metrics: Psnr vs.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Image quality metrics: Psnr vs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.704164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.704164Z digest=sha256:1e64e22aadcbf48cff3e5459e57f57ea101d582d8e74c0a09dbbc8eb18feab44

Observation 3ffd4885-155c-4e81-abc5-f638a89c64f9 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Animate anyone: Consistent and controllable image-to-video synthesis for character animation, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.591407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:07.790558Z digest=sha256:90d46511efa40dbcaf502900a1e39b548b275f72a3c464eb4dd9d9b53a463137

Observation 6cca46f8-22f5-49c0-8595-7d013c542b4c · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.893234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.893234Z digest=sha256:1b6d4e18947db7f2823594010b60cf24756cd52390328095ae2143f5b72ffc5c

Observation 68788735-1475-4339-83b2-c8416bce71c5 · outbound

This paper cites Auto-Encoding Variational Bayes.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Auto-Encoding Variational Bayes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.000441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.000441Z digest=sha256:5b9eb941aa1820a1a3abfec995b1aa645f8793564b475e1a8d351d36c457c6df

Observation 0636aea7-7fcd-4e80-ab1d-5f2b8b18daee · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.078452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.078452Z digest=sha256:65a48a8e200efcdaae2619972b1fde81bc6ed1f24b41ec21a3536c20a86a8b5e

Observation 94aeeebb-3717-4c42-a712-b7608a74bc19 · outbound

This paper cites Cyberhost: Taming audio- driven avatar diffusion model with region codebook atten- tion, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Cyberhost: Taming audio- driven avatar diffusion model with region codebook atten- tion, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.422000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:08.186136Z digest=sha256:7b6ced1467a7fcf1c9203b7cc2f5240556e1777c6252edb292e84080a798b435

Observation 8b9877f9-1d4e-4e5c-8881-6f54e769660f · outbound

This paper cites an unresolved cited work.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.256681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.256681Z digest=sha256:41e5715c4b1b5dc94ba58946cfb93a6fa945fcbca38b79252c27bd2cab044a49

Observation 03003a7f-4f5d-4b34-a0f3-b99ca0722cc2 · outbound

This paper cites Contextual gesture: Co- speech gesture video generation through context-aware ges- ture representation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Contextual gesture: Co- speech gesture video generation through context-aware ges- ture representation

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:15:11.430723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:08.352676Z digest=sha256:cf74b93e08a876ad954be28b855663f4c722fd974763175a05342977f53d18ad

Observation 914f27cf-23b6-424b-8de6-87d67fa033e7 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.419865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.419865Z digest=sha256:3a02f6220d0efad9ab9c7b4f336f405137180ecb88e0d3d00c986f054c620d0b

Observation 33cb5f90-de74-40bb-b61e-aa12aaf6b464 · outbound

This paper cites Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.240161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:08.496270Z digest=sha256:d82c1658bea178df8bb1f0dabbb10811598dd52dcefc664a7f1bad76db88815d

Observation 6414bd58-f282-4f0c-af20-0790b722467f · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Echomimicv2: Towards striking, simplified, and semi-body human animation, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.075281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:08.571863Z digest=sha256:2d6005dfe116c44792b492585020dc1dfc90484cbcb4e09f376ca5a81668dd55

Observation 9449d49d-a198-4ea6-9fd0-c1a9a8aaf336 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Scalable diffusion models with transformers, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.645607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.645607Z digest=sha256:df97f74fe68d0dd6109408a99dac65e3fec4e647b83fb8d8105d985f2fa33264

Observation 38041d61-9b3a-4bfc-848d-e573dd9856d0 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Movie Gen: A Cast of Media Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.717848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.717848Z digest=sha256:6c390e05bdbc53d60a60cc7dfa01f851b37ee380f102243d911000570d1bffd0

Observation 84f008ff-f55e-4b9d-ad11-44d7ae67f778 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation A lip sync expert is all you need for speech to lip generation in the wild

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.881768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:08.839958Z digest=sha256:4c4e0d51c4e2430b57d6a7d707d05cb90a932ddf09c0bf6c308ffa9e0098e8b8

Observation 8c08dc18-2a0d-4db9-b2f4-476d9d44ea38 · outbound

This paper cites Pirenderer: Controllable portrait image generation via semantic neural rendering.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Pirenderer: Controllable portrait image generation via semantic neural rendering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.713064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:08.925812Z digest=sha256:ca534407ab733998f180cad641f9885f8669878cca208ebf6a39796cc54fd584

Observation 5625fc53-4e99-440b-939b-e52b4f97cfe4 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation High-resolution image syn- thesis with latent diffusion models, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:09.011434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:09.011434Z digest=sha256:900e74391d1a0e66b2e99c9ff8a70dc953c65de2525343d77feae4a569804cd3

Observation 498e7418-3eac-4960-8d38-f4479fc0f4c5 · outbound

This paper cites wav2vec: Unsupervised pre-training for speech recognition, 2019.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation wav2vec: Unsupervised pre-training for speech recognition, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.496846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.072609Z digest=sha256:67f09d3d82a6477565cff242cb1a098a91478be01eaab01a50d95e55c8c221d7

Observation f5d2bb0d-953e-41fe-9e96-649ae4172440 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.329345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.172872Z digest=sha256:f7e60004a2497fad868c6fadc236cba0096f957cd16612d1b240d2110f15f4e8

Observation 8e2d2a80-799f-4a57-99fb-650e9b17cf60 · outbound

This paper cites First order motion model for image animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation First order motion model for image animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.166555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.265180Z digest=sha256:bc9ac97def3d7afb4ceafcc49acddbf246da6c3706a6e3d6215d357ed0b6a10d

Observation 49d5da74-6f0e-4b97-8f8c-0bc1656a0a9b · outbound

This paper cites Motion Representations for Ar- ticulated Animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Motion Representations for Ar- ticulated Animation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.031824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.329739Z digest=sha256:5218b58c3c8ca81cc125ae437cc15deb69da0a2da82f06b7455d0bac4c6f8d90

Observation cf84800a-afdf-476d-bcb7-4a007e66234c · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:09.426874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:09.426874Z digest=sha256:c33857fea40202103608a644af8a7f761e4b066261727db2a0f8db45894f38ba

Observation 528cedd2-1037-4c67-b9a4-764d013fc53e · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Roformer: Enhanced transformer with rotary position embedding, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.875265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.518721Z digest=sha256:4ab334a8b862e5beac44dddf7b1e8d53e3924af807aebca8730c4478225432b2

Observation be10bb11-5a2c-458e-b027-b03348b3a6a7 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:09.615461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:09.615461Z digest=sha256:706049f08e0439934a1df07dc57b0d0e37fb618b98576fc2b4f033b1489e427f

Observation 36dedce8-562c-4504-aa1a-9d56c1613bc0 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.731414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.702705Z digest=sha256:f173ab1d766744687eeb90300b6782426e208dc4e0e284496d41305c6d1389a7

Observation de828712-2de9-4f5b-b3ed-3cf877ba24f6 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.590772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.831850Z digest=sha256:3f43a0d97767c524014547ea9262da2e090880e98c48ef68c52d96834aead454

Observation bceb05c7-6e09-49cb-942b-76401c21950e · outbound

This paper cites Video generation models as world simulators.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Video generation models as world simulators

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.446574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:09.956943Z digest=sha256:fbc5d1c29b2e66fccd60158d3a7e944a35c506db10b788e4a5fda8ef9b9625e2

Observation 0954348a-96ba-465f-8671-0716054334d5 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.073906Z digest=sha256:5e1af477688ef3f91b9af71bb8008ccb73aae78fe97be36dbed5d4db323ebc43

Observation 007fc260-8475-4c0d-bb16-53c913a2e4a1 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Image quality assessment: from error visibility to structural similarity

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.154461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.154461Z digest=sha256:dbacee00eab3a9babb20d8a65d57955960ccc93ea55d68b5583c5ab28a6cd568

Observation 8c9d5c35-c01b-446f-a24b-8cbb6fee46b9 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.264492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.264492Z digest=sha256:44c72cea1f8707bfb522598571e6e97fced969f06d66e51e4b6ee5e072168802

Observation e33b25f5-35a6-477b-9fe0-d98e240dc3eb · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.341196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.341196Z digest=sha256:3b1ff7fca83618ecbfb982b7c4b110fb0823c939bb0c745c768fb9c14aa4c8f6

Observation 64b1ee34-c744-4964-a379-e8052d7693e3 · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Vasa-1: Lifelike audio-driven talking faces generated in real time

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.266086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:10.413500Z digest=sha256:c56847a61ba59da7a589cfeb97f408374e57c2f76bf1c11d25627b3e39b25f4c

Observation af5707fc-ad13-4069-9ce8-02a9563587b9 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Magicanimate: Temporally consistent human image animation using diffusion model, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.105902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:10.491551Z digest=sha256:aedea1dc14343dda0a192c97aa979df2abdef761d76e2388548a4536f0a77d01

Observation a9259f75-51ca-49e6-a6f9-8c08150e3ff9 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Cogvideox: Text-to-video diffusion models with an expert transformer, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:11.976643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:10.553677Z digest=sha256:80989449f8a2ad42c20e2d34cdbe52ce8cd97cd7d73376508a303db114aae150

Observation d0b21054-d17c-493a-adb5-5546a8766b84 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.627200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.627200Z digest=sha256:87d59df9bdb04f64e399324b1ea95c55309f4888c95f4d78b3d6abce04f96558

Observation 422cedc2-df94-4135-a8be-fff9975fc0f2 · outbound

This paper cites Metaportrait: Identity-preserving talking head gener- ation with fast personalized adaptation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Metaportrait: Identity-preserving talking head gener- ation with fast personalized adaptation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:11.807988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:10.687896Z digest=sha256:fcb94d023461c580808144f8680518b38b75277b2b9277f175f9e6f3a3b4552a

Observation 6da27382-e2ac-40e7-b179-747dc59baf85 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:11.646420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:10.780635Z digest=sha256:3b69a620220436e8fb93b978cf849c5f90716c66c7d8d4309196d2cfd8f8e98a

Observation a37699d5-0da1-4905-8fc2-53590d6e5432 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.839552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.839552Z digest=sha256:ae8fd7ba492463c3bf09100324b333b0d110d3a7f1337035be9da71414022c2d

Observation 09f95457-9aa6-4471-a486-6428f67a40ea · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Open-sora: Democratizing efficient video production for all, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.896178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.896178Z digest=sha256:89fa3aecdcb7052654fc708671dd7805e07f72b357da90aa1cebf1982c7e0941

Observation 8d7cad81-4fa1-4d7b-9da8-944d9fcc959c · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.988275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.988275Z digest=sha256:58fda0f6a5219a4fefb982c82ea99d9947667fa8eab058d1947d7d2fd5e31b40

Pith citing papers

Observation 3aecc621-db5b-43ca-abdf-97e1784feef9 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.381509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:3a757d8361bad291917eaa0b664872e5c2c6d31867ea81224378bd754b675cef

Observation e7d1445d-3efc-48f8-866b-590dff817efe · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:42.261377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:42.261377Z digest=sha256:366362fbe2cd5f04f71debade682d23d3d3ac7ed2dcdbb6ba647518a3156b032