Pith. sign in

Paper Citation Record · LEDGER

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

As of 14 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.22065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22065 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:10.988275Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:39:42.261377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:58:51.378798Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 734f4ac1-2d33-47ab-8fe3-40ea6d5b65de · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:16.063967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:06.783016Z digest=sha256:2d8b5af98d9fc6b2adfceba7fe99777820e142e860b035634494d54872656b90

Observation 74d212de-7d3e-42de-bd6e-ce13f7bf98f7 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:06.817759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:06.817759Z digest=sha256:469cdfa9c4b377feab1ab3541b3911e5b8a0042f284f74c59fa2a6554753dfc5

Observation 2a41ba8e-4dd0-466a-969b-6a826e4cb14e · outbound

This paper cites VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:06.874990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:06.874990Z digest=sha256:70784720d0d440e2449636b01afa77f25d2c8b0b6339ea8aa90312be53a44981

Observation 42158464-7561-4772-8727-c5bc0af53700 · outbound

This paper cites Headgan: One-shot neural head synthesis and editing.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Headgan: One-shot neural head synthesis and editing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.876786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:06.971448Z digest=sha256:b99e8a3e62c3060a0dce4d36e1e54587d8e46df7f9d9527a83280e5dc9d3e600

Observation cce68330-9e08-4413-b3d1-f85e84fe076d · outbound

This paper cites Emoportraits: Emotion-enhanced multimodal one-shot head avatars.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Emoportraits: Emotion-enhanced multimodal one-shot head avatars

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.695439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:06.987881Z digest=sha256:0b3add9c4542b84794492ec3c71e2ab3cd5418ee98a175f57652082cf49d644a

Observation affa1dd5-84ed-4997-8bd7-1dfa608acfaf · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Scaling rectified flow trans- formers for high-resolution image synthesis, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.566897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:07.110717Z digest=sha256:fdb132aa4def23877083332dd67ef340f6f5a5572b27eec39bc0ab43907b7c20

Observation 6879313c-bf07-46fe-860a-d286b111d2d4 · outbound

This paper cites High-fidelity and freely controllable talking head video generation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation High-fidelity and freely controllable talking head video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.354198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:07.185335Z digest=sha256:fb55d6499d3176dea0b3e193f9964090e1baeafaf06da10dec6324c95c4f51a3

Observation cb997797-a4cf-4486-b97a-b997fb18f8ee · outbound

This paper cites Generative adversarial networks.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Generative adversarial networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.250503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.250503Z digest=sha256:54a3bd567bd6cfffc13605db8b1d719ec01c1d0aa746f13feb5f1c6c4a2f4c0f

Observation cc17ce50-34b9-42a5-9e90-afb49ffba882 · outbound

This paper cites Densepose: Dense human pose estimation in the wild.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Densepose: Dense human pose estimation in the wild

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:15.146492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:07.340755Z digest=sha256:3e704fbd799acbb402f01bfaffbb95c23f9a3c4aca23af0b369fef859889362d

Observation 6aaf8965-6a13-4c4f-8949-6642e08a1285 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.444437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.444437Z digest=sha256:f071ed034ca971fe8778df2a77a982e342c4b4d01bf7ddaf71e71d887f628bc9

Observation 00df9e36-3b7a-4e9c-b392-daa4396af7f0 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.988980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:07.547096Z digest=sha256:55f3de838244ae04895046dfbd75264a837eaf2a3f8625d3fa37d5365d671078

Observation b03dd559-c26f-4097-8167-74dedf99b2e1 · outbound

This paper cites Ltx-video: Realtime video latent diffusion, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Ltx-video: Realtime video latent diffusion, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.772701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:07.614477Z digest=sha256:25190a6158e7684e8c75722b41f788b9c49f58fa6d120ac476f72912ddab2f5f

Observation 8c18808e-9e06-4a0f-93ec-6ce309bbc1f5 · outbound

This paper cites Image quality metrics: Psnr vs.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Image quality metrics: Psnr vs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.704164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.704164Z digest=sha256:1e64e22aadcbf48cff3e5459e57f57ea101d582d8e74c0a09dbbc8eb18feab44

Observation 3ffd4885-155c-4e81-abc5-f638a89c64f9 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Animate anyone: Consistent and controllable image-to-video synthesis for character animation, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.591407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:07.790558Z digest=sha256:6eb5394ccf41736f331c154f9d2869bcf2c5e7a6baae589e410464e1dfd4ddf9

Observation 6cca46f8-22f5-49c0-8595-7d013c542b4c · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:07.893234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:07.893234Z digest=sha256:1b6d4e18947db7f2823594010b60cf24756cd52390328095ae2143f5b72ffc5c

Observation 68788735-1475-4339-83b2-c8416bce71c5 · outbound

This paper cites Auto-Encoding Variational Bayes.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Auto-Encoding Variational Bayes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.000441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.000441Z digest=sha256:33fbe4b40ac35e5ab5339d67b49c5ecdf3819e56f8e9674826ac588c1a4e00c7

Observation 0636aea7-7fcd-4e80-ab1d-5f2b8b18daee · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.078452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.078452Z digest=sha256:65a48a8e200efcdaae2619972b1fde81bc6ed1f24b41ec21a3536c20a86a8b5e

Observation 94aeeebb-3717-4c42-a712-b7608a74bc19 · outbound

This paper cites Cyberhost: Taming audio- driven avatar diffusion model with region codebook atten- tion, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Cyberhost: Taming audio- driven avatar diffusion model with region codebook atten- tion, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.422000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:08.186136Z digest=sha256:f36ce937ccdaf733cb146cbc71d0ab84424a31d4c3b64776b84c2265579f076e

Observation 8b9877f9-1d4e-4e5c-8881-6f54e769660f · outbound

This paper cites an unresolved cited work.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.256681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.256681Z digest=sha256:41e5715c4b1b5dc94ba58946cfb93a6fa945fcbca38b79252c27bd2cab044a49

Observation 03003a7f-4f5d-4b34-a0f3-b99ca0722cc2 · outbound

This paper cites Contextual gesture: Co- speech gesture video generation through context-aware ges- ture representation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Contextual gesture: Co- speech gesture video generation through context-aware ges- ture representation

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:15:11.430723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:08.352676Z digest=sha256:77cf1111b23d336651fa32611983fe7d1c189bb79d1e1b475e675da95f2ca620

Observation 914f27cf-23b6-424b-8de6-87d67fa033e7 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.419865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.419865Z digest=sha256:3a02f6220d0efad9ab9c7b4f336f405137180ecb88e0d3d00c986f054c620d0b

Observation 33cb5f90-de74-40bb-b61e-aa12aaf6b464 · outbound

This paper cites Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.240161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:08.496270Z digest=sha256:002700e29b3721e61cced4e0b16ff730a3bf7de489e6ed1cdf4db8e712fa3a8a

Observation 6414bd58-f282-4f0c-af20-0790b722467f · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Echomimicv2: Towards striking, simplified, and semi-body human animation, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:14.075281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:08.571863Z digest=sha256:383466f6c333f4a8611675825d28e2bfd97fa49a408a9346782319dd21ed3a95

Observation 9449d49d-a198-4ea6-9fd0-c1a9a8aaf336 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Scalable diffusion models with transformers, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.645607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.645607Z digest=sha256:df97f74fe68d0dd6109408a99dac65e3fec4e647b83fb8d8105d985f2fa33264

Observation 38041d61-9b3a-4bfc-848d-e573dd9856d0 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Movie Gen: A Cast of Media Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.717848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.717848Z digest=sha256:6c390e05bdbc53d60a60cc7dfa01f851b37ee380f102243d911000570d1bffd0

Observation 84f008ff-f55e-4b9d-ad11-44d7ae67f778 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation A lip sync expert is all you need for speech to lip generation in the wild

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.881768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:08.839958Z digest=sha256:278e534b76d44f9563bcca16358bcc430c44257e19b3f9f87e5b72db708562b7

Observation 8c08dc18-2a0d-4db9-b2f4-476d9d44ea38 · outbound

This paper cites Pirenderer: Controllable portrait image generation via semantic neural rendering.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Pirenderer: Controllable portrait image generation via semantic neural rendering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.713064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:08.925812Z digest=sha256:75be8dddefe2734d839dcef97b04d624e04cc6f5131b12fff6e4eab62dc7f123

Observation 5625fc53-4e99-440b-939b-e52b4f97cfe4 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation High-resolution image syn- thesis with latent diffusion models, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:09.011434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:09.011434Z digest=sha256:900e74391d1a0e66b2e99c9ff8a70dc953c65de2525343d77feae4a569804cd3

Observation 498e7418-3eac-4960-8d38-f4479fc0f4c5 · outbound

This paper cites wav2vec: Unsupervised pre-training for speech recognition, 2019.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation wav2vec: Unsupervised pre-training for speech recognition, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.496846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.072609Z digest=sha256:c0a24850f35bd1240f20d13489076b2e304390594d5f92c01d1563c8fbc59037

Observation f5d2bb0d-953e-41fe-9e96-649ae4172440 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.329345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.172872Z digest=sha256:014c707a3e6542a42714696cdc9289d9d7843b1c78504ec3331510efe4d872ca

Observation 8e2d2a80-799f-4a57-99fb-650e9b17cf60 · outbound

This paper cites First order motion model for image animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation First order motion model for image animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.166555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.265180Z digest=sha256:695d4fba55d979b1740d5afed9f7a375f4d02531e8ffd0629f106d1b944ff9d7

Observation 49d5da74-6f0e-4b97-8f8c-0bc1656a0a9b · outbound

This paper cites Motion Representations for Ar- ticulated Animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Motion Representations for Ar- ticulated Animation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:13.031824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.329739Z digest=sha256:bb4e78a0613c76ca4465593db3723c247767672fece2dd16020f8c6355236130

Observation cf84800a-afdf-476d-bcb7-4a007e66234c · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:09.426874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:09.426874Z digest=sha256:c33857fea40202103608a644af8a7f761e4b066261727db2a0f8db45894f38ba

Observation 528cedd2-1037-4c67-b9a4-764d013fc53e · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Roformer: Enhanced transformer with rotary position embedding, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.875265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.518721Z digest=sha256:d544775dbe6748c8d8143bf0ec0bb0218591512a4da26c7b6c275bb17107514b

Observation be10bb11-5a2c-458e-b027-b03348b3a6a7 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:09.615461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:09.615461Z digest=sha256:36f7d4964ea51e40d5d9820219fac21a91afd5335923b8dd7a17bac287efb717

Observation 36dedce8-562c-4504-aa1a-9d56c1613bc0 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.731414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.702705Z digest=sha256:0b0944e099d39621278620b2169c697f477ba670b44da9aac8bf60a644fde740

Observation de828712-2de9-4f5b-b3ed-3cf877ba24f6 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.590772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.831850Z digest=sha256:e279f039e0f0569ccc8eb59bf2fd0b3a35e772b5cf5827ecdc89435d2c8c7390

Observation bceb05c7-6e09-49cb-942b-76401c21950e · outbound

This paper cites Video generation models as world simulators.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Video generation models as world simulators

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.446574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:09.956943Z digest=sha256:1354a6dc53c95ea4b02166cc27d6d0a79564287bf83d0c23a7ffec608f018cad

Observation 0954348a-96ba-465f-8671-0716054334d5 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.073906Z digest=sha256:33a202f0dffb2efdf71be7af632c2967613b332a318609266f65e62b50928dd1

Observation 007fc260-8475-4c0d-bb16-53c913a2e4a1 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Image quality assessment: from error visibility to structural similarity

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.154461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.154461Z digest=sha256:dbacee00eab3a9babb20d8a65d57955960ccc93ea55d68b5583c5ab28a6cd568

Observation 8c9d5c35-c01b-446f-a24b-8cbb6fee46b9 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.264492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.264492Z digest=sha256:44c72cea1f8707bfb522598571e6e97fced969f06d66e51e4b6ee5e072168802

Observation e33b25f5-35a6-477b-9fe0-d98e240dc3eb · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.341196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.341196Z digest=sha256:3b1ff7fca83618ecbfb982b7c4b110fb0823c939bb0c745c768fb9c14aa4c8f6

Observation 64b1ee34-c744-4964-a379-e8052d7693e3 · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Vasa-1: Lifelike audio-driven talking faces generated in real time

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.266086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:10.413500Z digest=sha256:74db183457618954c4e639ceaa87216ab959f9e65c75cea2c7df2e11f071f824

Observation af5707fc-ad13-4069-9ce8-02a9563587b9 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model, 2023.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Magicanimate: Temporally consistent human image animation using diffusion model, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:12.105902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:10.491551Z digest=sha256:9be046d136e9c9af597282989512722835d68c5715823270b738c090581e1e29

Observation a9259f75-51ca-49e6-a6f9-8c08150e3ff9 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Cogvideox: Text-to-video diffusion models with an expert transformer, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:11.976643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:10.553677Z digest=sha256:e987eaa8ec29cbe0242e6eaa125a0b1e02e15210ccb5d265608ae544150c4a52

Observation d0b21054-d17c-493a-adb5-5546a8766b84 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.627200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.627200Z digest=sha256:87d59df9bdb04f64e399324b1ea95c55309f4888c95f4d78b3d6abce04f96558

Observation 422cedc2-df94-4135-a8be-fff9975fc0f2 · outbound

This paper cites Metaportrait: Identity-preserving talking head gener- ation with fast personalized adaptation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Metaportrait: Identity-preserving talking head gener- ation with fast personalized adaptation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:11.807988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:10.687896Z digest=sha256:98d16990833cce6cee3c80bedfa96ecf0933b795f9836b92225b4b04e01ab7b7

Observation 6da27382-e2ac-40e7-b179-747dc59baf85 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:11.646420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:15:10.780635Z digest=sha256:7522ca87c46469a770b44e110feaf2771b8fe4b2cf20d0a42b5c7ff82a5f3c73

Observation a37699d5-0da1-4905-8fc2-53590d6e5432 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.839552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.839552Z digest=sha256:ae8fd7ba492463c3bf09100324b333b0d110d3a7f1337035be9da71414022c2d

Observation 09f95457-9aa6-4471-a486-6428f67a40ea · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Open-sora: Democratizing efficient video production for all, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.896178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.896178Z digest=sha256:89fa3aecdcb7052654fc708671dd7805e07f72b357da90aa1cebf1982c7e0941

Observation 8d7cad81-4fa1-4d7b-9da8-944d9fcc959c · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.988275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.988275Z digest=sha256:58fda0f6a5219a4fefb982c82ea99d9947667fa8eab058d1947d7d2fd5e31b40

Pith citing papers

Observation 3aecc621-db5b-43ca-abdf-97e1784feef9 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.381509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:81cce576d82c4204562314b033b5e654ddd87c954468863324c8efaee9018532

Observation e7d1445d-3efc-48f8-866b-590dff817efe · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:42.261377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:42.261377Z digest=sha256:366362fbe2cd5f04f71debade682d23d3d3ac7ed2dcdbb6ba647518a3156b032