Pith. sign in

Paper Citation Record · LEDGER

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2508.08891.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08891 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:24:08.721929Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved27
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de2bf21f-4036-459e-8002-dcefadac5b4b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.359067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.359067Z digest=sha256:9415e7fc61672997f68f05c1555b1cf76d06cc029507809b05ddc0f6ebc78f36

Observation f01b6fc9-2761-4af7-8c24-c0df7620a36c · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Emerg- ing properties in self-supervised vision transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.426808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.426808Z digest=sha256:8322e7b983ca525cb90610bae3a62492d05b967c1c1a9392694ccbe1169cbef7

Observation 66a2cb35-9ef2-4aa1-990f-4f17cadf014f · outbound

This paper cites Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.525128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.525128Z digest=sha256:91459695d0817b05074e648a85a9d771504f7bab3800edf0026b90cfc7bfdb19

Observation 9d7248bf-9fae-4233-a7e6-de05a57562a0 · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.595256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.595256Z digest=sha256:cb04ecc5e470c4b147164d6ca7de7b3d56ee6843994ca47a6862386d635eaec6

Observation d160ab63-8d60-43af-b7dd-56f2ed06878c · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.723175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.723175Z digest=sha256:b33b783f1e759b55bc249faa5c8e8eb059c31e3466895afae699834ef905ea30

Observation d42da7a8-dbf2-4411-9906-0afc46933a68 · outbound

This paper cites Gemini 2.5 pro: Our most intelligent ai model, 2025.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Gemini 2.5 pro: Our most intelligent ai model, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.556123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:05.819589Z digest=sha256:191679c9df40a546d11c6ae0ff6be8e5fb57f95e73309cfd77618f7db6b73bef

Observation dd4d3973-3b7d-4bba-98d6-54fc4a4c9f97 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.416785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:05.908061Z digest=sha256:8d5280ad708fd229644b2c8bd38c63389eeb49a7e1df629dab7482c324c2d8cd

Observation 02cab4ec-693a-4221-b10c-8befa4cd10b2 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.035484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.035484Z digest=sha256:2d211dbac52c8f75873c8849ab954a05a50a59b7688f3535d069a35948bf61d8

Observation 34596d63-1629-49b9-8b82-950be13908ba · outbound

This paper cites Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.127789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.127789Z digest=sha256:0226a09cd67dec9ff285a291e7fac17fe5877ecfbf4f7f1ba2edc3dd0d70b4f0

Observation 5fa83f4b-3bcc-4a19-bffc-ba006d6dbe76 · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.186484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.186484Z digest=sha256:7f2a6e351ee86bc8ddf465a6f18bf4fcd55772cafe53399216861b184054c5df

Observation 6471fbcf-7a4d-4bf6-921e-b3a5270daf23 · outbound

This paper cites The accu- racy of psnr in predicting video quality for different video scenes and frame rates.Telecommunication systems, 49:35– 48, 2012.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos The accu- racy of psnr in predicting video quality for different video scenes and frame rates.Telecommunication systems, 49:35– 48, 2012

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.226142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:06.279489Z digest=sha256:b24a7e6804ff15bcd0365f15aef27f44255867a4ff0ac2046093a2a594cc7812

Observation e54293de-0b29-48d9-924f-b7446fb7133d · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.340492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.340492Z digest=sha256:10299ad157c837d34ae9517e38e193ab9c1988b1db162add9fe5457942a91e6f

Observation 80926fae-08ce-4f87-9008-0757494756b5 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.431313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.431313Z digest=sha256:2141f2a43e45e64cd2776ce22e5a37813fce21fd59579f492c41d071438572d9

Observation 893ea3b1-3f2f-4b33-880b-bca847ba4a81 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Musiq: Multi-scale image quality transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.048008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:06.551175Z digest=sha256:bfab702079cb30b73fc2fa93d10a99fce1dbb8335ef2ff257b5a537d683202eb

Observation 994d9fab-d6d1-4b4d-8c98-01ac622ec4b9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.688755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.688755Z digest=sha256:90dfa6fac6c3b9ab42750f818b36c3276a6f4e67efe03894e5906088553699e0

Observation 27d6c694-364a-4966-928b-b35d9c2e7396 · outbound

This paper cites Luma ray 2 video model.https : / / lumalabs.ai/ray.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Luma ray 2 video model.https : / / lumalabs.ai/ray

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.854778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:06.807169Z digest=sha256:2ddc3bc3e00804de436c4281663f85aaf6c8fa03eef3468af8754df1876050e1

Observation 8bcc75af-195b-48d1-bd86-58844b0ad169 · outbound

This paper cites Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.700967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:06.887819Z digest=sha256:b7d04d1cdce4c92af404925573927b2854a2e51e6bd39332bbe0ef2a3b3d1327

Observation 98cb3211-2ac2-4a1d-8998-c10475b07b4b · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.088948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.088948Z digest=sha256:0232de6ea6915e44146d52001078750671be07e44581c77830c06e968f079cea

Observation 57901d56-65ea-4a19-9d9f-23e6f2a6ce66 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.155337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.155337Z digest=sha256:1b5d354783aa93b28f134631d5de5bad66ad19f29580c01b0ddaae3b4e36c50d

Observation 57d2b5d0-0628-4f84-b681-3eafd24e774d · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.374361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:07.257087Z digest=sha256:0cdc6768f6e8f4c5e3e45751b9b0b6ed8ff7874387869904f9d6b553266c7a82

Observation 8685b4a7-33cc-4bf9-98ef-1e5cffc6b000 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.313868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.313868Z digest=sha256:55d2a94a5d9c9a4af538f1f7f86e4b47eb434a6b739cf2bed8ddd0930636eec7

Observation c88e2b08-679f-473c-a951-557d39ae9b46 · outbound

This paper cites StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:24:09.334830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:07.379667Z digest=sha256:f1d03e8809a391553709ed7668e5ea16067e5c37dcd1d9d589cc9e9d71a02b2b

Observation 9da70060-4f55-40ab-a22f-0f3d457aee67 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.462449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.462449Z digest=sha256:1b55abbae4a585ecf3a9770a37dd014b5591dccfebce55f3e08a983c9af5e8bc

Observation b7948f18-721e-4235-828d-3ad5a26b8791 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos A lip sync expert is all you need for speech to lip generation in the wild

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.148523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:07.587910Z digest=sha256:918645b6da206cdab726675e82f0e9d82dbf0872e5889da6d693c66561eeabd3

Observation 67dfffa3-33ec-4b33-a529-05935706537a · outbound

This paper cites Versatile Multimodal Controls for Expressive Talking Human Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Versatile Multimodal Controls for Expressive Talking Human Animation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:24:09.176859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:07.672135Z digest=sha256:740c0e4f98ace4364925e68adbd09a45ba25c0392b49a9c2026815c0f7f8dfac

Observation 5101bf6c-1045-4805-b9a0-3fa55900749e · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.747573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.747573Z digest=sha256:9263bdd4e5e3f751e39d3b625f9b8286b1289f63fe8cdf693d47a12b64a0416b

Observation 8f735ff8-6cea-431c-bd36-ad67c6eb923c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.813333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.813333Z digest=sha256:c2627b6330731d65bd66d05516c97041c690051c12ec6f155c82ef8f56e941fd

Observation aeca28fe-9986-42bb-a0a5-5681bdb3fc2b · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.908609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.908609Z digest=sha256:f240ed310483abd9b3b9dbac582de8563b08bef065825871d9d57a1f561aa626

Observation cd24ca58-02cc-4f59-afc6-e10a8b08301d · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:09.922883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:07.986728Z digest=sha256:81b62198fc78f3534924b2e76fc79ca8336e54aeebed9c215e75f107df303db7

Observation 5a0d1fb3-327d-4405-8d2e-8cd94feae4db · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Raft: Recurrent all-pairs field transforms for optical flow

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.082383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.082383Z digest=sha256:99c1b00c3443a897af88b555c5e7bb54a805d9a1701968cb5ed942354c5c3a2b

Observation cf2d0577-8566-4b21-bb82-73ba931ad1ae · outbound

This paper cites StableAnimator: High-Quality Identity-Preserving Human Image Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos StableAnimator: High-Quality Identity-Preserving Human Image Animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.163993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.163993Z digest=sha256:f8f4a53271a3f8467f85aec2a49c0135ff20202a57993e4e3ace1d99ea637a35

Observation 42b3d4fe-6134-4f02-b366-e5a92e2e859a · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.229848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.229848Z digest=sha256:8f4aee369efb3151bef19a28b30a20019451f4a07dd97453ea033019d79f276f

Observation fc65f202-0682-40f7-ace3-c363d5aed292 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Wan: Open and Advanced Large-Scale Video Generative Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.292449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.292449Z digest=sha256:8c62cc2b7ee7b13c793ab477a925ef06318c9e15a6eaeb7406b6de0dff9158a9

Observation 3182e70d-7162-457e-a16b-6345e95beead · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.International Journal of Computer Vision, pages 1–20, 2024.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Lavie: High-quality video generation with cascaded latent diffusion models.International Journal of Computer Vision, pages 1–20, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:09.761627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:08.373562Z digest=sha256:ee32311f3a465050dd2e6319e88c321ccbd45c677e3f18ddda927fde72fdce77

Observation f3bcd8ed-ba27-4c22-85a8-309c428e58a6 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.460168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.460168Z digest=sha256:fe1f31936ad3330d96c71cd645328a3733f41ad8480631ab62af9f242a3b86bd

Observation 6e2d9b1d-8293-4304-a815-28494b7920de · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991,.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.543292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.543292Z digest=sha256:242233cf3b593b217e499aab631f0aa8741ba1572c15089a36f71f953457feb6

Observation d6d1aa7b-b701-498d-aa61-9fe116bdff55 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.608636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.608636Z digest=sha256:42781d5976c1997752ae37ccf5b7345f4688503f76d0ba0ac00ebd68e888d692

Observation adcbc9e8-b530-4c38-8aea-e0c81a9ad333 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.669726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.669726Z digest=sha256:a9f6bfe70448da699c8975e8fbcb9917bfb0e461fe992d9c4234db40d3ea03da

Observation d0689d79-1e84-4c1b-b2cf-a75876eef380 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.721929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.721929Z digest=sha256:e4d6e67693f628f3e2a34146cbc5e395ecd4965f167655b68832dea268e0821c

Observation 22c0ad9a-3c42-4827-90bf-7a15d8360f3d · outbound

This paper cites an unresolved cited work.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T21:24:10.530104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T21:24:06.990235Z digest=sha256:84c9516af2254f0c4a51b90f441f9bfcfca5d36eb53bddb40dc444790c071a07

Pith citing papers

No inbound Pith citation observations are available.