Pith. sign in

Paper Citation Record · LEDGER

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

As of 19 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2508.06511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06511 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:43:27.553611Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:01:54.160626Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:13:27.684326Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy57
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23e7f1ae-2355-4938-bc93-9c8cc521c53c · outbound

This paper cites Spatio- temporal energy-guided diffusion model for zero-shot video synthesis and editing,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Spatio- temporal energy-guided diffusion model for zero-shot video synthesis and editing,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.779195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:20.438409Z digest=sha256:79f7f35a2a3761b04324f80cf045dea6e7cf164d1bf255c600979b9f7091a081

Observation 61d0c4ad-4374-47cc-be30-e068769bfe9b · outbound

This paper cites Tvg: A training-free transition video generation method with diffusion models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Tvg: A training-free transition video generation method with diffusion models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.565940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:20.607686Z digest=sha256:896d2f421f32675eab020819780d8e5af1e845bc7008c25f43f97d9a05b594fb

Observation 0d1cd090-bf90-4703-a118-f3a4e125a3dc · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:20.763128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:20.763128Z digest=sha256:23ae06be410ee2f4b5c9e4ad0f7a641dd1ba35408f3b8d6b046fc6c86f8cbc1a

Observation 76b08039-2173-4305-a945-c8532ab21bea · outbound

This paper cites Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.365671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:20.872571Z digest=sha256:d7cedfe70dab1bd3a67b1d4f5cdee9aec3bcd9f123c58d41eeee7d8cd33f8fe2

Observation 15db9d35-dc96-49b4-b9bb-608b5e91ebef · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Wonderjourney: Going from anywhere to everywhere,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.228910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:20.959506Z digest=sha256:4081f2867b6338feae9ca5b1d7ec2fc1a58082a2e52639cce49f37249a6c70db

Observation ac4c71df-2336-42b2-87ca-4dfcb924d932 · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.998274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.076719Z digest=sha256:80e8ea2041001e3c180a9a19ffe5f05f613e599ebd1cea64ea6150efb397b746

Observation a84b4c09-8d90-4c46-b4de-15bb8d934705 · outbound

This paper cites Audio-semantic enhanced pose-driven talking head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Audio-semantic enhanced pose-driven talking head generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.769317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.154548Z digest=sha256:038387a6be1f8507898fad7afc9c6c9e4ebcf042a18629acd85674f9dd3d8d35

Observation b29b8937-4508-40fc-9517-030d68a8c97b · outbound

This paper cites Alleviating one- to-many mapping in talking head synthesis with dynamic adaptation context and style adapter,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Alleviating one- to-many mapping in talking head synthesis with dynamic adaptation context and style adapter,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.587619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.325531Z digest=sha256:199ed4dd761ed674e4b0886813748e9e57e20bb2854ab50871838c4af2819a4a

Observation 76152d6b-50ad-45db-a1ca-01b56ee10fb1 · outbound

This paper cites Stochastic latent talking face generation toward emotional expressions and head poses,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stochastic latent talking face generation toward emotional expressions and head poses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.403599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.435507Z digest=sha256:23d30323cc9b4da9b083bd9df8246677dbc542f6d9ee3bab08c3cf044d3fad64

Observation d19d0b5e-9631-4218-a31d-96a41128f5e8 · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait image animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo2: Long-duration and high-resolution audio-driven portrait image animation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.249313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.543133Z digest=sha256:6dde55a871e9d2c13c3f0d2b770704c0a929dae345b5a562a67948dc21e4dfdb

Observation 452e0136-51ca-4206-a242-b16880eb494e · outbound

This paper cites Out of time: automated lip sync in the wild,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Out of time: automated lip sync in the wild,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.073999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.646913Z digest=sha256:b3ae2b2938751ba94ce4a38c591c46ddded9a2da1c014607374c37f49f575cf5

Observation d97cfdfd-9130-4f64-aa87-18fec9500e84 · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Styletalk++: A unified framework for controlling the speaking styles of talking heads,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.881020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.773061Z digest=sha256:e405d4a5312337fac63b42925e1c13c482a1cdb4201e520c4f7ca1658d8cf276

Observation 1788aa76-1a4c-4014-80e0-d6f57b38682f · outbound

This paper cites Multimodal inputs driven talking face generation with spatial–temporal dependency,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Multimodal inputs driven talking face generation with spatial–temporal dependency,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.666264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:21.885332Z digest=sha256:4e9333a2521b25042820aef65f27b2ff0f13c78b59ad5334561cca894af4d600

Observation 6d903a5b-71d8-4679-b2f9-6a477f61c1e7 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation A lip sync expert is all you need for speech to lip generation in the wild,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.466007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.054746Z digest=sha256:228070e8eb41cbd1508845bdb5c2cbd2e506f27ed12c32d717c2c07b39660348

Observation 6d23a5b6-e5de-4f06-b5a4-c1ed4db6f079 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits anima- tion,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Difftalk: Crafting diffusion models for generalized audio-driven portraits anima- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.226831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.113727Z digest=sha256:9524b00c3b2d5683ef3eebb41d96cb969b5ea7f7ba2abb0167c2ff8288ea6809

Observation 717e74f8-781b-4d37-8072-472b6c6aad71 · outbound

This paper cites InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:22.208132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:22.208132Z digest=sha256:e912fb2c43df37b543bf8054b474615be0cbde819b08383fad455ce93849ec75

Observation 88dafea1-90d0-4a4f-91e9-b5bf13a1187e · outbound

This paper cites Moee: Mixture of emotion experts for audio-driven portrait animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Moee: Mixture of emotion experts for audio-driven portrait animation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.088689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.291883Z digest=sha256:188b09a9ca43e88d70087628446a33455824337a93d7717b528d9ffcc35210f1

Observation d38c91a5-5f69-4b5b-a381-49458102688f · outbound

This paper cites Face recognition based on fitting a 3d mor- phable model,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Face recognition based on fitting a 3d mor- phable model,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.845916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.352937Z digest=sha256:f144568c2f3f9d82d60e8536163d10128ac55210d109c601a04b7a57b1589532

Observation 35c75f7c-e5a3-489b-8dce-e3561e78b74a · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.645001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.461146Z digest=sha256:cf97c0c2ab477ee6f5844037f09429dd605bb35ded03120be97d5bc9d01c33ca

Observation b95cff9c-01c3-4fe1-92c9-21f702ab2b5e · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.470591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.572249Z digest=sha256:b70da6e8ea1da00a0f61c95f4e1c4c981ea7e7adb2fa0637b17ca275cfc73614

Observation 791123d2-719c-4b7b-9375-4d5aab3d8325 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speaking styles,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Styletalk: One-shot talking head generation with controllable speaking styles,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.296235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.625988Z digest=sha256:9e7a476a945e2fdf725d9fa2a315e84baefb3da8e026c29c5e8d53e55c726699

Observation d8ebe8a7-1bc3-4514-86e7-30e255eb9b84 · outbound

This paper cites Style2talker: High-resolution talking head generation with emotion style and art style,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Style2talker: High-resolution talking head generation with emotion style and art style,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.101665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.709334Z digest=sha256:8d24fe88f0f8cb1cda49e99348c1d366240e4d73153c4b7afffd0c6c0e6f5e7b

Observation 0dbb7566-bc70-4c97-a2e8-b83e276edda7 · outbound

This paper cites Edtalk: Efficient disentanglement for emotional talking head synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Edtalk: Efficient disentanglement for emotional talking head synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.858085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.797427Z digest=sha256:62fe8e706fb633164b0ff541bf689d4e5c73d57263fed7b5c00a55bb0ffc6d19

Observation c14acd06-90e2-43af-a9dc-f3c400550026 · outbound

This paper cites Say anything with any style,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Say anything with any style,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.701760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.909984Z digest=sha256:7aa30029fdcb095ac8b587e978b159916947befa9342a288ef50cf2aa73aa626

Observation 0d175107-6644-4bba-96a0-2bc4906775bd · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.476957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:22.942861Z digest=sha256:10a5c3ead13e320d397aa0a89c002288a85bca26a9d86f2294ce23f23d1441fd

Observation b24e3996-28ca-44c3-81fa-d4db7c263a02 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.307021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.033694Z digest=sha256:9c3f51739ae6c7a49f798e1334e6e6399c7f3801e9f9750fb7b322a8609a0431

Observation 46cb5150-1c18-4cd5-a097-6fe6d9085ed0 · outbound

This paper cites Real3d-portrait: One-shot realistic 3d talking portrait synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Real3d-portrait: One-shot realistic 3d talking portrait synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.138099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.090722Z digest=sha256:e66a5887a63131236c5e678f53026494e24efa82dfdfaf26f76a6bc76ec85efa

Observation 7c348509-6106-48d2-b11a-2e557bc8570f · outbound

This paper cites Scalable diffusion models with transformers,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Scalable diffusion models with transformers,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.907031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.147505Z digest=sha256:648b2e8966cb3289b4729f9604341fede878229d863fbeba48d57a1890d8b6ec

Observation 6c8df61a-d7fe-40cd-9420-e1bc79c41d92 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Cogvideox: Text-to-video diffusion models with an expert transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.727718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.241228Z digest=sha256:fb7e13428fd3806fc794858d157ab38ff3b2c74d7525292ccc13af940b9bd4c0

Observation fd53d714-bd38-4778-b867-a7e1fb121158 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.332696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.332696Z digest=sha256:489c2e4270c82d9ede5b98460358b90bb5b3eb5b981383a1aaf36ee39331a29a

Observation 64874710-1360-4236-96b1-bf5fe39e4031 · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Efficient emotional adaptation for audio-driven talking-head generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.566500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.427386Z digest=sha256:432cdebd028eff28f5097f86ff14ed6a6e7a93e350ac266555b280d7f1f0d196

Observation 86f3dfd8-79ed-4ef6-b804-61215e02fe25 · outbound

This paper cites Talkclip: Talking head generation with text-guided expressive speaking styles,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Talkclip: Talking head generation with text-guided expressive speaking styles,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.404635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.500353Z digest=sha256:30c0b928f4cd85c4b605ed12c94fecc3ef31a1f1bae71fc34f88aeb5eff9e558

Observation 9faeff50-dbcb-4bb2-8b88-134aec1e951d · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Whisperx: Time-accurate speech transcription of long-form audio,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.589831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.589831Z digest=sha256:134dcf405d233637192d42b41aa933c321343dca8ba4c7b4224af55adbde100f

Observation 499289a0-6f66-421d-af3e-61575593d220 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.684599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.684599Z digest=sha256:76f38963a44c23f1bdf8d045f69d7a6855007b9525da0d58f69bf13fe74f9ed5

Observation 9da71d78-ef20-4e73-bb0d-8d84202ec70f · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.782760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.782760Z digest=sha256:9cd2bab804394cd5a7af8c29287a18aa6daf88753ef9f4c108925913cf1849f6

Observation 85e08829-a9c3-4242-8017-f0cb3d426677 · outbound

This paper cites Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.244426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.845045Z digest=sha256:9f1c476c92769903a5bf04b46af8678d3d1e0c88a30fd2146eb18e15e8805987

Observation 73f9ec82-391c-47bf-95dd-a89513ed5f6a · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Dinov2: Learning robust visual features without supervision,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.038179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.915048Z digest=sha256:aef5968eac8fb5ab644aab60b9198068b578bccdb7d12e70aa1bdf7ccf3ab7d1

Observation 0c8471b9-20c6-470d-bb11-dddd135e1d93 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.847464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:23.985124Z digest=sha256:7121bb787649ef77875151fdeba82c2308d3e889257a3841711598590c8104df

Observation 8be51ba6-c19e-4c1e-a37d-b2156889883f · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation CelebV-HQ: A large-scale video facial attributes dataset,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.636580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.080905Z digest=sha256:8371a23ce100f9ff167afc66de7845b4d1cd73ea7c4252ad0a790b6a28c14776

Observation fe30ef19-9d24-4883-be37-144cd3bed17a · outbound

This paper cites Hierarchical feature warping and blending for talking head animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hierarchical feature warping and blending for talking head animation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.433129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.238790Z digest=sha256:3924d2ab09de2245655ecbe31287694bb3093115dcb6bf334890bfe32f4d9e78

Observation 75d24d69-e53d-4113-a87b-be33cd9f0c26 · outbound

This paper cites Denoising diffusion probabilistic models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Denoising diffusion probabilistic models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:24.356376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:24.356376Z digest=sha256:3b2d43660107332175020787764be3c388dfc66f3ac64ed7f7e1ae457d7c0035

Observation 7a07d779-5799-4c25-bbcc-a72fdf3677c9 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Diffused heads: Diffusion models beat gans on talking-face generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.261330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.523648Z digest=sha256:a49845bbf98e25f3060611c4b8aa9f5f6b9a4ccb94b2b3f43bcb8afd292cc10a

Observation e1da9b9c-3384-4687-9fd7-cb1f31ca2b0f · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Loopy: Taming audio-driven portrait avatar with long-term motion dependency,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.085060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.615081Z digest=sha256:ad01b5567404fc58ee5465675c26b1e3d0b2c11431e9b7fe53ab67e8c87d9429

Observation 977c85de-1616-4d4b-978d-883bfa9f98d4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:24.682758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:24.682758Z digest=sha256:e91e9cbadbfdfc7cc0e218670aceef2315f69e4db2fe9fb6ea743637e0a1f8b9

Observation ca2ad9aa-9256-48cc-b7c2-f94119409e40 · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Efficient emotional adaptation for audio-driven talking-head generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.863138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.780806Z digest=sha256:0ca9343c2e1540ca9466535a5eda23bfaf5ae089a75831241c79a8527f4e8752

Observation bbde5224-6056-47f0-aa05-f09de5ef64a0 · outbound

This paper cites Emmn: Emotional motion memory network for audio-driven emotional talking face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Emmn: Emotional motion memory network for audio-driven emotional talking face generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.669046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.877887Z digest=sha256:06f3d17689d6b8ec2f3372c82ad88fe3a566f9a9f6b3f2ff92ee3cf35a47db4a

Observation 001f16cd-4748-4652-a42f-0e9d09616507 · outbound

This paper cites Talking face gener- ation with audio-deduced emotional landmarks,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Talking face gener- ation with audio-deduced emotional landmarks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.452342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:24.940996Z digest=sha256:b14021192868b2e8551bc15aea0e6a8fa08f1c2cb1934877d50543fb6ad3ce0b

Observation 82de0386-a1b7-4c06-8e82-c8e3a89ccabe · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.294588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:25.046119Z digest=sha256:3d95d961b080afd7e83a5a779176d547d767b1c5e053b0684985c3ba505d3f24

Observation e06e0b86-af4b-4930-8ac4-68a5f866dc0d · outbound

This paper cites Neural discrete representation learning,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Neural discrete representation learning,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.103654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:25.180962Z digest=sha256:55db5b5296a7ea4b4cc68d36ef357bd9cff0ba0eadb023f81d17b92c63e6a670

Observation 1df25f53-5755-424b-b07b-c463d8e2d844 · outbound

This paper cites Progressive disentangled representation learning for fine-grained controllable talking head synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Progressive disentangled representation learning for fine-grained controllable talking head synthesis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.839573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:25.288675Z digest=sha256:8fee22e94d07b56ef744bfc5102ca522e66606916976dd090500fc6e8a31b62c

Observation 4841976a-bf37-4262-a744-c86f37d48914 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.371332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.371332Z digest=sha256:ddababf91fad24d09872e5af12fc480bf27ed4d82d9e2c71f451d7dc9623fc6f

Observation 12e53d0b-d73e-49f8-9cd5-eaf8efb3e117 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Animate anyone: Consistent and controllable image-to-video synthesis for character animation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.685967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:25.473087Z digest=sha256:1e086635c26eff1c3b9759ae575b0835fe0ab2e2ee9212cd5f445f940ab11c95

Observation d8204951-6415-48ad-8c77-e803742d4b90 · outbound

This paper cites Megactor- σ: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Megactor- σ: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.479263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:25.616433Z digest=sha256:1b669524b00969ee6240dba8dff7c2cc01f3736b2a3c4b718fb88b74da917874

Observation fedecec2-efc7-46d6-9efa-406140d0c4be · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Vasa-1: Lifelike audio-driven talking faces generated in real time,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.267860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:25.797404Z digest=sha256:87a6f20c17fa2888c28b57d719a0fce3ee009bbae469bd8f9c10c7845323d3b9

Observation d810176f-352f-4331-9db9-d1a337d234a0 · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation High- resolution image synthesis with latent diffusion models,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.877250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.877250Z digest=sha256:f86792278cd731065c4c574010d36a90c2952666f55e38f21cf1020fefb02c6b

Observation 1395334b-d60f-4ab7-82ed-0b771d9c44b9 · outbound

This paper cites Easyanimate: A high-performance long video gen- eration method based on transformer architecture,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Easyanimate: A high-performance long video gen- eration method based on transformer architecture,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.953112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.953112Z digest=sha256:ac0ce7748b84cf05727460326b7b335a09b29863b9bb4aa7a98f9e1f5fbbb40f

Observation c4f017e6-91b8-4be8-aa3d-ed68e2a76e37 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:26.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:26.010676Z digest=sha256:d4ac1b694d5f84c86d4c87e1fd0d425cf3eb17949b16cb5594d735d64aa1e163

Observation 99c38c09-4ff1-48b9-b933-742b6633e492 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.081629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.094954Z digest=sha256:1eaca547535ee007bd0a3993ec244ae201510ed4d4427ba3111324a5576ef432

Observation 07ce0a9e-4aa9-42c5-9f5c-51fbb3c3124a · outbound

This paper cites Effective whole-body pose estimation with two-stages distillation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Effective whole-body pose estimation with two-stages distillation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.879551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.179445Z digest=sha256:287dc69ef0f76aecae6d32dedbab240622921fd591690ad373c4717f7663dd62

Observation 97dc7f49-0224-44b6-967f-5e19fca39027 · outbound

This paper cites Stylecrafter: Enhancing stylized text-to-video generation with style adapter,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stylecrafter: Enhancing stylized text-to-video generation with style adapter,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.655657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.283689Z digest=sha256:431578c8e20996e294e8fa84c337cc303f92009787bd8db9be39fc3265791a58

Observation b0e833b1-9f78-4b4d-97ba-b86315713c97 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Learning transferable visual models from natural language supervision,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.491485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.373131Z digest=sha256:db828868b80855c93ac2964c8c8ef4b17dc45885c8ac7928be76a1bdbed1a592

Observation 0bc9ccad-1c30-45bf-bf48-2b8db2da1e75 · outbound

This paper cites Vision transformer with quad- rangle attention,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Vision transformer with quad- rangle attention,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.346309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.457543Z digest=sha256:c67ba5494a0b397f91b5b08ff18d1079f568a8f7e99b2ab005bed513588d2651

Observation 10ac8c89-d956-4722-8868-8d3baa0e54ed · outbound

This paper cites Celebv-text: A large-scale facial text-video dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Celebv-text: A large-scale facial text-video dataset,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.150137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.612613Z digest=sha256:4180471edc74daae4f38e5e4247c9a9ff618ba04fb9da412b74a46b51fa482ea

Observation eb870c7b-82ee-490b-9c7b-310cc136010f · outbound

This paper cites DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:26.732825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:26.732825Z digest=sha256:d5c91b0135dc544c603e01c4a2cbd4aa570e0a040077d365805e1482675b2302

Observation cfad3b58-5566-457a-8896-b88f521e39f0 · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.940438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:26.853076Z digest=sha256:7a125a77cf4bbd4b32b6e688147d4957bdf124c29feafbeec11fd35ba65d8af6

Observation 74c75899-b98c-4529-acdf-2cdec7496d4e · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:27.004117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:27.004117Z digest=sha256:de28824901ef22eb408b7241b22c73324b7500324afb41d323d2d885890d68c4

Observation 6d984e16-3249-45e2-a821-dcc5d65678b2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.758591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:27.132662Z digest=sha256:33969edb9571e79606a74c14de7f751a5835ccf7db4cc72e33e95ae768145262

Observation d756d998-225c-4e56-8b4b-280b72fec00a · outbound

This paper cites Video-to-video synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Video-to-video synthesis,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.559570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:27.290711Z digest=sha256:dab8406fdc5ad0d690c572cd47ad3fd577f47aa272ae13df9113b2e09141e489

Observation 7db45693-961c-4dc3-afe1-336b58182499 · outbound

This paper cites Ani- mating arbitrary objects via deep motion transfer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Ani- mating arbitrary objects via deep motion transfer,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.352098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:27.374229Z digest=sha256:8aece50b15fcd683bf96ff6fc443dd53d4ccee78b7102634f6aea674d3545c0d

Observation fb8e6625-bf93-43b8-9a23-5b65c68f696c · outbound

This paper cites Seeing what you said: Talking face generation guided by a lip reading expert,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Seeing what you said: Talking face generation guided by a lip reading expert,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.145129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:27.454581Z digest=sha256:455580817c306309e7389e23475177132763e29b286bd3621e48b356e0be4132

Observation 64953955-4a54-4cc1-a894-2bdbaeebdd65 · outbound

This paper cites Towards robust blind face restoration with codebook lookup transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Towards robust blind face restoration with codebook lookup transformer,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:27.914596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T12:43:27.553611Z digest=sha256:851fd015054cfd8cee38d8c51e771dabf0bca29c339171edb1d0a8224aa6f15b

Pith citing papers

Observation d773345f-2744-4f8c-a9d3-c70dc31562d1 · inbound

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning cites this paper.

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.685726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T13:03:27.312548Z digest=sha256:5562605752399850a6edff8f46346c8b67a1bfc9e9fcd738fd03f255d2fd5111

Observation be47e19f-0ab2-41bd-8864-6cb4b794748e · inbound

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models cites this paper.

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:54.160626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:01:54.160626Z digest=sha256:c5e1533e470b1ebd4b775739f8e26bee125af302df125935b10312f41156a1d8