Pith. sign in

Paper Citation Record · LEDGER

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2508.06511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06511 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:43:27.553611Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:01:54.160626Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:13:27.684326Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy57
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23e7f1ae-2355-4938-bc93-9c8cc521c53c · outbound

This paper cites Spatio- temporal energy-guided diffusion model for zero-shot video synthesis and editing,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Spatio- temporal energy-guided diffusion model for zero-shot video synthesis and editing,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.779195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:20.438409Z digest=sha256:afc487d630c1595f48c0100b704a72a481ed9371732b5e4e5413254daa8a8b5f

Observation 61d0c4ad-4374-47cc-be30-e068769bfe9b · outbound

This paper cites Tvg: A training-free transition video generation method with diffusion models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Tvg: A training-free transition video generation method with diffusion models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.565940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:20.607686Z digest=sha256:55c8ca4e71f9e383d01a57af9671375c7df49dd927806129c9975699d42ec6f1

Observation 0d1cd090-bf90-4703-a118-f3a4e125a3dc · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:20.763128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:20.763128Z digest=sha256:b5dfd072cd8fba561e659d3182beebfa595f5c424b812c06c9139985aa1fa717

Observation 76b08039-2173-4305-a945-c8532ab21bea · outbound

This paper cites Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.365671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:20.872571Z digest=sha256:0d295abb632ae53eb2cb85bb64189855c11765f3a8613ee59c0750d488ada3b0

Observation 15db9d35-dc96-49b4-b9bb-608b5e91ebef · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Wonderjourney: Going from anywhere to everywhere,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.228910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:20.959506Z digest=sha256:89ac5b01ef8f39e4cab84dccd5ec6ed235dd06567e5e23949944511e6376c39d

Observation ac4c71df-2336-42b2-87ca-4dfcb924d932 · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.998274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.076719Z digest=sha256:cf7cb71add7b897a232dc1b3d79756ae25d811945f3889b5a49ecf883e57bfee

Observation a84b4c09-8d90-4c46-b4de-15bb8d934705 · outbound

This paper cites Audio-semantic enhanced pose-driven talking head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Audio-semantic enhanced pose-driven talking head generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.769317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.154548Z digest=sha256:d3a992e64d4e32103b534be4ce0198ffe6f7314edef650b327e79374bb85a3fe

Observation b29b8937-4508-40fc-9517-030d68a8c97b · outbound

This paper cites Alleviating one- to-many mapping in talking head synthesis with dynamic adaptation context and style adapter,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Alleviating one- to-many mapping in talking head synthesis with dynamic adaptation context and style adapter,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.587619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.325531Z digest=sha256:203341fe449f921c6e60cc3af2501c169d56bcebdcde71c5944106b6a920cc58

Observation 76152d6b-50ad-45db-a1ca-01b56ee10fb1 · outbound

This paper cites Stochastic latent talking face generation toward emotional expressions and head poses,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stochastic latent talking face generation toward emotional expressions and head poses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.403599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.435507Z digest=sha256:3d5a5c6783e890439e40d7122e881ef2169c84cc2c352e12ac8d4985dd2942c1

Observation d19d0b5e-9631-4218-a31d-96a41128f5e8 · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait image animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo2: Long-duration and high-resolution audio-driven portrait image animation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.249313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.543133Z digest=sha256:12a7a96ee037a792211de4165919f4048f6106384ac6e7d93f26ac6e62f9cae0

Observation 452e0136-51ca-4206-a242-b16880eb494e · outbound

This paper cites Out of time: automated lip sync in the wild,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Out of time: automated lip sync in the wild,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.073999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.646913Z digest=sha256:db5b62adac9c26bc1b269aca5a7e696f9bf59f6a0a6b4ff53f467050ed2bcce8

Observation d97cfdfd-9130-4f64-aa87-18fec9500e84 · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Styletalk++: A unified framework for controlling the speaking styles of talking heads,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.881020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.773061Z digest=sha256:66ee7638045bd045bf9c56d1d96231ad73c5bf77ed5f5426c65ba2a67625b4dc

Observation 1788aa76-1a4c-4014-80e0-d6f57b38682f · outbound

This paper cites Multimodal inputs driven talking face generation with spatial–temporal dependency,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Multimodal inputs driven talking face generation with spatial–temporal dependency,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.666264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:21.885332Z digest=sha256:00d8aa33885572ca51df470d6116b4447082c5c8b9ff2e9c75a9d9944b8a5ac1

Observation 6d903a5b-71d8-4679-b2f9-6a477f61c1e7 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation A lip sync expert is all you need for speech to lip generation in the wild,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.466007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.054746Z digest=sha256:778cf8e4a8202d758014377fff383564775a1ed27d2103b35b8ffda5fa7ee047

Observation 6d23a5b6-e5de-4f06-b5a4-c1ed4db6f079 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits anima- tion,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Difftalk: Crafting diffusion models for generalized audio-driven portraits anima- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.226831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.113727Z digest=sha256:9727f509135c9e81c96c61d17f17b4851b38ec4b0a9bea489c5bb9dd34d84715

Observation 717e74f8-781b-4d37-8072-472b6c6aad71 · outbound

This paper cites InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:22.208132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:22.208132Z digest=sha256:92b6ffefc55a7d80947fc18541afb2a98bf2bee19d9ac846e90a6555fe56871b

Observation 88dafea1-90d0-4a4f-91e9-b5bf13a1187e · outbound

This paper cites Moee: Mixture of emotion experts for audio-driven portrait animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Moee: Mixture of emotion experts for audio-driven portrait animation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.088689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.291883Z digest=sha256:35abd4856544174deaa3be657ac83cf5a62d360f2b50708ec0395901356fd822

Observation d38c91a5-5f69-4b5b-a381-49458102688f · outbound

This paper cites Face recognition based on fitting a 3d mor- phable model,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Face recognition based on fitting a 3d mor- phable model,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.845916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.352937Z digest=sha256:596f859d9c39fe25f6bc8de7fe6deed607788fb0a7c5997f6611fabc19414a91

Observation 35c75f7c-e5a3-489b-8dce-e3561e78b74a · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.645001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.461146Z digest=sha256:5adf21ada0648b9338b5be8f65ebb180092d7a36824c2d4f65803f5df63ba4f7

Observation b95cff9c-01c3-4fe1-92c9-21f702ab2b5e · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.470591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.572249Z digest=sha256:fecac00af27211549a5b0cc35713f54449514e8ec55d5f120988100eeaaf6362

Observation 791123d2-719c-4b7b-9375-4d5aab3d8325 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speaking styles,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Styletalk: One-shot talking head generation with controllable speaking styles,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.296235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.625988Z digest=sha256:a182d750ee6c949241ee30dea5384043fe0754b34eeebf466dd34ab68d61b9e1

Observation d8ebe8a7-1bc3-4514-86e7-30e255eb9b84 · outbound

This paper cites Style2talker: High-resolution talking head generation with emotion style and art style,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Style2talker: High-resolution talking head generation with emotion style and art style,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.101665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.709334Z digest=sha256:735805b473c970d94e85c73375d660436100c46caeb7dce5a57df2d96e1a9145

Observation 0dbb7566-bc70-4c97-a2e8-b83e276edda7 · outbound

This paper cites Edtalk: Efficient disentanglement for emotional talking head synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Edtalk: Efficient disentanglement for emotional talking head synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.858085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.797427Z digest=sha256:1cb3289f336bef5bfc2afbb1cdfe8860e923a48e52c47ed5c43bc20250d3ecba

Observation c14acd06-90e2-43af-a9dc-f3c400550026 · outbound

This paper cites Say anything with any style,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Say anything with any style,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.701760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.909984Z digest=sha256:65a2ca78b8bf3b21cde0ac57f3211f9ba35b7558d737f91723a4a62b60710745

Observation 0d175107-6644-4bba-96a0-2bc4906775bd · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.476957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:22.942861Z digest=sha256:b0da2c172fdbac1a325dc4533e9c52a8230bb056a96db84059bbeec3375019d5

Observation b24e3996-28ca-44c3-81fa-d4db7c263a02 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.307021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.033694Z digest=sha256:5835475c0174c79d0866c945e662bd92c8c524a7bf68a12ed9c2561320483790

Observation 46cb5150-1c18-4cd5-a097-6fe6d9085ed0 · outbound

This paper cites Real3d-portrait: One-shot realistic 3d talking portrait synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Real3d-portrait: One-shot realistic 3d talking portrait synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.138099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.090722Z digest=sha256:40cb5529197910280494f1ffbc22a2fcf795d12afb3af2995f2cd618427a8b4b

Observation 7c348509-6106-48d2-b11a-2e557bc8570f · outbound

This paper cites Scalable diffusion models with transformers,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Scalable diffusion models with transformers,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.907031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.147505Z digest=sha256:3cb10b0e4df065d198fdb47c6e6a053d5e94ff3fe85f781cc82ce248ae6cf775

Observation 6c8df61a-d7fe-40cd-9420-e1bc79c41d92 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Cogvideox: Text-to-video diffusion models with an expert transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.727718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.241228Z digest=sha256:9c8f3a5b85cdffdb0c2b3942ae377665df887ce6cc53b4e2bd97cf20561b1057

Observation fd53d714-bd38-4778-b867-a7e1fb121158 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.332696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.332696Z digest=sha256:49dc01166659d5df0e926f000ea2a179fe5db1e421568237e9f982c91d30d40d

Observation 64874710-1360-4236-96b1-bf5fe39e4031 · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Efficient emotional adaptation for audio-driven talking-head generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.566500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.427386Z digest=sha256:01094cc593308169c9dd834416a0b19d5486db80cd01404e9cbe7ae3ba82aadc

Observation 86f3dfd8-79ed-4ef6-b804-61215e02fe25 · outbound

This paper cites Talkclip: Talking head generation with text-guided expressive speaking styles,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Talkclip: Talking head generation with text-guided expressive speaking styles,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.404635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.500353Z digest=sha256:c744a9bbfc71bc5be91f633c7c5940b4a7659da7f518a420e82597156129acea

Observation 9faeff50-dbcb-4bb2-8b88-134aec1e951d · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Whisperx: Time-accurate speech transcription of long-form audio,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.589831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.589831Z digest=sha256:4e80a367297355ccab164a54618b6891f146e1a6ec63e0dbf4aa439ec16e0c6d

Observation 499289a0-6f66-421d-af3e-61575593d220 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.684599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.684599Z digest=sha256:9beda31f46becc79c01c37a0f4ccd7df53b552b6fe5e809166ee7dbd1b5191fd

Observation 9da71d78-ef20-4e73-bb0d-8d84202ec70f · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.782760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.782760Z digest=sha256:dbfa4ff351dc9bb5b23ab55be89e977277b06be9d57ea024db14f64e5ac745cc

Observation 85e08829-a9c3-4242-8017-f0cb3d426677 · outbound

This paper cites Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.244426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.845045Z digest=sha256:ed7ca695930673ef92ec13e2f53edb0c909a49155c8aca51b5f073b362f37cc7

Observation 73f9ec82-391c-47bf-95dd-a89513ed5f6a · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Dinov2: Learning robust visual features without supervision,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.038179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.915048Z digest=sha256:ec8a0df4f3569fb27b077ee79d09deeca8360cda364882275bca57df2ba37f90

Observation 0c8471b9-20c6-470d-bb11-dddd135e1d93 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.847464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:23.985124Z digest=sha256:8b4cb618ce862659d292e9cfb8e753a1b03bf6de747d78077a7bbda092f71220

Observation 8be51ba6-c19e-4c1e-a37d-b2156889883f · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation CelebV-HQ: A large-scale video facial attributes dataset,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.636580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.080905Z digest=sha256:4a65ba101e8647799c6cb4fedf3e39137cb6242f147ccd2c6d9bfc12627971ed

Observation fe30ef19-9d24-4883-be37-144cd3bed17a · outbound

This paper cites Hierarchical feature warping and blending for talking head animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hierarchical feature warping and blending for talking head animation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.433129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.238790Z digest=sha256:7d4f555010d18fc96dc705716ea55b534d598943b99018e1f26c668cd8f063a8

Observation 75d24d69-e53d-4113-a87b-be33cd9f0c26 · outbound

This paper cites Denoising diffusion probabilistic models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Denoising diffusion probabilistic models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:24.356376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:24.356376Z digest=sha256:6ccbb62bd58239696cd2779af148973139a689850c565c9482d0f5879fa2b712

Observation 7a07d779-5799-4c25-bbcc-a72fdf3677c9 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Diffused heads: Diffusion models beat gans on talking-face generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.261330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.523648Z digest=sha256:bdd6fac30e058b7f1991a667e8c14b56682dc8e640025975293270a733b31935

Observation e1da9b9c-3384-4687-9fd7-cb1f31ca2b0f · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Loopy: Taming audio-driven portrait avatar with long-term motion dependency,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.085060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.615081Z digest=sha256:1a935b6d3ee04e34deb99c4b8986d86f8aa8efa1a30723baf5dac325e1d09d8e

Observation 977c85de-1616-4d4b-978d-883bfa9f98d4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:24.682758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:24.682758Z digest=sha256:82dc4937ceeec6635a3437902be37732306fff9c0f2afbac5f1e9e24525d1048

Observation ca2ad9aa-9256-48cc-b7c2-f94119409e40 · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Efficient emotional adaptation for audio-driven talking-head generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.863138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.780806Z digest=sha256:213cf9b0d923c63ecb1fdf199e18c830c0c910d66f64ab69092867ef650d2e90

Observation bbde5224-6056-47f0-aa05-f09de5ef64a0 · outbound

This paper cites Emmn: Emotional motion memory network for audio-driven emotional talking face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Emmn: Emotional motion memory network for audio-driven emotional talking face generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.669046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.877887Z digest=sha256:41ae62e255b97b5b22605afaebba0175a83a29d7a98a29f754d5976c059c4d68

Observation 001f16cd-4748-4652-a42f-0e9d09616507 · outbound

This paper cites Talking face gener- ation with audio-deduced emotional landmarks,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Talking face gener- ation with audio-deduced emotional landmarks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.452342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:24.940996Z digest=sha256:89ee4bdc33c7374cfa18bf6b3c08b1c9a1277067ec57cbce67b9ca8a45b9dcb8

Observation 82de0386-a1b7-4c06-8e82-c8e3a89ccabe · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.294588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:25.046119Z digest=sha256:25f9a8bb358de3c704e557f7f51800964e02a604fbdf9e4423bb66171c432f25

Observation e06e0b86-af4b-4930-8ac4-68a5f866dc0d · outbound

This paper cites Neural discrete representation learning,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Neural discrete representation learning,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.103654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:25.180962Z digest=sha256:81cf2da80c5162e7c02f02c4149f320810f24d22fa65c48fae0560648c857b28

Observation 1df25f53-5755-424b-b07b-c463d8e2d844 · outbound

This paper cites Progressive disentangled representation learning for fine-grained controllable talking head synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Progressive disentangled representation learning for fine-grained controllable talking head synthesis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.839573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:25.288675Z digest=sha256:a0089dd28d0669e7bf593dcb5d43fe3e64a89ec85a2a038bc2e171dab1d83f01

Observation 4841976a-bf37-4262-a744-c86f37d48914 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.371332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.371332Z digest=sha256:9193cf2f70c15173fc11fedd31b246a5396185388256b5a9cee4567e687dd1e1

Observation 12e53d0b-d73e-49f8-9cd5-eaf8efb3e117 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Animate anyone: Consistent and controllable image-to-video synthesis for character animation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.685967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:25.473087Z digest=sha256:dd3b3697c353b725fb884922b7e88204f875928b2d49e335342cdbb109b5b9fb

Observation d8204951-6415-48ad-8c77-e803742d4b90 · outbound

This paper cites Megactor- σ: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Megactor- σ: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.479263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:25.616433Z digest=sha256:32dececf30be08fc76c580ccbb7f679dc14e42e0e7c27c8ee11fdd47b1287e0e

Observation fedecec2-efc7-46d6-9efa-406140d0c4be · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Vasa-1: Lifelike audio-driven talking faces generated in real time,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.267860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:25.797404Z digest=sha256:fa834f4f1caf4a236538a6c6ea1c001a516c6e4c81eb250c30c514a311b60ac9

Observation d810176f-352f-4331-9db9-d1a337d234a0 · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation High- resolution image synthesis with latent diffusion models,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.877250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.877250Z digest=sha256:e3be7f82ed2c740be3a21ca95e2359b10cd2c06d584702718a8d9e9564fac493

Observation 1395334b-d60f-4ab7-82ed-0b771d9c44b9 · outbound

This paper cites Easyanimate: A high-performance long video gen- eration method based on transformer architecture,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Easyanimate: A high-performance long video gen- eration method based on transformer architecture,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.953112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.953112Z digest=sha256:1debee24398dac6bf8649a6c60e43d6dea76d5fc99d3f280ad45557754860c65

Observation c4f017e6-91b8-4be8-aa3d-ed68e2a76e37 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:26.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:26.010676Z digest=sha256:d8042f299f81061869577a0e34d51c9fd3b02ba5719437c47eb78a2c67ed2d60

Observation 99c38c09-4ff1-48b9-b933-742b6633e492 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.081629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.094954Z digest=sha256:170f05f1950acb982f358317923b26e44143f06141988eb5dc4535bb44b9a875

Observation 07ce0a9e-4aa9-42c5-9f5c-51fbb3c3124a · outbound

This paper cites Effective whole-body pose estimation with two-stages distillation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Effective whole-body pose estimation with two-stages distillation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.879551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.179445Z digest=sha256:9b43156fa56bdb89e2b8020e591d75d7e09425d5876183898f7735065da34969

Observation 97dc7f49-0224-44b6-967f-5e19fca39027 · outbound

This paper cites Stylecrafter: Enhancing stylized text-to-video generation with style adapter,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stylecrafter: Enhancing stylized text-to-video generation with style adapter,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.655657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.283689Z digest=sha256:566e257a0f71ff1148026dde471db3c03b0c7008e950adb191999177e9c330d9

Observation b0e833b1-9f78-4b4d-97ba-b86315713c97 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Learning transferable visual models from natural language supervision,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.491485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.373131Z digest=sha256:36cc931d14a7e2434e29b420df19106eee61d6d4fe2de3c4d5f991132813bba3

Observation 0bc9ccad-1c30-45bf-bf48-2b8db2da1e75 · outbound

This paper cites Vision transformer with quad- rangle attention,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Vision transformer with quad- rangle attention,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.346309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.457543Z digest=sha256:f5807bc26bbf81f3291cf1bcfe41b8ea7ce84ce1269b827c6fe4f1f84d603074

Observation 10ac8c89-d956-4722-8868-8d3baa0e54ed · outbound

This paper cites Celebv-text: A large-scale facial text-video dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Celebv-text: A large-scale facial text-video dataset,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.150137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.612613Z digest=sha256:3676ac366e268a173a62c05371b067bbe6a2ebc903500b3245b93bf45b794522

Observation eb870c7b-82ee-490b-9c7b-310cc136010f · outbound

This paper cites DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:26.732825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:26.732825Z digest=sha256:0cd71d24840e8c15ea6488066f5de880464c7d8625df04eb0867cd65fa43e681

Observation cfad3b58-5566-457a-8896-b88f521e39f0 · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.940438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:26.853076Z digest=sha256:a1ecf10045e0a64a9f3d74cda1787158cb653baac3abcddf021be88dbed2907a

Observation 74c75899-b98c-4529-acdf-2cdec7496d4e · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:27.004117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:27.004117Z digest=sha256:6c561fd0180335ac26270c3df597706030e425be4bc048bab80d32f1598ca21b

Observation 6d984e16-3249-45e2-a821-dcc5d65678b2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.758591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:27.132662Z digest=sha256:c73063c7a66ab1d73b7d7de30fa52ea5d9256d337a2a3b7647aab3b3bcf208b8

Observation d756d998-225c-4e56-8b4b-280b72fec00a · outbound

This paper cites Video-to-video synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Video-to-video synthesis,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.559570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:27.290711Z digest=sha256:b601b2cc69fc49959c9e0d42f70f2e57c6775de58bdf8dc6e9cc1ddb25541eb1

Observation 7db45693-961c-4dc3-afe1-336b58182499 · outbound

This paper cites Ani- mating arbitrary objects via deep motion transfer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Ani- mating arbitrary objects via deep motion transfer,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.352098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:27.374229Z digest=sha256:a2ccf521043350def5f5648f8766343e392014908b77e6179bd802504f0ec0ed

Observation fb8e6625-bf93-43b8-9a23-5b65c68f696c · outbound

This paper cites Seeing what you said: Talking face generation guided by a lip reading expert,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Seeing what you said: Talking face generation guided by a lip reading expert,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.145129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:27.454581Z digest=sha256:23aa4a12408f14ebf683863dd33b09ff28b74660631b68b255f057651777ee36

Observation 64953955-4a54-4cc1-a894-2bdbaeebdd65 · outbound

This paper cites Towards robust blind face restoration with codebook lookup transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Towards robust blind face restoration with codebook lookup transformer,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:27.914596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:43:27.553611Z digest=sha256:15fcec3b84588e714bec184dac2d20c7c3dd53df84f3060fadceed3d73324873

Pith citing papers

Observation d773345f-2744-4f8c-a9d3-c70dc31562d1 · inbound

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning cites this paper.

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.685726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:03:27.312548Z digest=sha256:5ed71c77fc7800160d2e71e1396d65f6d4d5eb6ab0dd40cdb1025e0b7aaea147

Observation be47e19f-0ab2-41bd-8864-6cb4b794748e · inbound

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models cites this paper.

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:54.160626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:01:54.160626Z digest=sha256:c1f46ffdfff56a1677288f68e49502a905702265f084a5624a567e57331e3123