Pith. sign in

Paper Citation Record · LEDGER

TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2508.13618.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13618 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:32:26.313485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.037365Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aff5c502-e45a-46f6-8ca0-34a01eb77b0b · inbound

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection cites this paper.

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:48:58.439681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T03:48:26.807495Z digest=sha256:1d91541087d7d47a8da37fad8c41d62703a60d3ff1b9aee863d3c66cfe12aa60

Observation a61b82db-94ac-4902-bdaa-b89186533c04 · inbound

EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents cites this paper.

EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.701727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T07:13:14.016037Z digest=sha256:ddefa859906af34b4e37f182353e0ab935c154fc7e4d0519153cdcc2d59bccc3

Observation b8967c69-f332-4dc5-8517-e487b831e57a · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.838171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:0be482213874949b7ebf4b035fb2a473f5039821872d9515a12b7a7d1310710a

Observation 57d2f305-9fe3-4d7a-bfb5-c781f9646639 · inbound

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection cites this paper.

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:04.220395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T14:04:52.065878Z digest=sha256:af5db50ed72772a9a0b9a43105556bf204df793c99d1bdd6a698eeb9419ed2e0

Observation 8f3c463c-7e35-4739-ae30-61409daab282 · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.373765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:cb59b03f0588b755a298507a8c5bc14cfe645bfd9d26ae61302c48ee257c3a14

Observation 80cd7ae3-9b44-4ac3-94bc-49f1ca17d1c4 · inbound

Loki: Representation over Architecture for Diffusion-Based Portrait Animation cites this paper.

Loki: Representation over Architecture for Diffusion-Based Portrait Animation TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:24:49.843928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T15:23:48.899214Z digest=sha256:07ce538f7aa964b6a7bd1b6c8f146531129de4e448fe8aa01da339f1d501e87f

Observation 6b01a8fd-5bb2-4d0f-a464-9b771debbcb7 · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.912836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:72f3e6b8442d09ed4dbc07474cbb3bd8c108e728273f54a6e5dff012a93da9df

Observation da29839b-403c-4925-a329-2f9f826e875a · inbound

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning cites this paper.

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.691764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:03:27.312548Z digest=sha256:3cc6f5839eb29b1d975a0919750306c4810d80db7daeb92cf8853c9938828320

Observation 07d31efe-f448-4693-8484-ed62606324eb · inbound

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation cites this paper.

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:56:39.445907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T08:38:44.906160Z digest=sha256:b8bf3e05dc63f4d605b9efd97cc8ea6c5f1f8964aceeac20ef3fa1a304ee72ef

Observation 53b7fc41-61f6-4ec3-9328-2fed4d84fb54 · inbound

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization cites this paper.

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.038924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T13:40:14.506288Z digest=sha256:77f050aba02f11185addabe4d0a26ccb31e4bae94d0a90cd008aaa92bde2dd05

Observation 5a1053a8-3470-49b4-8b6b-7998d5137aad · inbound

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization cites this paper.

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:32:26.313485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:32:26.313485Z digest=sha256:feaec446b303e10c65f8d293211d0afd647cf482c6f44376ae6cb11decbcfdaa