Pith. sign in

Paper Citation Record · LEDGER

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

As of 21 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2605.31294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31294 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T22:45:39.443440Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 238f65e8-cfd0-4a26-9957-7a4733319a1c · outbound

This paper cites an unresolved cited work.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:8d149a283c0a8ef981d7011177e70ada0f028bb097bdb4e4cfeb48bd91005c40

Observation 62fac434-eb11-4576-8a5c-137dd0a46e75 · outbound

This paper cites Distributed by Warner Bros.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Distributed by Warner Bros

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:b2149343c6e2444598671a0a8e7e2c80a47167c051dc5e0ea0f4bbd9b01af896

Observation dbd75caa-c90c-48b8-bc4b-f1eb4b076a1b · outbound

This paper cites Gesturediffu- clip: Gesture diffusion model with clip latents.ACM Trans.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Gesturediffu- clip: Gesture diffusion model with clip latents.ACM Trans

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:c5c2b278c370124fe56f20946a455f49572637ce8eb213c4074f44ffd9c18258

Observation d5779646-f04e-4321-b3a1-3406597830be · outbound

This paper cites wav2vec 2.0: a framework for self-supervised learning of speech representations.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens wav2vec 2.0: a framework for self-supervised learning of speech representations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:a21ca76e46c2b1c951c8851a928cdd231868996b1c7911a2a454be346014cd51

Observation c3532c69-855b-40a1-b1c8-d71ee4263583 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:5b6cc753e2d26a5629420fa5ecf06d44dfd7571f0cd365e470ee2f25c1f8ab5d

Observation a3a26bec-6bf7-4ffc-934d-f9004d3e69be · outbound

This paper cites Cohen and Dominic W.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Cohen and Dominic W

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:73cd73517c05025bbd187667c5d0441783fa10d113d8e3b12cefd181d8b3ac30

Observation 2e791cee-e7e1-48d0-a22e-f06f8b8e55d7 · outbound

This paper cites Emotional speech-driven animation with content-emotion disentangle- ment.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Emotional speech-driven animation with content-emotion disentangle- ment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:49c9bfd151807f0efcc6ccee08f47bb7d8c3031e0dc75397cc6c424763c84476

Observation 4772aacd-727b-45ab-8376-35d1485ee341 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:25:59.875366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:a52097706e46bdb6d0d329b107dc8c482ef3b8b5ae90f4e9850c03480a2dd3e5

Observation 3572fb52-23a1-45dc-8585-94bfc1396daa · outbound

This paper cites CosyV oice: A Scalable Multilin- gual Zero-shot Text-to-speech Synthesizer based on Super- vised Semantic Tokens.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens CosyV oice: A Scalable Multilin- gual Zero-shot Text-to-speech Synthesizer based on Super- vised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:2580d2eec13c1061ddeeac58d95f2b26ef1f3662c7093656b72b51461965c235

Observation ab9c697c-b568-4e20-8332-32449bde583a · outbound

This paper cites High Fidelity Neural Audio Compression.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:40e1add4aa620e270e9d4500dbd7427941fefd295cbe9f54df809b273f36693d

Observation 54215726-8171-4e03-b9e1-15f1807f9e64 · outbound

This paper cites JALI: an animator-centric viseme model for expres- sive lip synchronization.ACM Transactions on Graphics, 35 (4):1–11, 2016.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens JALI: an animator-centric viseme model for expres- sive lip synchronization.ACM Transactions on Graphics, 35 (4):1–11, 2016

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:c3ddae7b9f2bde296090f57dae9cd165e65539515c6edcba1828148b9b28a404

Observation 7c1e6fb2-4506-4646-b577-744c0bd4a6df · outbound

This paper cites Jali-driven expressive facial animation and multilin- gual speech in cyberpunk 2077.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Jali-driven expressive facial animation and multilin- gual speech in cyberpunk 2077

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:dba15a423e26b4e79c01cee9bba3c1a1ab339786bc42cbfd2559e3c69b4751d8

Observation d237f36a-daad-46c6-a01e-272f84860120 · outbound

This paper cites Papka, Sanjif Shanmugavelu, Darshan Gandhi, Hengyu Zhao, Dun Ma, Kiran Ranganath, Rick Weisner, Jiunn-yeu Chen, Yuting Yang, Natalia Vas- silieva, Bin C.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Papka, Sanjif Shanmugavelu, Darshan Gandhi, Hengyu Zhao, Dun Ma, Kiran Ranganath, Rick Weisner, Jiunn-yeu Chen, Yuting Yang, Natalia Vas- silieva, Bin C

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:0dbed339867e791ed6c5c3e0a4b1fb2d020268b8b5dd46556f85a669e005f805

Observation 3cd5da28-e73a-4841-b2a2-47e90f21f24b · outbound

This paper cites Faceformer: Speech-driven 3d facial anima- tion with transformers.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Faceformer: Speech-driven 3d facial anima- tion with transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:e4b5f5240e6d122af4b7ea13e0bfd6893d963a025b50fead0058137d1df17eb6

Observation a43b7391-d3e0-4f63-9793-c00a05313c7a · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:25:59.872963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:298df7bf4ca0a85d053b0f6a4d71b28d313d1b99476b4ab1e49901fb61c9fa3a

Observation abb55630-1a12-4b13-8d71-a9c68bac0533 · outbound

This paper cites Tiny is not small enough: High quality, low- resource facial animation through hybrid knowledge distil- lation.ACM Trans.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Tiny is not small enough: High quality, low- resource facial animation through hybrid knowledge distil- lation.ACM Trans

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:3b85fb4862e5df4f7572203451d30f8f8c75eee0ccbbe65577eb0bad375b39f8

Observation 4f6a6edd-df17-4537-864c-70369f204574 · outbound

This paper cites Classifier-free diffusion guidance, 2022.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Classifier-free diffusion guidance, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:bc14acd2b3cf4dcbebad66c8df45d373ad99fb4966424b6f696588b415172aae

Observation 05df52ea-6355-4b7f-907a-066a59850658 · outbound

This paper cites HuBERT: Self-Supervised Speech Representa- tion Learning by Masked Prediction of Hidden Units.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens HuBERT: Self-Supervised Speech Representa- tion Learning by Masked Prediction of Hidden Units

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:1b4e1138c2e34e0f88283dcc1173c8e2bb7980f60d194d73cbb3ccaa324e511d

Observation a41fa6a5-1451-481f-a4ac-bd38485d0055 · outbound

This paper cites Speed- aware audio-driven speech animation using adaptive win- dows.ACM Transactions on Graphics, 44(1):1–14, 2024.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Speed- aware audio-driven speech animation using adaptive win- dows.ACM Transactions on Graphics, 44(1):1–14, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:6a3cc8fe039cd3eb12e55a37763e20688fad2e81535f5ed2f684f8316114091b

Observation efda1815-324f-4f86-b2a5-ed21a5b461e8 · outbound

This paper cites Audio-driven facial animation by joint end- to-end learning of pose and emotion.ACM Trans.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Audio-driven facial animation by joint end- to-end learning of pose and emotion.ACM Trans

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:7913a7ad6134578dc5a7844f2f076ce3ef9d2d7746e6d5ad8e1f0c1387059429

Observation 2ca5f8fd-2dd8-44eb-8ea6-d08fca1f832f · outbound

This paper cites Audio Driven Real-Time Facial Animation for Social Telep- resence.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Audio Driven Real-Time Facial Animation for Social Telep- resence

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:0c61ed286c70c230cdc67273df3858a2aeac0302efb2fdd19cbfac8956678948

Observation 3d9d66ad-16ec-46bf-a3b5-bd40822786e5 · outbound

This paper cites Ditto: Motion-Space Diffusion for Control- lable Realtime Talking Head Synthesis.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Ditto: Motion-Space Diffusion for Control- lable Realtime Talking Head Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:db79d0839c0816169e2053a84e8c33617997f1a51403bce0491c713b89900d9f

Observation 79d377d4-9e4a-42c7-98b9-25a8f6d0c32e · outbound

This paper cites an unresolved cited work.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:60f59d3660294f1fbeca0d45b6870e2efad84f8adaa21133ed725685f72538ac

Observation 8aa55d09-f1d0-469c-b029-c07b142a810c · outbound

This paper cites an unresolved cited work.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:51b64cd93e947d70cc6c45622e77153ba97beb759b1b75993be8d724dff12c64

Observation f3410bfb-594d-426b-bda8-2721f33ee2d6 · outbound

This paper cites Medtalk: Multimodal controlled 3d facial animation with dynamic emotions by disentangled embed- ding.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Medtalk: Multimodal controlled 3d facial animation with dynamic emotions by disentangled embed- ding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:ae79a975bca967f2341807440594f50ff163e003b66c4e31fd2c31606d4a4e27

Observation 79612e0c-73df-4be6-bc9c-b3919089cbdd · outbound

This paper cites an unresolved cited work.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:4a14085a2efc035749c4686c178e1efe7c1f6a7184edcb58b88351e91a30678e

Observation 07dfb8f3-ff33-4cea-abdc-4f6a4d9a3e96 · outbound

This paper cites Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:f3a475970fc2af1344dd876773d8745c124bd504baa877475f2798ee3cf9ca7a

Observation 9be75481-cf9d-4e5e-8798-347fafcece5d · outbound

This paper cites S3: Speech, Script and Scene driven Head and Eye Animation.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens S3: Speech, Script and Scene driven Head and Eye Animation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:7ad74bfba44fe9369755841dc8186e32f1d47d817f4a97bf7a18508b99e46de6

Observation 6d82dc12-564a-4c37-b2ed-70bb7f0fd9e2 · outbound

This paper cites Model See Model Do: Speech-Driven Facial Animation with Style Control.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Model See Model Do: Speech-Driven Facial Animation with Style Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:10b7a923cb6d0555d867dd78a593df15f771b0d414a51df6610996c3c2c0851e

Observation 291209c5-c984-4049-9279-e81a58391cd1 · outbound

This paper cites VOCAL: V owel and Consonant Layering for Expres- sive Animator-Centric Singing Animation.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens VOCAL: V owel and Consonant Layering for Expres- sive Animator-Centric Singing Animation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:16243d370b6802cdc77e282a4b80a2f9980ea34090f2805b35dbe8940ae14fe5

Observation 7b52de19-93b2-47e4-8cf0-d8f8c09167df · outbound

This paper cites Emotalk: Speech-driven emotional disentanglement for 3d face anima- tion.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Emotalk: Speech-driven emotional disentanglement for 3d face anima- tion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:d93acf39645cedaf9550f4061e761c42d5e1647b087b8d3f4024544181b7d599

Observation 32263147-47ed-47dd-9f1b-642f2e349f99 · outbound

This paper cites MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:25:59.878239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:5da40cff71bf5073e0396c80996d05aacde3a996bf49a6ff82926d6867274f79

Observation 967994d3-5c40-4bf5-842d-817651597e48 · outbound

This paper cites Facediffuser: Speech-driven 3d facial animation synthesis using diffusion.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Facediffuser: Speech-driven 3d facial animation synthesis using diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:54de9cf50c361fac58948d765cfbb26665ece6a906d1728f29224b7911a02ed4

Observation 90065ea2-fb1a-4410-8037-b2c763e14fe5 · outbound

This paper cites Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.ACM Transactions on Graphics (TOG), 43(4):1–9, 2024.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.ACM Transactions on Graphics (TOG), 43(4):1–9, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:9e9e730c305a979f1b3c83e35023cfce14f9e8e5603174af82d5d4b479fa582d

Observation 52bde93e-af66-4805-b399-951e31defd30 · outbound

This paper cites Turn-taking and Backchannel Pre- diction with Acoustic and Large Language Model Fusion.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Turn-taking and Backchannel Pre- diction with Acoustic and Large Language Model Fusion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:3bf62d67436a531c5610fc975f111a3dfaee7f464b7bea14da7bc190014cbe14

Observation 44551fb8-e48a-4d2a-ae1d-92cf84f5b19e · outbound

This paper cites Mini-Omni2: Towards Open- source GPT-4o with Vision, Speech and Duplex Capabilities.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Mini-Omni2: Towards Open- source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:65cb61a24300bc5444562a885be1798dfb51ef48735e9fdb026f591b3242822c

Observation 0104c4ab-a40e-4fce-9807-77c703ecd44d · outbound

This paper cites CodeTalker: Speech- Driven 3D Facial Animation with Discrete Motion Prior.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens CodeTalker: Speech- Driven 3D Facial Animation with Discrete Motion Prior

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:a5e6c2bb2c8e0e8843247e60d68bb583708bfeffed876a54d60c521792887233

Observation 3451b55e-0307-477a-9c11-d0d79c910593 · outbound

This paper cites Qwen3-Omni Technical Report.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Qwen3-Omni Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:aef8723f874587bb92b19c8eb04fcaa1acf3e1b6612e637314783fd7242777e1

Observation 5f6bdf2e-4e14-4bb1-b8fc-ead13aa08ea3 · outbound

This paper cites SoundStream: An End- to-End Neural Audio Codec.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens SoundStream: An End- to-End Neural Audio Codec

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:185a510ad60dbc87a23f290826eb0593d653878167578e85a2023befd57e2a33

Observation 81eba8c8-9cd8-481e-9cd2-2950c9081ecf · outbound

This paper cites GLM-4-V oice: Towards Intelligent and Human-Like End-to- End Spoken Chatbot.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens GLM-4-V oice: Towards Intelligent and Human-Like End-to- End Spoken Chatbot

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:3e849f8ae0217784d2f723fabd13b6668e980a4b47595fbf5e0c64db8d0a8aa1

Observation 88007d3d-7bfe-4564-ac77-bf6a81eed46a · outbound

This paper cites SpeechGPT: Empow- ering Large Language Models with Intrinsic Cross-Modal Conversational Abilities,.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens SpeechGPT: Empow- ering Large Language Models with Intrinsic Cross-Modal Conversational Abilities,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:d9b485273639fad546d4d25499a29e77f489c24a621719014d0fc5c3357b851b

Observation 6e21ec83-4b02-455c-bbde-4dc4cebada0c · outbound

This paper cites MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling,.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:38786d590e3beae9fd583d8b84ffffc05e10addf6fa8a4409e1741d378cfd9a0

Observation bce2bbe5-805f-4ba4-b43a-2e806622f345 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T22:45:39.443440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:961cba630310326d3561ed563190239b6e9c86192f38eb9b7d3319bbaa0a44d6

Observation 6463b96a-c5fd-4f12-8615-7bbbc3e40e76 · outbound

This paper cites Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:25:59.870529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:a863891c960876a871d86fadaa7b3c3483343aa6386cdca1db5b9ad3a184ac76

Pith citing papers

No inbound Pith citation observations are available.