Pith. sign in

Paper Citation Record · LEDGER

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis

As of 21 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.13209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13209 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:47:09.073686Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f79ae570-cab8-4f83-ab06-75077cbb4ac3 · outbound

This paper cites Simulation-based learning in higher education: A meta-analysis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Simulation-based learning in higher education: A meta-analysis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.137518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.809431Z digest=sha256:caa0bc3b4e764d5ca4270867a9c1e1bbaa281f8dd7415d398e25afad7e7f257f

Observation 9dd10fdb-49bc-4d6e-8a30-eba6a04a0831 · outbound

This paper cites Psychological foundations of emerging technologies for teaching and learning in higher education.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Psychological foundations of emerging technologies for teaching and learning in higher education

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.116225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.815259Z digest=sha256:3933e6bb9469357dad69ae7e497895ba9192542cffc44fbf1f693c4688df959f

Observation f15c7172-b18b-41a4-8d1f-96d407ea625f · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:10.097669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.821437Z digest=sha256:e06dc087f6e13178c5b90e7f25bf62c9a4f4536526cc30fb80bfe1c246dc479a

Observation b8c58379-42cb-4f6a-b13a-ba795041be66 · outbound

This paper cites Designing effective training programs for investigative interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Designing effective training programs for investigative interviewers of children

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.079749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.826723Z digest=sha256:00dfee7c1c7caef943440d21dfb77298262afc669fa1567b15b0ffb8f7873624

Observation 7fb1a43d-1ef9-482a-9fa9-76a5447ead54 · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:10.058801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.831853Z digest=sha256:3ea0cc0e030b8078d0af9ae237728b809774a965818ef99f51e75b74710b5af5

Observation 06bbf19c-0514-4af6-aea9-aa1f47d53b63 · outbound

This paper cites Interviewing children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Interviewing children

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.037742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.837457Z digest=sha256:0192903d4bf7070ecca3190c6f01e6241bc5736fb0fb60c0d6aa851744bc01b8

Observation d63462fa-9032-4607-a731-0f2b2fd92dff · outbound

This paper cites Tell me what happened: Questioning children about abuse.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Tell me what happened: Questioning children about abuse

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.013244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.845623Z digest=sha256:988a65726c6b8cfbc98a1c0680206b7ac2132771896bd8ad738bd983d0d5331a

Observation 8de70706-2588-468e-9620-a250123246a0 · outbound

This paper cites An overview of mock interviews as a training tool for interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis An overview of mock interviews as a training tool for interviewers of children

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.996698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.854431Z digest=sha256:efa4f0ba295bc929a11c3d4d12ac1008dcf608d5a5b26b7bd8fc95d4d0bd514b

Observation ef11a872-8f79-49ef-a3da-61ca9fdfec2f · outbound

This paper cites Towards an ai-driven talking avatar in virtual reality for investigative interviews of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Towards an ai-driven talking avatar in virtual reality for investigative interviews of children

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.979634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.860126Z digest=sha256:6f2796dab5b3f1bf0e9cc9a385c8299ea6410ff4352b198b9c23eea455c571aa

Observation e6547aa4-dfc5-4c06-ad48-0e81540c28da · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:09.962601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.866481Z digest=sha256:e4185ac9dc8bc6adda0c1a85740424791bc08b00692f00910f3e06c60f9a86f1

Observation 459553b4-ab63-4bf2-8d48-a903b32abb04 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.871543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.871543Z digest=sha256:666f9558366cec0276be5200266916f71155ca169e7580854717bbd23b76af8c

Observation 49f82dc4-9230-4af4-b7c9-93105dbfb566 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Robust speech recognition via large-scale weak supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.876364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.876364Z digest=sha256:f1f823c958248e8ccc2c782daac72392d9d787058f9eea06a887068f58f7eb4c

Observation 9afcc583-79bf-4bae-8a57-dd3d1b84e27f · outbound

This paper cites Whisper afe for talking heads generation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Whisper afe for talking heads generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.923125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.881311Z digest=sha256:5c9d9b74b1337c8e892689adfb35f35284e94b6e746e9daa25a387020ef364af

Observation 1e4731ba-445f-4ebe-85c7-2beadc0f5829 · outbound

This paper cites Technological acceptance of an avatar based interview training application: The development and technological acceptance study of the avbit application., 2021.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Technological acceptance of an avatar based interview training application: The development and technological acceptance study of the avbit application., 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.906522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.886036Z digest=sha256:be507a0872cfb2e45a53405dc5f2464e8bb137cf612bbce20c3ecfb835d38ddf

Observation 254f0f89-aefd-464a-8222-e8ef696b67b7 · outbound

This paper cites A field assessment of child abuse investigators’ engagement with a child-avatar to develop interviewing skills.Child Abuse & Neglect, 143:106324, 2023.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis A field assessment of child abuse investigators’ engagement with a child-avatar to develop interviewing skills.Child Abuse & Neglect, 143:106324, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.889013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.890840Z digest=sha256:c7974448a4c349c7703c709f71a7e83f6925c37f5c2249eb73761815b3aeb4a6

Observation 685af020-1d76-4329-ba5f-1a0148af9b01 · outbound

This paper cites Evaluation of a comprehensive interactive training system for investigative interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Evaluation of a comprehensive interactive training system for investigative interviewers of children

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.871318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.895764Z digest=sha256:cd8f3013c8e79be3571f5fd5138e5fa59ded51bf14bf94d1041cff38c3570ca7

Observation 702d9421-bf07-49a6-9884-87ad94a959c6 · outbound

This paper cites Training in investigative interviews of children: Serious gaming paired with feedback improves interview quality.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Training in investigative interviews of children: Serious gaming paired with feedback improves interview quality

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.842122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.900547Z digest=sha256:62a3aaa5ecfcc46bd074634bc7c873b95215c43a6ff5e8d94d758e514c70e32f

Observation 61322ef2-4dae-47ec-8f1a-1dda5755b770 · outbound

This paper cites How to prepare for conversations with children about suspicions of sexual abuse? evaluation of an interactive virtual reality training for student teachers.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis How to prepare for conversations with children about suspicions of sexual abuse? evaluation of an interactive virtual reality training for student teachers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.823380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.905396Z digest=sha256:ccea69ef916355dc34c8a5db7c08a3c108a56fac2cb1e63a8157eb5c5cd84037

Observation 6d6295ca-6422-4315-b270-9bbbae55ebbd · outbound

This paper cites A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.805367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.910586Z digest=sha256:0f7980114d6249be14c1e15e2c8d168e0d933d9dd61433b36d73669cf025cb95

Observation c2fae4ba-10c1-41fd-8864-375aade7d8dc · outbound

This paper cites Live speech portraits: real-time photorealistic talking-head animation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Live speech portraits: real-time photorealistic talking-head animation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.784817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.915545Z digest=sha256:1ec5bc6a0a04a3bf73096caa6db3baf30cda707139668412365c165848a940cb

Observation 20e87e22-84b4-49f7-9d7e-c043b3586c9c · outbound

This paper cites Generative pre-training for speech with autoregressive predictive coding.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Generative pre-training for speech with autoregressive predictive coding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.764253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.921810Z digest=sha256:ecf8461acacb22fb87c8b20eaec3007ebba21eb5b522f98a41a5008acc124d0a

Observation 968615b6-4508-44d6-aeb5-d66cba41690e · outbound

This paper cites RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.927249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.927249Z digest=sha256:56fb131717d5a5583f21d1d7d3e14558127cba42646c66ec93f4ef1a29a03575

Observation eccb74c8-ffda-4f2a-b55a-d225e16f4298 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis 3d gaussian splatting for real-time radiance field rendering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.934320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.934320Z digest=sha256:9c2162e277a90f5603648464c4fdd8febe557f3efba0b19c1a405615bce4a391

Observation 73e3db43-c7ed-4143-8bb4-e062bdbeafc8 · outbound

This paper cites GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.940014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.940014Z digest=sha256:232639dedfc9c18e230ffba21ca136cc1c2eadb3efec1b029ee25a774d475fe1

Observation b3479857-ac82-4187-914f-25a0c087d6b8 · outbound

This paper cites Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.733110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.945932Z digest=sha256:db3045fd8ba94c456e343584aea72477dbc17c85befdf014a9861201fbc9dcce

Observation 5d2d5425-b974-4633-b021-4a6d0777df82 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.951700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.951700Z digest=sha256:23e8c1dbc2c50b60ad568be1fd3f47b89cdfed16d9e9215a6477b5de1da0d554

Observation 340d058e-1a90-42b8-a0eb-a1ba860a06db · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.713679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.957073Z digest=sha256:f16a40cca716e01939bf6f421b93e900dc104006a628a20f4f1c72e22fb6162f

Observation 6a0bc73a-2840-4c52-a266-a925773d2d0e · outbound

This paper cites Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.693229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.962409Z digest=sha256:e980d5331a0c20c1130bba1e120b8a9b7f6bb9abe545e341f980d9534763a650

Observation 64c5e362-c91c-447e-b9b9-abd19f84ba16 · outbound

This paper cites GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.967948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.967948Z digest=sha256:42874dc36db7c33bb52652011a3b0d54aa100ea28c08d12f6c57099899c368fa

Observation d4842ba4-5216-4a80-80a5-f2b6cf9ca685 · outbound

This paper cites R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.974918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.974918Z digest=sha256:443fa52ccb46df2207afe6ebe47fc08a45147bd6309786bfb932b91db3af4ba0

Observation 332784ed-7baa-4373-bca8-fab16268a54e · outbound

This paper cites Deep speech 2: End-to-end speech recognition in english and mandarin.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Deep speech 2: End-to-end speech recognition in english and mandarin

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.980477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.980477Z digest=sha256:14daec2895f11d4fe95dcfcbce453c1031303ed5970a44c144eed45ab0f36c38

Observation fb0395bf-b997-4d87-82f8-3aae258faa57 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.663290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.985782Z digest=sha256:5068139474087d4c51cfc5f0b1fae98e07500535bc147afd63875b935d44d4ef

Observation cf61b4c4-1e46-434f-b3a0-86b91a4766d9 · outbound

This paper cites Hubert: How much can a bad teacher benefit asr pre-training? In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6533–6537.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Hubert: How much can a bad teacher benefit asr pre-training? In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6533–6537

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.645909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:08.991193Z digest=sha256:fc869149bd190f27adcc3ed58823a7ea216b6409c74e9d2e1747e112b7c58717

Observation 8f0216d4-df7f-4a21-a22a-ae3c0b888dcc · outbound

This paper cites Bidirectional recurrent neural networks.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Bidirectional recurrent neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.997783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.997783Z digest=sha256:7b01d138dd8741b4407120dfdadfb7b2898f9c5982198abb4c90ba219de67117

Observation 89851c0b-c735-49f0-8656-174646668c90 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gaussian Error Linear Units (GELUs)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.003205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.003205Z digest=sha256:57abb6932409b2f0dd6f96e8d8ff414ff7bc0383ffd31f552a6fc665e3fa8b98

Observation 058cc7c1-7830-4a7c-9327-49db50e7c999 · outbound

This paper cites vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.012586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.012586Z digest=sha256:02d27b73d9bf7e8ca5ba73067c0552c88404e7406d25d230c8a5b719d895ce8b

Observation 0dcb4ce5-457d-45d7-b455-1c25e382afc7 · outbound

This paper cites Product quantization for nearest neighbor search.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Product quantization for nearest neighbor search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.018545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.018545Z digest=sha256:7a9f3526b6a8e429a3710ff9a4e5c39422485b778ab5547609ac9c1dd406fb65

Observation bb09211e-51af-44f9-bdc5-38409a41b11b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Librispeech: an asr corpus based on public domain audio books

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.026473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.026473Z digest=sha256:b202940f4e867029dc310df323eefcce4e5c9f9cadbe74eef23ba947a2f30142

Observation 07afb9dd-47a8-4890-a6f3-30820aa4e43a · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Libri-light: A benchmark for asr with limited or no supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.588803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:09.033426Z digest=sha256:32efbcaa16f938acb070ab44f091a91e65374b46ce74f29dff3c68132a620d16

Observation f6ef059b-e217-4d56-a14b-a04a74f80190 · outbound

This paper cites Spanbert: Improving pre-training by representing and predicting spans.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Spanbert: Improving pre-training by representing and predicting spans

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.560601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:09.038620Z digest=sha256:220273ec93edbdda60da1bd30038b71159f4e7bb8c667c3f07b0ef94378056cd

Observation 3a3b8606-b8d1-438e-92b2-cead2e4d942e · outbound

This paper cites Aws polly.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Aws polly

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.527753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:09.043469Z digest=sha256:5fbc923d2c84046b7facbc9d4f546e1b6ea21541be6473d636df73cd3d4d924c

Observation 25de3783-3ddb-49b7-823b-b5f12ffaaf4a · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis The unreasonable effectiveness of deep features as a perceptual metric

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.048662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.048662Z digest=sha256:d2d79d68949200e03710f0486662643c3f0c9f754f64a0291b515d719d91b2c7

Observation 35fb33ed-2fd1-432c-a89f-493632187426 · outbound

This paper cites Lip movements generation at a glance.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Lip movements generation at a glance

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.347158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:09.053891Z digest=sha256:daac1ed132084273fba577eececbe53fe1ca8a524c14aa849a22b9db81a6ae24

Observation fc2861b9-2d0f-4d83-aeea-a5060dc8bc2d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.058524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.058524Z digest=sha256:18d0bdaa9fbe5b1c6b3b138985f15efdb708f166e2ae799dfd8c1d69e350701e

Observation feaa8353-5d57-414f-a279-d85d128132f7 · outbound

This paper cites Openface: an open source facial behavior analysis toolkit.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Openface: an open source facial behavior analysis toolkit

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.313668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:09.063648Z digest=sha256:7496172f16d66a2bd692f0cb6d1bb13310e3a8188efc4d37ddba06112d779cc7

Observation 41c8fe3c-3a32-4f46-a06b-c9de70833c37 · outbound

This paper cites Out of time: automated lip sync in the wild.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Out of time: automated lip sync in the wild

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.068356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.068356Z digest=sha256:a7b4ec4a5541964f4132f5eb2a79f3d3b433c61fbd5b7045023c0790a4018b90

Observation 7fb85829-6862-4cb8-9218-679ab8d50885 · outbound

This paper cites The uncanny valley [from the field].

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis The uncanny valley [from the field]

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.272448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T16:47:09.073686Z digest=sha256:d44649cbd6c0f606ea655aa676053938560cd2ad2bf224f009a06ab580929417

Pith citing papers

No inbound Pith citation observations are available.