Pith. sign in

Paper Citation Record · LEDGER

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis

As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.13209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13209 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:47:09.073686Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f79ae570-cab8-4f83-ab06-75077cbb4ac3 · outbound

This paper cites Simulation-based learning in higher education: A meta-analysis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Simulation-based learning in higher education: A meta-analysis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.137518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.809431Z digest=sha256:d911c9ae30c50387b4155355ece600cc88411299a471c14202bdafeea83e330e

Observation 9dd10fdb-49bc-4d6e-8a30-eba6a04a0831 · outbound

This paper cites Psychological foundations of emerging technologies for teaching and learning in higher education.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Psychological foundations of emerging technologies for teaching and learning in higher education

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.116225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.815259Z digest=sha256:bb4293b07def456014c3a4ada6c67ab04bfcc69b68aedbd6cf776f5118e27c24

Observation f15c7172-b18b-41a4-8d1f-96d407ea625f · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:10.097669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.821437Z digest=sha256:f157017c641dfe54e09dcb624fc2bb6a276ed4f8f1648c0c2f8887b1e2ba3cba

Observation b8c58379-42cb-4f6a-b13a-ba795041be66 · outbound

This paper cites Designing effective training programs for investigative interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Designing effective training programs for investigative interviewers of children

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.079749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.826723Z digest=sha256:01c674327effbd782b4a68ecd3b0b69dbe5207599a1a13e12b40108b0e1074ce

Observation 7fb1a43d-1ef9-482a-9fa9-76a5447ead54 · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:10.058801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.831853Z digest=sha256:7b89be3bfef9285247a80a1058bf37049583288795219021ac46cad1b1bc6177

Observation 06bbf19c-0514-4af6-aea9-aa1f47d53b63 · outbound

This paper cites Interviewing children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Interviewing children

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.037742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.837457Z digest=sha256:c392ff8d586669ae921e20959ae8c2aa5943b3bf8471d2547a016861289cb988

Observation d63462fa-9032-4607-a731-0f2b2fd92dff · outbound

This paper cites Tell me what happened: Questioning children about abuse.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Tell me what happened: Questioning children about abuse

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.013244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.845623Z digest=sha256:2cb515282ec648574021ae1e6a123d79623a4b26391f7e628b89bb975db7b3e5

Observation 8de70706-2588-468e-9620-a250123246a0 · outbound

This paper cites An overview of mock interviews as a training tool for interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis An overview of mock interviews as a training tool for interviewers of children

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.996698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.854431Z digest=sha256:6fa85157ed97327996b82feebbcac530411034e12f838a7f3566fbe02e87410c

Observation ef11a872-8f79-49ef-a3da-61ca9fdfec2f · outbound

This paper cites Towards an ai-driven talking avatar in virtual reality for investigative interviews of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Towards an ai-driven talking avatar in virtual reality for investigative interviews of children

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.979634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.860126Z digest=sha256:ffea7bb08e12d8faa15747c2b550193fd5b7faf33bb9cc2cffd1139741cba4f6

Observation e6547aa4-dfc5-4c06-ad48-0e81540c28da · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:09.962601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.866481Z digest=sha256:6ba1113c25608345ed89dfadd1af515b095d0553b344f1214b78648a9e73a06f

Observation 459553b4-ab63-4bf2-8d48-a903b32abb04 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.871543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.871543Z digest=sha256:666f9558366cec0276be5200266916f71155ca169e7580854717bbd23b76af8c

Observation 49f82dc4-9230-4af4-b7c9-93105dbfb566 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Robust speech recognition via large-scale weak supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.876364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.876364Z digest=sha256:f1f823c958248e8ccc2c782daac72392d9d787058f9eea06a887068f58f7eb4c

Observation 9afcc583-79bf-4bae-8a57-dd3d1b84e27f · outbound

This paper cites Whisper afe for talking heads generation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Whisper afe for talking heads generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.923125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.881311Z digest=sha256:12e8c875df0f48523abce751d55fabde8bc58d8d33767973552d2e5e24933b03

Observation 1e4731ba-445f-4ebe-85c7-2beadc0f5829 · outbound

This paper cites Technological acceptance of an avatar based interview training application: The development and technological acceptance study of the avbit application., 2021.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Technological acceptance of an avatar based interview training application: The development and technological acceptance study of the avbit application., 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.906522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.886036Z digest=sha256:1866f13d2df0bd576a8645f66bf116971811df6e8f6b6785edac6d446be74157

Observation 254f0f89-aefd-464a-8222-e8ef696b67b7 · outbound

This paper cites A field assessment of child abuse investigators’ engagement with a child-avatar to develop interviewing skills.Child Abuse & Neglect, 143:106324, 2023.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis A field assessment of child abuse investigators’ engagement with a child-avatar to develop interviewing skills.Child Abuse & Neglect, 143:106324, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.889013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.890840Z digest=sha256:9484bce3325c98e9884bec4d7d3f5b35d4117a0b560bd2cfa2fe6ce3d7bb6a0c

Observation 685af020-1d76-4329-ba5f-1a0148af9b01 · outbound

This paper cites Evaluation of a comprehensive interactive training system for investigative interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Evaluation of a comprehensive interactive training system for investigative interviewers of children

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.871318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.895764Z digest=sha256:5c1586a0c1fd7c08bdffcaaa87dca58daf2d0afe361fff587e1453419e8f10e3

Observation 702d9421-bf07-49a6-9884-87ad94a959c6 · outbound

This paper cites Training in investigative interviews of children: Serious gaming paired with feedback improves interview quality.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Training in investigative interviews of children: Serious gaming paired with feedback improves interview quality

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.842122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.900547Z digest=sha256:11604de1fe1acf25728ce5642f124a77a9a8ac0a4ca1d67c91bb179164dda175

Observation 61322ef2-4dae-47ec-8f1a-1dda5755b770 · outbound

This paper cites How to prepare for conversations with children about suspicions of sexual abuse? evaluation of an interactive virtual reality training for student teachers.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis How to prepare for conversations with children about suspicions of sexual abuse? evaluation of an interactive virtual reality training for student teachers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.823380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.905396Z digest=sha256:93d3a00c0faa099fa9d872fa1f7613519ab1d9aa9841372dddde253791a40430

Observation 6d6295ca-6422-4315-b270-9bbbae55ebbd · outbound

This paper cites A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.805367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.910586Z digest=sha256:ea3d25f641a0e0d68c6c1eb5efd6339621cad215f5cd7b7660b297e94247a158

Observation c2fae4ba-10c1-41fd-8864-375aade7d8dc · outbound

This paper cites Live speech portraits: real-time photorealistic talking-head animation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Live speech portraits: real-time photorealistic talking-head animation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.784817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.915545Z digest=sha256:7f49fdda608953fdf5a2c4560deb90fcbf1394cddd0cd5ee0c11aaabcd0e469b

Observation 20e87e22-84b4-49f7-9d7e-c043b3586c9c · outbound

This paper cites Generative pre-training for speech with autoregressive predictive coding.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Generative pre-training for speech with autoregressive predictive coding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.764253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.921810Z digest=sha256:eccf5e58829d1845a34024c8df9f9a29c312a5efdffae74b44360e301729e141

Observation 968615b6-4508-44d6-aeb5-d66cba41690e · outbound

This paper cites RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.927249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.927249Z digest=sha256:56fb131717d5a5583f21d1d7d3e14558127cba42646c66ec93f4ef1a29a03575

Observation eccb74c8-ffda-4f2a-b55a-d225e16f4298 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis 3d gaussian splatting for real-time radiance field rendering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.934320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.934320Z digest=sha256:9c2162e277a90f5603648464c4fdd8febe557f3efba0b19c1a405615bce4a391

Observation 73e3db43-c7ed-4143-8bb4-e062bdbeafc8 · outbound

This paper cites GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.940014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.940014Z digest=sha256:232639dedfc9c18e230ffba21ca136cc1c2eadb3efec1b029ee25a774d475fe1

Observation b3479857-ac82-4187-914f-25a0c087d6b8 · outbound

This paper cites Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.733110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.945932Z digest=sha256:4f11454a79d6ba23123a74c40337abce19fe2fc99c40c88387855f0bb7e2414b

Observation 5d2d5425-b974-4633-b021-4a6d0777df82 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.951700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.951700Z digest=sha256:23e8c1dbc2c50b60ad568be1fd3f47b89cdfed16d9e9215a6477b5de1da0d554

Observation 340d058e-1a90-42b8-a0eb-a1ba860a06db · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.713679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.957073Z digest=sha256:4f4efc5373d7cffd8acaf279d9ea3097593f442e7696f16fed2ca512915546a9

Observation 6a0bc73a-2840-4c52-a266-a925773d2d0e · outbound

This paper cites Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.693229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.962409Z digest=sha256:feae37a5ede13193e6ede954495f28a1b543dcf874ff2b5b2c9e7f3995808471

Observation 64c5e362-c91c-447e-b9b9-abd19f84ba16 · outbound

This paper cites GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.967948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.967948Z digest=sha256:42874dc36db7c33bb52652011a3b0d54aa100ea28c08d12f6c57099899c368fa

Observation d4842ba4-5216-4a80-80a5-f2b6cf9ca685 · outbound

This paper cites R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.974918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.974918Z digest=sha256:443fa52ccb46df2207afe6ebe47fc08a45147bd6309786bfb932b91db3af4ba0

Observation 332784ed-7baa-4373-bca8-fab16268a54e · outbound

This paper cites Deep speech 2: End-to-end speech recognition in english and mandarin.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Deep speech 2: End-to-end speech recognition in english and mandarin

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.980477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.980477Z digest=sha256:14daec2895f11d4fe95dcfcbce453c1031303ed5970a44c144eed45ab0f36c38

Observation fb0395bf-b997-4d87-82f8-3aae258faa57 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.663290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.985782Z digest=sha256:633e1081c04b7729e0d82696240f7ead621f7c672e2b4987b906ce122a7a7662

Observation cf61b4c4-1e46-434f-b3a0-86b91a4766d9 · outbound

This paper cites Hubert: How much can a bad teacher benefit asr pre-training? In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6533–6537.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Hubert: How much can a bad teacher benefit asr pre-training? In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6533–6537

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.645909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:08.991193Z digest=sha256:4d46bfba5a03b5b3bbee87dc764346a9a3edbab7391dd5daa7445a0f38fd9c0e

Observation 8f0216d4-df7f-4a21-a22a-ae3c0b888dcc · outbound

This paper cites Bidirectional recurrent neural networks.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Bidirectional recurrent neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.997783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.997783Z digest=sha256:7b01d138dd8741b4407120dfdadfb7b2898f9c5982198abb4c90ba219de67117

Observation 89851c0b-c735-49f0-8656-174646668c90 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gaussian Error Linear Units (GELUs)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.003205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.003205Z digest=sha256:57abb6932409b2f0dd6f96e8d8ff414ff7bc0383ffd31f552a6fc665e3fa8b98

Observation 058cc7c1-7830-4a7c-9327-49db50e7c999 · outbound

This paper cites vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.012586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.012586Z digest=sha256:02d27b73d9bf7e8ca5ba73067c0552c88404e7406d25d230c8a5b719d895ce8b

Observation 0dcb4ce5-457d-45d7-b455-1c25e382afc7 · outbound

This paper cites Product quantization for nearest neighbor search.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Product quantization for nearest neighbor search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.018545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.018545Z digest=sha256:7a9f3526b6a8e429a3710ff9a4e5c39422485b778ab5547609ac9c1dd406fb65

Observation bb09211e-51af-44f9-bdc5-38409a41b11b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Librispeech: an asr corpus based on public domain audio books

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.026473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.026473Z digest=sha256:b202940f4e867029dc310df323eefcce4e5c9f9cadbe74eef23ba947a2f30142

Observation 07afb9dd-47a8-4890-a6f3-30820aa4e43a · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Libri-light: A benchmark for asr with limited or no supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.588803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:09.033426Z digest=sha256:c820c55eb914d9afe542f180138072c5898589a8e1e990cca0378521b7d7d86b

Observation f6ef059b-e217-4d56-a14b-a04a74f80190 · outbound

This paper cites Spanbert: Improving pre-training by representing and predicting spans.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Spanbert: Improving pre-training by representing and predicting spans

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.560601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:09.038620Z digest=sha256:1180df711277f79d7777343420f87c12f6288839735bc8204f65f0e27efbf760

Observation 3a3b8606-b8d1-438e-92b2-cead2e4d942e · outbound

This paper cites Aws polly.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Aws polly

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.527753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:09.043469Z digest=sha256:2cb20333919c77178f6aa9231225570978f77d57b7d690fae896dd3e79d724d0

Observation 25de3783-3ddb-49b7-823b-b5f12ffaaf4a · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis The unreasonable effectiveness of deep features as a perceptual metric

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.048662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.048662Z digest=sha256:d2d79d68949200e03710f0486662643c3f0c9f754f64a0291b515d719d91b2c7

Observation 35fb33ed-2fd1-432c-a89f-493632187426 · outbound

This paper cites Lip movements generation at a glance.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Lip movements generation at a glance

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.347158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:09.053891Z digest=sha256:5874c403a000d4d1abb47d5b06f4af6be57541ce4522c2043b988f641adc5a6f

Observation fc2861b9-2d0f-4d83-aeea-a5060dc8bc2d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.058524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.058524Z digest=sha256:18d0bdaa9fbe5b1c6b3b138985f15efdb708f166e2ae799dfd8c1d69e350701e

Observation feaa8353-5d57-414f-a279-d85d128132f7 · outbound

This paper cites Openface: an open source facial behavior analysis toolkit.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Openface: an open source facial behavior analysis toolkit

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.313668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:09.063648Z digest=sha256:7090dcd2b22930a943e796b212c6b4fce0aa68b1b10044f2edd1473377e61b5e

Observation 41c8fe3c-3a32-4f46-a06b-c9de70833c37 · outbound

This paper cites Out of time: automated lip sync in the wild.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Out of time: automated lip sync in the wild

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.068356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.068356Z digest=sha256:a7b4ec4a5541964f4132f5eb2a79f3d3b433c61fbd5b7045023c0790a4018b90

Observation 7fb85829-6862-4cb8-9218-679ab8d50885 · outbound

This paper cites The uncanny valley [from the field].

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis The uncanny valley [from the field]

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.272448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T16:47:09.073686Z digest=sha256:d7a7116698372f83ba23c2ced6f891d95e087af24a5c54768acf680912f1ac17

Pith citing papers

No inbound Pith citation observations are available.