Pith. sign in

Paper Citation Record · LEDGER

A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2111.02735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.02735 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:16:45.044605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:52.037436Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e54207ef-db25-4e0b-9637-4f8ba3a6ea1d · inbound

Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information cites this paper.

Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:45.044605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:45.044605Z digest=sha256:955e7bea2582f76d2f69da21898acb9eba5821090c28d63e999e6307bec5a335

Observation 3064ba90-d203-4f60-9c64-601241a5492f · inbound

Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment cites this paper.

Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:20:49.211984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:18:57.564840Z digest=sha256:acda1929bbf25a45f0d54e3f18945faba25d7d2d4eccbd63135dec26663b7c1b

Observation 8d87a066-029c-4c5f-908a-7b80a81e2f22 · inbound

Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers cites this paper.

Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:25:28.434258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:25:28.434258Z digest=sha256:434c3d4e247cb2c11b7f076ffa545ce852ef94851aa1a9cfaf62076f2810ca69

Observation 899ee35a-95c7-4794-ad69-431a93b98030 · inbound

A Dataset for Automatic Assessment of TTS Quality in Spanish cites this paper.

A Dataset for Automatic Assessment of TTS Quality in Spanish A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:26.543744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:26.543744Z digest=sha256:158f198ab57014ad812daeacbc9ae25d47d68c2a8c6c6e5da5e5b5d20ea89998

Observation e1be3bd1-be96-4ce8-886e-db345248959c · inbound

"How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases cites this paper.

"How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:56.065305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:56.065305Z digest=sha256:266bb0249db677fed0d0ccbf5626d3b50b1af6d0d9223813b7c39d0860cc610b

Observation 8ca5be82-cb43-4ad0-b3cf-8df346c28fc9 · inbound

Deep Learning Approaches for Multimodal Intent Recognition: A Survey cites this paper.

Deep Learning Approaches for Multimodal Intent Recognition: A Survey A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:34.000510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:34.000510Z digest=sha256:4d96b6842360c98aefc4e2d107013e7332c85692299a166115c7a03c68b3918e

Observation 20f435b3-9089-4f31-b9c2-5defe4d5ecfd · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:40.648163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:40.648163Z digest=sha256:9c7206f3d00dc4a7df60ee6f4dfe8b58471019b8aaf5e12313a88734deb91310

Observation e093735d-c928-47c2-a74d-e3511988e94b · inbound

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition cites this paper.

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T19:01:24.191393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:01:24.191393Z digest=sha256:44b8fd17d8ea8eb5bea6e662781e97d8bd082be1c28210cb2ae90effc0f55df9

Observation 1e3e4584-8631-45dc-bd6b-8fad767f7957 · inbound

Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition cites this paper.

Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:11:42.905106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T18:07:34.965356Z digest=sha256:1b3df4ceb60ad74229d273a694cd576f55f62b9c536231609d4c4976845c74f2

Observation c17e9a26-d61c-41d8-b06f-acc2587c7b15 · inbound

Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model cites this paper.

Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T18:22:53.721498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:22:53.721498Z digest=sha256:e0d93f5faee49e50d1a5ccea9548132c8612aaff228c2f2f40081ca4753f3241

Observation f0ef815e-bd07-4ce5-b522-f86c01704c4e · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.330890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:796a0dd5d4af03831d83101195aa461acda8201cc6df1c8c7a93f2affe648d9c

Observation b8aecfa9-82b3-4b4e-a113-30e3e51a9ede · inbound

Selective Capability Unlearning in End-to-End Spoken Language Understanding cites this paper.

Selective Capability Unlearning in End-to-End Spoken Language Understanding A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.102746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:54:46.225878Z digest=sha256:1ede67288617b3ada5c50701530c139b3af88afee78660c82bc73e62dd857a7b

Observation 2df84290-3037-4920-8856-03cad9dd6cfc · inbound

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL cites this paper.

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.502926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:17:41.860360Z digest=sha256:88d3efb9957d7e78c4b18142c9aff153ab201f260130a6b0940d35a9a55fbd7b

Observation 95228737-b994-48b2-968a-b13813d00a89 · inbound

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition cites this paper.

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:24:57.565737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T04:44:59.279562Z digest=sha256:25d1404dd3e124ae05a607f574f606caa3f467156f46e717d535495adc09095e

Observation 5890da5e-07d5-4cb0-a070-6e64d930fa15 · inbound

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective cites this paper.

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:34:43.128546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T07:33:22.898615Z digest=sha256:6aa37f2584442818ad1f756a5cbceea1fdb0c3b065a8a1530f3103f04ab59e5d

Observation 2b030a15-2b87-45b2-8590-d35b6fbd9979 · inbound

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts cites this paper.

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:52.061776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T01:53:00.127636Z digest=sha256:ccaa2be92fae47982d6add7efbf603d32ef6e4d9c2323d4809930ade943618fa