Pith. sign in

Paper Citation Record · LEDGER

Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2104.03502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.03502 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:23:42.651816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:34:43.114766Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 42e07762-53d0-40bd-9304-00a22760f0bc · inbound

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities cites this paper.

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:42.651816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:42.651816Z digest=sha256:75fe3f931296c2e372d72fdbc7608299e64352751de9daeed26fff0f238eb940

Observation 6840b528-b5f6-4b71-a05d-fc00a2c0dff3 · inbound

Synthetic Audio Helps for Cognitive State Tasks cites this paper.

Synthetic Audio Helps for Cognitive State Tasks Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:41:51.174050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:41:51.174050Z digest=sha256:8d7d260fd8c7e357cc2e3bfca9ee4194516cd922637f618956627627150febb1

Observation a71d5443-d719-4183-a034-43b3a7a0effe · inbound

Investigating the Impact of Word Informativeness on Speech Emotion Recognition cites this paper.

Investigating the Impact of Word Informativeness on Speech Emotion Recognition Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:15.791716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:15.791716Z digest=sha256:d49fe3031196eb4076944faf1b911259608095433944572048ad51379431eac2

Observation f15749ac-0c24-496d-aae4-a5f7766a748f · inbound

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? cites this paper.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.459509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.459509Z digest=sha256:745196c30531acd166f467f469ee87081d7b985ba83343c797b783a27a5454a8

Observation 87888b61-f0a7-4ea2-8609-c1f8381350d5 · inbound

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews cites this paper.

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:39.878332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:39.878332Z digest=sha256:cada7b23c62c938f938e3b9dd21630b57525a220ed2d05a0e806fb26ff830652

Observation 8f82d945-779b-4bf2-bb6a-5dbef9fe54ec · inbound

Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond cites this paper.

Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:23.587518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:23.587518Z digest=sha256:6199c1a8334ac880a84a4ae1ba7170aab6085de44303f4875d7235fee6751806

Observation fedc5dc6-d724-47b7-b347-6f96fc92f8c8 · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.317218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:5db741c6cf6058b8a1239d6139b920d29b033800947bb8ac28326ecf9684a295

Observation 290e3121-0c63-41c9-93dc-8c059de5046f · inbound

Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction cites this paper.

Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.849407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:58:49.423628Z digest=sha256:9556b6055141a8598511a5cba8d6991ea4909e23d044701d44eedd9c5b7378a2

Observation 1f25fb35-91ca-4465-b66b-7067feb29066 · inbound

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures cites this paper.

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.082710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:14:20.892333Z digest=sha256:a1f68155ab1c7bff44d4ab35c817775ff6246510c6cfddbd13f94d7d8771dcc7

Observation 6c93ab97-6a5e-4c93-bfe7-0cc20f230153 · inbound

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition cites this paper.

SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:24:57.691826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T04:44:59.279562Z digest=sha256:01824866aa67ba065075717c803338f26ca9b0f8a7467ec22e79926862060815

Observation 1c122ce8-1fa7-490d-8cb0-98cddfc424c3 · inbound

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study cites this paper.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:a2daf481945501eecfbda3bb0d50b59e51098c0507a07ea9199270cd46ab7664

Observation 3677867c-4038-4993-976c-4a15694b23b4 · inbound

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment cites this paper.

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:46.343211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:46.343211Z digest=sha256:9b5190c13208993bc822422a5f59f37719efeeb3a5311ef2c41935968f85885e

Observation d94f2228-de5b-4331-8f5c-64da30fd43e3 · inbound

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective cites this paper.

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:34:43.117158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-08T07:33:22.898615Z digest=sha256:81c92e24e8c3dbd422c4f9fd38fb11b5ca4c58b9c1355578632e421a8c936ae6