Pith. sign in

Paper Citation Record · LEDGER

Learning Emotion-Invariant Speaker Representations for Speaker Verification

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.18498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18498 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:33:14.007487Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:33:13.912562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:33:14.053965Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d272e83d-db89-464e-a96d-f5c90b0553ec · outbound

This paper cites Currently, SV systems that use low- dimensional speaker representations extracted from deep learning- based speaker encoders have become the dominant approach in this field.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Currently, SV systems that use low- dimensional speaker representations extracted from deep learning- based speaker encoders have become the dominant approach in this field

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.358571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.908650Z digest=sha256:e2ec444372bcfe71bb059f32af0f26ec5a2dd03b4bdaa4ab350645f7f1bf9596

Observation 1c53738c-ad4d-4cb8-b7b5-c8af21ec7dfb · outbound

This paper cites Learning Emotion-Invariant Speaker Representations for Speaker Verification.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Learning Emotion-Invariant Speaker Representations for Speaker Verification

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:33:14.058991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.912562Z digest=sha256:89b98270d1a811a290930c8b2c05d4e3b7f61f8c46540232e32a8f41ee7d9bf4

Observation ef6aa4ac-7b3b-43de-9116-967227f13384 · outbound

This paper cites Datasets Our models are first pre-trained on the V oxCeleb [22] and then fine- tuned on the Dusha [18] dataset to evaluate the performance of SV.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Datasets Our models are first pre-trained on the V oxCeleb [22] and then fine- tuned on the Dusha [18] dataset to evaluate the performance of SV

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.346234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.916593Z digest=sha256:e228dfb660b8ed87c573ada11be8c6cac833615b099b1b525f1a52db7beb7085

Observation e4b5f3c6-29e9-4647-b9dd-79d75c8cae9a · outbound

This paper cites Performance of the baseline system The performance of the pre-trained model on V oxCeleb1 is presented in Table 3, with the results of the ECAPA model referenced from [3].

Learning Emotion-Invariant Speaker Representations for Speaker Verification Performance of the baseline system The performance of the pre-trained model on V oxCeleb1 is presented in Table 3, with the results of the ECAPA model referenced from [3]

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:33:14.335678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.920331Z digest=sha256:e4156867af7189060ff583231d7c24a18917d5f5415f3bc0b527ae7b44d2f601

Observation 51d46830-4f89-40c2-9dd7-eb89b07fa94e · outbound

This paper cites We first verified that emotional utterances degrade SV performance, with cross-emotion test trails performed worse than same-emotion test trails.

Learning Emotion-Invariant Speaker Representations for Speaker Verification We first verified that emotional utterances degrade SV performance, with cross-emotion test trails performed worse than same-emotion test trails

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.323742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.923606Z digest=sha256:9123aa8509bb7f29300c317e1596c475a4e84f1ba79028993af07a35aa0ec1fb

Observation e6b4887c-d265-4784-9cd3-313a6928f3c6 · outbound

This paper cites X-vectors: Robust dnn em- beddings for speaker recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification X-vectors: Robust dnn em- beddings for speaker recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.312011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.926783Z digest=sha256:8482803e8de511a283e0a161429e7b008ad6b8f3bcd2af4e39861d912e056447

Observation 4c2a3cbe-4731-48f2-836e-011dc408225f · outbound

This paper cites BUT System Description to VoxCeleb Speaker Recognition Challenge 2019.

Learning Emotion-Invariant Speaker Representations for Speaker Verification BUT System Description to VoxCeleb Speaker Recognition Challenge 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.929697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.929697Z digest=sha256:459668aeca50f292bd2bf8786ab7f97f4311823473bd42aa5f3e8497884bb83a

Observation acb2c15c-c308-4464-85c9-2b6052663fdb · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.300839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.933411Z digest=sha256:1f86e0b725b3138d6a7802193d50d775203450882f5128170c277fe40a398af0

Observation 451a8d6f-545d-4d5f-966b-bf4bb031e276 · outbound

This paper cites MFA-Conformer: Multi-scale Feature Aggregation Con- former for Automatic Speaker Verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification MFA-Conformer: Multi-scale Feature Aggregation Con- former for Automatic Speaker Verification,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.289827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.937296Z digest=sha256:43ee0c10e7e8aa9acaf52013e9991b6e8426218b4ce5b881a6b623bddd92a222

Observation 662e5a3b-e0d7-424e-bc6b-256c079af3e4 · outbound

This paper cites Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.279108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.940222Z digest=sha256:6fbb9910f4a87f55303d741aae4f5016affe277caa14724dcf1019da5e9b419b

Observation 447260bd-c7cd-4507-916b-ec3ed4f55f6a · outbound

This paper cites In Defence of Metric Learn- ing for Speaker Recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification In Defence of Metric Learn- ing for Speaker Recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.267645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.944520Z digest=sha256:117c874c5343d4d3887aa2a848114bf7e6fecbf8a0baf121729cb54700b013f1

Observation 49ed5cc9-3b10-4dea-a8fb-32e56cc92aeb · outbound

This paper cites Multi-query multi-head attention pooling and inter-topk penalty for speaker verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Multi-query multi-head attention pooling and inter-topk penalty for speaker verification,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.256580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.948506Z digest=sha256:2951a87b50ec89c41fa39950ab61e63348e4d56c24260a366b1c04318f09ad07

Observation ca578001-9df4-4d34-95a1-ec41e6d25e45 · outbound

This paper cites Explor- ing binary classification loss for speaker verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Explor- ing binary classification loss for speaker verification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.242783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.952200Z digest=sha256:6881bfa678226793d17550f253fa3e7f9a5fee2f225b502131ea0b996ed46df4

Observation 6cf5c2e2-72ec-481d-a0c4-de75ad522a61 · outbound

This paper cites Nplda: A deep neural plda model for speaker verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Nplda: A deep neural plda model for speaker verification,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.231624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.956178Z digest=sha256:fffd43cecf4924400fc799bd3a6522c59d5ae664e507e7483ed57f31f87fcd2c

Observation ca1df279-d672-42a6-8fad-01f644036535 · outbound

This paper cites Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.219413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.959639Z digest=sha256:47adb122ea5b7dce5c875cc75938b907f2f369e16112edce55ce3840a9ba3d3d

Observation cc0dba39-3f10-4b23-a105-ad9b91c58656 · outbound

This paper cites Attention back-end for automatic speaker verification with multiple enrollment utterances,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Attention back-end for automatic speaker verification with multiple enrollment utterances,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.207711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.962936Z digest=sha256:b8dba8c920a1e019b50bfc568e14f7dce155471c4beeaea1f8802ff7c1b6f92f

Observation 8fb1c116-b121-4d24-ac15-c9e94d28bc45 · outbound

This paper cites Prob- abilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddings,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Prob- abilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddings,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.196784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.966314Z digest=sha256:682ef8c9a298a744deb4aadb151731d1a894be36c1150f926b2e34a3227afe6c

Observation 7a8bbea9-f537-430f-9007-e5b796bfa35a · outbound

This paper cites A study of speaker verification performance with expressive speech,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification A study of speaker verification performance with expressive speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.185802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.968866Z digest=sha256:ba660dbff83baf92108b381989402a2aedcee1377fddad27f58abef48163cce2

Observation 975d0073-5454-4f20-b597-0ea549de3dd3 · outbound

This paper cites x-vectors meet emotions: A study on dependencies between emotion and speaker recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification x-vectors meet emotions: A study on dependencies between emotion and speaker recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.175812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.971447Z digest=sha256:5ea03538a2b7c92398880f3dafd6881fff5061c2d583a74c2e4e76ae13aedca6

Observation 355036ba-0493-4a31-888b-0724812840b3 · outbound

This paper cites Emotion attribute projection for speaker recognition on emo- tional speech,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Emotion attribute projection for speaker recognition on emo- tional speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.163436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.974634Z digest=sha256:b67399af054a6cf2347dfa78051d1e4896c3e98582e682d850e77afc9308fd9e

Observation cf971125-39da-41ee-9c76-acbc9fee0ae7 · outbound

This paper cites Segment- Level Effects of Gender, Nationality and Emotion Informa- tion on Text-Independent Speaker Verification,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Segment- Level Effects of Gender, Nationality and Emotion Informa- tion on Text-Independent Speaker Verification,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.151537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.977992Z digest=sha256:40fd6d0872eea45d547b968337a3f4d51546211508053df6589b8ca85ed90715

Observation 0a3ed5c0-f612-479b-98ef-9108cdf3f6e5 · outbound

This paper cites Instance-based Temporal Normalization for Speaker Verifi- cation,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Instance-based Temporal Normalization for Speaker Verifi- cation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.139636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.981024Z digest=sha256:1f838fa83ee39afb486600f328357e6f496b18ca4a273caa1935518fbf54667d

Observation 37549ff7-4e44-4184-a0a5-a54e562cd1f2 · outbound

This paper cites Hy- brid Dataset for Speech Emotion Recognition in Russian Lan- guage,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Hy- brid Dataset for Speech Emotion Recognition in Russian Lan- guage,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.128075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.984661Z digest=sha256:cd40bf681d9281cfd079b6b2c0e7eab1b3473075040d1d78c37b43f61de010eb

Observation b7d2bb8d-0695-40f2-81fd-39aa33242bf7 · outbound

This paper cites Copypaste: An augmentation method for speech emotion recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Copypaste: An augmentation method for speech emotion recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.116391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.988353Z digest=sha256:c525b4970bf2bb34154596000882dfdf9d7fc90c22681c59fac27222af035851

Observation 992077d4-9643-4452-af41-65fa8401c116 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recogni- tion,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Arcface: Additive angular margin loss for deep face recogni- tion,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.991313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.991313Z digest=sha256:993cc38e5f97416b93d9973710b96fd8cf0e1899293924edeefd14bfced9cf9d

Observation 32601c19-731e-4846-9047-ba2e1f0537ab · outbound

This paper cites Sur- vey on speech emotion recognition: Features, classification schemes, and databases,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Sur- vey on speech emotion recognition: Features, classification schemes, and databases,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.100109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.993871Z digest=sha256:aa8be951a75176c7c53c043d0ebd9871698cdf5ac9a713bc9580cb464fc395ee

Observation e8787bbb-9e0f-4918-a1b3-ed414dc1f771 · outbound

This paper cites V oxceleb: Large-scale speaker verification in the wild,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification V oxceleb: Large-scale speaker verification in the wild,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.090707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.996952Z digest=sha256:6e4f7012f77b6edb0457baed9e50049acb5735a3133dc6360f1e0474b2dac0e2

Observation 9fe6b536-cc38-4501-b8de-3d7e303c9a97 · outbound

This paper cites Re- visiting the statistics pooling layer in deep speaker embedding learning,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Re- visiting the statistics pooling layer in deep speaker embedding learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.080892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.999565Z digest=sha256:7b5cf2a49084fc91e323089862cbd841de0935d5b3473f888cc769583c421e8d

Observation da5662a4-e6e1-4ce4-a827-e6df1c00a14b · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

Learning Emotion-Invariant Speaker Representations for Speaker Verification MUSAN: A Music, Speech, and Noise Corpus

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:14.003406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:14.003406Z digest=sha256:5080dd6d631d838c849526186cc690447d55cf85feb499e0e8ea1b2bf5a95988

Observation 647e1a2e-06b3-48ff-ac2e-916d31a40bdc · outbound

This paper cites A study on data augmen- tation of reverberant speech for robust speech recognition,.

Learning Emotion-Invariant Speaker Representations for Speaker Verification A study on data augmen- tation of reverberant speech for robust speech recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:33:14.070342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:14.007487Z digest=sha256:70334bdd5c64d4607b71b6da1c83d513a03923eade210d940af209ea541d5921

Pith citing papers

Observation 1c53738c-ad4d-4cb8-b7b5-c8af21ec7dfb · inbound

Learning Emotion-Invariant Speaker Representations for Speaker Verification cites this paper.

Learning Emotion-Invariant Speaker Representations for Speaker Verification Learning Emotion-Invariant Speaker Representations for Speaker Verification

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:33:14.058991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:33:13.912562Z digest=sha256:89b98270d1a811a290930c8b2c05d4e3b7f61f8c46540232e32a8f41ee7d9bf4