Pith. sign in

Paper Citation Record · LEDGER

Speech Separation using Neural Audio Codecs with Embedding Loss

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2411.17998.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17998 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:41:45.504981Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:48:34.841419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T16:46:37.685528Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69f94909-8948-4cc1-9db6-b80182fe6a11 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

Speech Separation using Neural Audio Codecs with Embedding Loss Some experiments on the recognition of speech, with one and with two ears,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.333593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.333593Z digest=sha256:6164db0a41b6898b61aecda23a7056e94bc7cbeab1812116aa21074abbd4b458

Observation e92cd3e9-9beb-480c-bd97-e27a7b691490 · outbound

This paper cites Atten- tion is all you need in speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Atten- tion is all you need in speech separation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.067206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.340401Z digest=sha256:d399ffe2055b2e03b25d6e2eb8432b1d5afcc69b2ce2d0ecfe0a806c2bdc3f7f

Observation ca44b4b5-c246-46a6-90e4-01c890318825 · outbound

This paper cites Spgm: Prioritizing local features for enhanced speech separation performance,.

Speech Separation using Neural Audio Codecs with Embedding Loss Spgm: Prioritizing local features for enhanced speech separation performance,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.049025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.345659Z digest=sha256:c2d8dd2487dcba63761c74e94d1c2744827b19cfa1b924c7633eee47955a1d8a

Observation d76f519d-8cbc-4d7c-ba3b-07822213962d · outbound

This paper cites Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.031180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.351303Z digest=sha256:b63e6f4565f2f3bd5f2ba25d3b9de09652a4d322e887acf3eb4e935d2f99de1f

Observation 599e8f6b-a545-44de-845a-4bcaeb6c6ede · outbound

This paper cites Gass: Generalizing audio source separation with large-scale data,.

Speech Separation using Neural Audio Codecs with Embedding Loss Gass: Generalizing audio source separation with large-scale data,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.011198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.356597Z digest=sha256:5891f416d4df612f55472e13deb59a243714d58eb355680b46c3960d11acab1a

Observation 2be18dab-e38d-47ce-bec7-d7f590469310 · outbound

This paper cites Exploring self-attention mechanisms for speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Exploring self-attention mechanisms for speech separation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.993979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.363038Z digest=sha256:d2cf107e3222c869bf887a38427ff428b7b89d41663f4a6c6c65cfc63994d180

Observation b0a861b9-44e3-463f-b060-bfd07c4d0cb3 · outbound

This paper cites TF- GridNet: Integrating full- and sub-band modeling for speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss TF- GridNet: Integrating full- and sub-band modeling for speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.974904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.369277Z digest=sha256:78cb5f57063112dbce46b6ce5914a864ea7f515c40252c9fa763a3c807389dd3

Observation fccdc939-376c-4d7b-94e1-9c641505d6a9 · outbound

This paper cites A neural state-space model approach to efficient speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss A neural state-space model approach to efficient speech separation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.956269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.374305Z digest=sha256:7f60835fa74aca7034e1750daa72d1273802ce37735a4beb38b9e9fe25c497c1

Observation df33a327-e700-4751-b367-6dd6c643002a · outbound

This paper cites Permutation invariant training of deep models for speaker-independent multi-talker speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Permutation invariant training of deep models for speaker-independent multi-talker speech separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.939179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.381233Z digest=sha256:ea2bda4236ae2b3c33b9af395a80a8d80fb92081abc56bafa6c5e437be533e4a

Observation e8243670-d1ec-49ac-9c85-5f5b5b4e546f · outbound

This paper cites Sdr–half-baked or well done?.

Speech Separation using Neural Audio Codecs with Embedding Loss Sdr–half-baked or well done?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.386290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.386290Z digest=sha256:c717a7c8eaa441fb14cb7a8673a54e1a02da6c668e284c8704e6889b2f4ce5b1

Observation d59cce01-1d76-4aa2-8e8a-43250db14d6b · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Speech Separation using Neural Audio Codecs with Embedding Loss Soundstream: An end-to-end neural audio codec,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.391032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.391032Z digest=sha256:8c12220a37b20293f9812820643cbb658106134ed38124e0150bc36f479b2c32

Observation 117c8613-a45b-4ed7-b081-799e295d7930 · outbound

This paper cites High Fidelity Neural Audio Compression.

Speech Separation using Neural Audio Codecs with Embedding Loss High Fidelity Neural Audio Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.395769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.395769Z digest=sha256:816985ad6dc103372142ab9050b425570205395d98efff89a5c8fec64d58aadc

Observation 06f22c58-3e30-4455-ba24-096d212cf527 · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

Speech Separation using Neural Audio Codecs with Embedding Loss High- fidelity audio compression with improved rvqgan,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.401593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.401593Z digest=sha256:bc7c42bd2a4a20a334ea01afbcdfb45a25b6ceccb18edb60710efc0c47522893

Observation 3b162e68-e8e0-4646-a278-0daec62c0e03 · outbound

This paper cites Vector quantization,.

Speech Separation using Neural Audio Codecs with Embedding Loss Vector quantization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.885444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.407138Z digest=sha256:f237f4defdbb52e835abdbe27bbf07072525f95ef7d6ddd22940194faf169426

Observation 1a76d6cc-4751-466a-8922-2d176bc48e85 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

Speech Separation using Neural Audio Codecs with Embedding Loss VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.414147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.414147Z digest=sha256:cfa9dd14af6f12453b03c2f8ea5dc98a5236b391a0be3e3861a950f78b0dfc64

Observation 10a6b267-3f43-42bb-ac11-4d53c1e4dbd9 · outbound

This paper cites Exploring the limits of decoder-only models trained on public speech recognition corpora.

Speech Separation using Neural Audio Codecs with Embedding Loss Exploring the limits of decoder-only models trained on public speech recognition corpora

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.420271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.420271Z digest=sha256:e05f633937a385d650e9017c03ae7578b0875b120c16d092ecf07f1c72bd23f0

Observation 826a5a72-243c-4839-8519-bb264089a574 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Speech Separation using Neural Audio Codecs with Embedding Loss Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.865570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.426939Z digest=sha256:35f8a23cb3a731e478845a2ca9541084ad7d6cefde4b432eb19a6a1c51d3becf

Observation e0a27129-931a-4f22-92d5-d9946d014e29 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

Speech Separation using Neural Audio Codecs with Embedding Loss SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.432669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.432669Z digest=sha256:3683be16440b086cc92fa1063f1cf0950c377e3d82af3405a7747b7ad3fa1d60

Observation 95d00d77-5065-4bec-bf4f-dfb640e74074 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

Speech Separation using Neural Audio Codecs with Embedding Loss Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.843407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.439015Z digest=sha256:c952c16ca01af4a274a9dc0602cacc13a9a6e37988c9de23067a9286e75aba65

Observation d14e5e27-71bf-4831-982e-9937e0dfdf04 · outbound

This paper cites The Interspeech 2024 challenge on speech processing using discrete units,.

Speech Separation using Neural Audio Codecs with Embedding Loss The Interspeech 2024 challenge on speech processing using discrete units,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.825518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.446408Z digest=sha256:47294413afe53f1c1022e784a32e707147695642101521687ef3929102c8bef7

Observation adcd272e-f277-4ebe-b3df-d0d60ae0a7c5 · outbound

This paper cites Towards audio codec-based speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Towards audio codec-based speech separation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.806074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.452343Z digest=sha256:19b584daaf5073e1883078a48068a1b52a472fd893217196f08641cfe709ee6e

Observation 5070b1e8-2b4a-497e-9a18-6239815fd0cf · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Speech Separation using Neural Audio Codecs with Embedding Loss Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.785316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.457424Z digest=sha256:9d3f5b9eb64607a69ef74bef838c11bfdad5dd8da6d0f143e019229fe0259696

Observation 70b8d970-c0fd-4784-8c9c-dcfc20c0ee0c · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech,.

Speech Separation using Neural Audio Codecs with Embedding Loss An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.764170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.462773Z digest=sha256:0dd9cc262977cecf373fafb484a32d3ddb594992254d8dc9e6f45593a3cb4f91

Observation d74a01bc-ed42-4287-9de5-ccc68e404b89 · outbound

This paper cites DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

Speech Separation using Neural Audio Codecs with Embedding Loss DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.745444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.469684Z digest=sha256:5ed3731b839d19a2573e0f3181a0e393e5ca0cc94abcc45eacf54215b0b37f76

Observation 89913a5f-146c-4aaf-b01e-3bf6813c3194 · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text-to-speech,.

Speech Separation using Neural Audio Codecs with Embedding Loss LibriTTS: A corpus derived from librispeech for text-to-speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.726927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.476166Z digest=sha256:748099a37c9f9c611de26b676912c3827d0cc03ea4f2bc57a72a429e7d7d49bb

Observation bd200c3c-2545-495b-b079-7407a23aee08 · outbound

This paper cites Espnet-codec: Comprehensive training and evaluation of neural codecs for audio, music, and speech,.

Speech Separation using Neural Audio Codecs with Embedding Loss Espnet-codec: Comprehensive training and evaluation of neural codecs for audio, music, and speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.707945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.482598Z digest=sha256:4112f7c16030ed38b55df1f2a8313a29a844d48ff8a3b94ba93c7909a11ce92b

Observation 8caa13cc-51dd-47b6-97e3-a59543962839 · outbound

This paper cites Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs).

Speech Separation using Neural Audio Codecs with Embedding Loss Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.487621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.487621Z digest=sha256:2d1a044552a6a2fa7dce01379ae76c07ebd654ceb27d905b69012caf1b0df468

Observation b5da3910-e288-4e96-8884-8e5f96a2d5f6 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Deep clustering: Discriminative embeddings for segmentation and separation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.689821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.493346Z digest=sha256:bf8243e97993e26e9064cd3fefd4a299cc57007c4c43e9cedcd05829acdb65d1

Observation 6493b892-ae27-4d6d-9a59-3702ec38a7fe · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Speech Separation using Neural Audio Codecs with Embedding Loss SpeechBrain: A General-Purpose Speech Toolkit

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.499310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.499310Z digest=sha256:2d8d92796085b874604e135df0e0632162b294b843eb481e5ab70bbbf42dab91

Observation db757385-ea69-4f37-8717-c8a68321252c · outbound

This paper cites Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem.

Speech Separation using Neural Audio Codecs with Embedding Loss Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:41:45.557966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:41:45.504981Z digest=sha256:ec0543180d491ab8ecf9911d49cec148c0f5eb64833cd2dab398c040a941683d

Pith citing papers

Observation 7f0dc690-f759-4476-99b6-7b7e3466b511 · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Speech Separation using Neural Audio Codecs with Embedding Loss

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:46:37.688387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T16:44:32.104114Z digest=sha256:4d4182415183d8fd2734dfd9579f22798e2b83372a4b9e97d56e6e53942f6785

Observation 22b9aa25-409d-4d79-9cf0-d4d9f6c09442 · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Speech Separation using Neural Audio Codecs with Embedding Loss

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T16:48:34.841419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:48:34.841419Z digest=sha256:3238aadc6ca14644eb5842a0c746e1107999c63c2db2f9b9a534f693204403a1