Pith. sign in

Paper Citation Record · LEDGER

Speech Separation using Neural Audio Codecs with Embedding Loss

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2411.17998.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17998 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:41:45.504981Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:48:34.841419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T16:46:37.685528Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69f94909-8948-4cc1-9db6-b80182fe6a11 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

Speech Separation using Neural Audio Codecs with Embedding Loss Some experiments on the recognition of speech, with one and with two ears,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.333593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.333593Z digest=sha256:6164db0a41b6898b61aecda23a7056e94bc7cbeab1812116aa21074abbd4b458

Observation e92cd3e9-9beb-480c-bd97-e27a7b691490 · outbound

This paper cites Atten- tion is all you need in speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Atten- tion is all you need in speech separation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.067206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.340401Z digest=sha256:5483ae4335c106bc7dc529331de656450b09b93609779b947261a9937d8fe1b0

Observation ca44b4b5-c246-46a6-90e4-01c890318825 · outbound

This paper cites Spgm: Prioritizing local features for enhanced speech separation performance,.

Speech Separation using Neural Audio Codecs with Embedding Loss Spgm: Prioritizing local features for enhanced speech separation performance,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.049025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.345659Z digest=sha256:d3fb34c6b5035cebe2ab5b603f710244b8f4c19d2610c5c7e1facdd23045f12e

Observation d76f519d-8cbc-4d7c-ba3b-07822213962d · outbound

This paper cites Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.031180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.351303Z digest=sha256:3f371c418988f0923b545f78d828d0027cdf7ac00e73d8978ac4f13a3822cdee

Observation 599e8f6b-a545-44de-845a-4bcaeb6c6ede · outbound

This paper cites Gass: Generalizing audio source separation with large-scale data,.

Speech Separation using Neural Audio Codecs with Embedding Loss Gass: Generalizing audio source separation with large-scale data,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:46.011198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.356597Z digest=sha256:565c40803d3e86236405510aad66093c650b8f44202513587c57f4ab2ba63897

Observation 2be18dab-e38d-47ce-bec7-d7f590469310 · outbound

This paper cites Exploring self-attention mechanisms for speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Exploring self-attention mechanisms for speech separation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.993979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.363038Z digest=sha256:125ccb3181108555d7b3ee09fd94b4a8858c77a9a8e6d88a0c0e4ae84638111e

Observation b0a861b9-44e3-463f-b060-bfd07c4d0cb3 · outbound

This paper cites TF- GridNet: Integrating full- and sub-band modeling for speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss TF- GridNet: Integrating full- and sub-band modeling for speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.974904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.369277Z digest=sha256:ce21b8a085a2a1d82d5a76b9e3ee787581c0d7ae4bcb9529d677c7260167da21

Observation fccdc939-376c-4d7b-94e1-9c641505d6a9 · outbound

This paper cites A neural state-space model approach to efficient speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss A neural state-space model approach to efficient speech separation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.956269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.374305Z digest=sha256:415ceab09ff37cdd7b9a950ac02153c615dbac3d8fc5b3fd8d50451c604e4c62

Observation df33a327-e700-4751-b367-6dd6c643002a · outbound

This paper cites Permutation invariant training of deep models for speaker-independent multi-talker speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Permutation invariant training of deep models for speaker-independent multi-talker speech separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.939179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.381233Z digest=sha256:331b62d307c6287a25734215432afc845747736b14a5c986219af76fe7396286

Observation e8243670-d1ec-49ac-9c85-5f5b5b4e546f · outbound

This paper cites Sdr–half-baked or well done?.

Speech Separation using Neural Audio Codecs with Embedding Loss Sdr–half-baked or well done?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.386290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.386290Z digest=sha256:c717a7c8eaa441fb14cb7a8673a54e1a02da6c668e284c8704e6889b2f4ce5b1

Observation d59cce01-1d76-4aa2-8e8a-43250db14d6b · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Speech Separation using Neural Audio Codecs with Embedding Loss Soundstream: An end-to-end neural audio codec,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.391032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.391032Z digest=sha256:8c12220a37b20293f9812820643cbb658106134ed38124e0150bc36f479b2c32

Observation 117c8613-a45b-4ed7-b081-799e295d7930 · outbound

This paper cites High Fidelity Neural Audio Compression.

Speech Separation using Neural Audio Codecs with Embedding Loss High Fidelity Neural Audio Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.395769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.395769Z digest=sha256:816985ad6dc103372142ab9050b425570205395d98efff89a5c8fec64d58aadc

Observation 06f22c58-3e30-4455-ba24-096d212cf527 · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

Speech Separation using Neural Audio Codecs with Embedding Loss High- fidelity audio compression with improved rvqgan,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.401593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.401593Z digest=sha256:bc7c42bd2a4a20a334ea01afbcdfb45a25b6ceccb18edb60710efc0c47522893

Observation 3b162e68-e8e0-4646-a278-0daec62c0e03 · outbound

This paper cites Vector quantization,.

Speech Separation using Neural Audio Codecs with Embedding Loss Vector quantization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.885444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.407138Z digest=sha256:cc4c5ff88b606c2cb0c71343284193557be3fe89ac3bea6d7940ed1f0a19b0ab

Observation 1a76d6cc-4751-466a-8922-2d176bc48e85 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

Speech Separation using Neural Audio Codecs with Embedding Loss VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.414147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.414147Z digest=sha256:cfa9dd14af6f12453b03c2f8ea5dc98a5236b391a0be3e3861a950f78b0dfc64

Observation 10a6b267-3f43-42bb-ac11-4d53c1e4dbd9 · outbound

This paper cites Exploring the limits of decoder-only models trained on public speech recognition corpora.

Speech Separation using Neural Audio Codecs with Embedding Loss Exploring the limits of decoder-only models trained on public speech recognition corpora

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.420271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.420271Z digest=sha256:e05f633937a385d650e9017c03ae7578b0875b120c16d092ecf07f1c72bd23f0

Observation 826a5a72-243c-4839-8519-bb264089a574 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Speech Separation using Neural Audio Codecs with Embedding Loss Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.865570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.426939Z digest=sha256:f7120af0a56caaac42b2c7a802fcd792552e4729c3a3e43ec4a54a70d154a933

Observation e0a27129-931a-4f22-92d5-d9946d014e29 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

Speech Separation using Neural Audio Codecs with Embedding Loss SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.432669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.432669Z digest=sha256:3683be16440b086cc92fa1063f1cf0950c377e3d82af3405a7747b7ad3fa1d60

Observation 95d00d77-5065-4bec-bf4f-dfb640e74074 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

Speech Separation using Neural Audio Codecs with Embedding Loss Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.843407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.439015Z digest=sha256:68d9dbfb2e8a5cf54c829a317524054a75d932d9dd8f235484c915d4435d62e2

Observation d14e5e27-71bf-4831-982e-9937e0dfdf04 · outbound

This paper cites The Interspeech 2024 challenge on speech processing using discrete units,.

Speech Separation using Neural Audio Codecs with Embedding Loss The Interspeech 2024 challenge on speech processing using discrete units,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.825518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.446408Z digest=sha256:4aaf279594ac73b9c307378089fc97819b001c315530b25a4a7c3c5665494971

Observation adcd272e-f277-4ebe-b3df-d0d60ae0a7c5 · outbound

This paper cites Towards audio codec-based speech separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Towards audio codec-based speech separation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.806074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.452343Z digest=sha256:93f4d04ffd79a4fbf732b9c4b7a2fcfba6aa88ec52a0590108c67d632dd2381c

Observation 5070b1e8-2b4a-497e-9a18-6239815fd0cf · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Speech Separation using Neural Audio Codecs with Embedding Loss Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.785316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.457424Z digest=sha256:b6c13a4d17bcc5bc6a417a3faf2190830aab3d94aa7a1f4668331519e28b5d28

Observation 70b8d970-c0fd-4784-8c9c-dcfc20c0ee0c · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech,.

Speech Separation using Neural Audio Codecs with Embedding Loss An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.764170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.462773Z digest=sha256:217e294f561c84e2926baa52e619b403d06a06f2bc021cf32ec50d69c8e2895b

Observation d74a01bc-ed42-4287-9de5-ccc68e404b89 · outbound

This paper cites DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

Speech Separation using Neural Audio Codecs with Embedding Loss DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.745444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.469684Z digest=sha256:f934c6c374712abb91952bbef9c976678de93acb5c04d025bd9445b3d500450c

Observation 89913a5f-146c-4aaf-b01e-3bf6813c3194 · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text-to-speech,.

Speech Separation using Neural Audio Codecs with Embedding Loss LibriTTS: A corpus derived from librispeech for text-to-speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.726927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.476166Z digest=sha256:6ed0fbf24a3fc5339ce91d7dc5093c628a648077f7bcce9efdf17019f30577e4

Observation bd200c3c-2545-495b-b079-7407a23aee08 · outbound

This paper cites Espnet-codec: Comprehensive training and evaluation of neural codecs for audio, music, and speech,.

Speech Separation using Neural Audio Codecs with Embedding Loss Espnet-codec: Comprehensive training and evaluation of neural codecs for audio, music, and speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.707945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.482598Z digest=sha256:b357043eb52c659fbc3308e90c04832ea2e93d8fa9219bbff8e7b38d9838e4ec

Observation 8caa13cc-51dd-47b6-97e3-a59543962839 · outbound

This paper cites Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs).

Speech Separation using Neural Audio Codecs with Embedding Loss Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.487621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.487621Z digest=sha256:2d1a044552a6a2fa7dce01379ae76c07ebd654ceb27d905b69012caf1b0df468

Observation b5da3910-e288-4e96-8884-8e5f96a2d5f6 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation,.

Speech Separation using Neural Audio Codecs with Embedding Loss Deep clustering: Discriminative embeddings for segmentation and separation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:41:45.689821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.493346Z digest=sha256:09bf62163e1e5c09c4911b006f2aed9d5c7cead019a7327d7c690433e9cdc8af

Observation 6493b892-ae27-4d6d-9a59-3702ec38a7fe · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Speech Separation using Neural Audio Codecs with Embedding Loss SpeechBrain: A General-Purpose Speech Toolkit

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.499310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.499310Z digest=sha256:2d8d92796085b874604e135df0e0632162b294b843eb481e5ab70bbbf42dab91

Observation db757385-ea69-4f37-8717-c8a68321252c · outbound

This paper cites Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem.

Speech Separation using Neural Audio Codecs with Embedding Loss Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:41:45.557966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:41:45.504981Z digest=sha256:46cd8d535fa1f17d6832d2c6c763a13e4f20dc7f3e102ce97f9d7d5e1f3ba4fd

Pith citing papers

Observation 7f0dc690-f759-4476-99b6-7b7e3466b511 · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Speech Separation using Neural Audio Codecs with Embedding Loss

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:46:37.688387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T16:44:32.104114Z digest=sha256:9d2462bb37e5820b0a6eb0501845f0112c20eab492f60c7903215fe198ece43b

Observation 22b9aa25-409d-4d79-9cf0-d4d9f6c09442 · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Speech Separation using Neural Audio Codecs with Embedding Loss

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T16:48:34.841419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:48:34.841419Z digest=sha256:3238aadc6ca14644eb5842a0c746e1107999c63c2db2f9b9a534f693204403a1