Pith. sign in

Paper Citation Record · LEDGER

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram

As of 14 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2411.11258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11258 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:47:51.689848Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact1
  • verified fuzzy53
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c47f90a2-7987-4b51-b815-cbf24766bb4a · outbound

This paper cites APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.288451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.288451Z digest=sha256:7719502eca230b7b526c66d4e1454d67e26048b1f6af1c493911f654435bfa3b

Observation 4b56eaca-498e-458e-ab90-63622f40485e · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:53.122993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.295087Z digest=sha256:8594c0d37a07e117189e1ed311d25ff9948cfbd8d37a78d1ae1cdc1e2f42519b

Observation e5a4ab15-641f-4883-a496-1466aaef0844 · outbound

This paper cites APNet: An all-frame-level neural vocoder incorpo- rating direct prediction of amplitude and phase spectra.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram APNet: An all-frame-level neural vocoder incorpo- rating direct prediction of amplitude and phase spectra.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.104762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.300069Z digest=sha256:c290ce1e60d2ab859fd88d423fc8be281fbb8dcf00cd65e38628527af6de66e6

Observation 941397a5-4586-4567-b8d9-13954a1f2a18 · outbound

This paper cites Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.089020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.305152Z digest=sha256:b9226e53780bbd52ec59be398a56c815952a39e84c608f20256f8ee3e51b1b38

Observation 8acb7ca0-f313-455d-9cff-e87352053c20 · outbound

This paper cites Long-frame-shift neural speech phase prediction with spectral continuity enhancement and interpolation error compen- sation.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Long-frame-shift neural speech phase prediction with spectral continuity enhancement and interpolation error compen- sation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.071416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.310478Z digest=sha256:1fcaae43cb75fba7db52ad0d409b594721058cc4c82d36e5ac86f3044d45a4bf

Observation 1ed81267-acbe-47a6-b11a-49bbf046eafd · outbound

This paper cites vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.315555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.315555Z digest=sha256:e22d0d71fd41fa74e946769574b9ee9e36452de7e1e181ce0d37a20ec4023794

Observation 697d32cd-6f82-431f-8350-602e496b66b5 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Audiolm: a language modeling approach to audio generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.053852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.321616Z digest=sha256:934b9cab01fb48b4c7c866ca5c0a0e1586c0bb59d2e2c619c3ba288b7c8f923f

Observation 9bbf6214-0846-452c-8ca2-5aa583c8fc1b · outbound

This paper cites Iso/mpeg-1 audio: A generic standard for coding of high-quality digital audio.Journal of the Audio Engineering Society , 42(10):780–792, 1994.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Iso/mpeg-1 audio: A generic standard for coding of high-quality digital audio.Journal of the Audio Engineering Society , 42(10):780–792, 1994

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.032029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.326911Z digest=sha256:52c971295821941fab704b66a188933fd279df6613b5fff68ba35c1ed5d156ab

Observation f992542b-b0d6-4b1a-801a-b6db205cd865 · outbound

This paper cites Crowdsourcing preference tests, and how to detect cheating.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Crowdsourcing preference tests, and how to detect cheating

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.014682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.331806Z digest=sha256:5f31c5e5cd8a6e643c01335213e1eff48f29d0ba1a4da4ad8476ca91b9a7860f

Observation 3cb54547-c291-4b60-aed2-294195bd0fe2 · outbound

This paper cites ViSQOL v3: An open source production ready objective speech and audio metric.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram ViSQOL v3: An open source production ready objective speech and audio metric

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.997282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.337005Z digest=sha256:ee04324d262614d98273e5725294fff82c305fdf3377f3e331f4c835b9a154c1

Observation 96a17bfa-ada3-48f8-ace9-120a324aa15e · outbound

This paper cites High fidelity neural audio compression.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram High fidelity neural audio compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.980738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.342222Z digest=sha256:b9f7daada44c73315b54c2c29ea79052889b7069b580d737ec5b4fcb820d413a

Observation 2b2e2821-aa0f-491d-8f20-500d520096b7 · outbound

This paper cites Overview of the evs codec architecture.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Overview of the evs codec architecture

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.964100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.347324Z digest=sha256:19a693451ae18e0ae3fadec469d3f9355346815ca5cc189482dc47d441e46697

Observation 76fcf08f-aaab-403a-9db2-d61b6e5653d0 · outbound

This paper cites Adversarial audio synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Adversarial audio synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.352694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.352694Z digest=sha256:f266c3cac5dd1abc66aa56d1c52068f64cfc8138643439c3e8d92972c14fe6ee

Observation 3a511ba9-fbc4-4dca-870e-1e75faf706da · outbound

This paper cites VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.357279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.357279Z digest=sha256:cf58cb6fc147c023c1114663818fa6f16931987eccd82ed4c810ce3aa0db31f7

Observation ed038733-e9fb-42d2-b396-b4060d4c58c0 · outbound

This paper cites Apnet2: High-quality and high-efficiency neural vocoder with direct prediction of amplitude and phase spectra.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Apnet2: High-quality and high-efficiency neural vocoder with direct prediction of amplitude and phase spectra

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.937121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.363643Z digest=sha256:426814c37b6d00a60e9626970049c30f6a1f601c0ce59bd13d9b720bb06d04b1

Observation d6403e04-e83a-4a2c-992e-ab1a71551444 · outbound

This paper cites Generative adversarial nets.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Generative adversarial nets

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.921959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.369873Z digest=sha256:2750424b26409ce3803ef4413efdea3cc1f1248ea711e68e3262e16a224a5830

Observation c04a2b58-0840-49c9-b290-0f5e592b06a0 · outbound

This paper cites Long short-term memory.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Long short-term memory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.905221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.375651Z digest=sha256:fcace0d478a52b76e2683fd0317b588afc504742f3712f67f27084bf2fd3d2a3

Observation 4865663c-5a24-4fed-8618-85e055c53f84 · outbound

This paper cites A spectral energy distance for parallel speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A spectral energy distance for parallel speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.889445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.381364Z digest=sha256:561ca49e03e1e7bfe78508d14511825266e0437ba58ed3db34f7d38114776234

Observation c6074c2d-65be-49b0-aeb0-ebc65e1dcb41 · outbound

This paper cites A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:47:51.902113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.386579Z digest=sha256:9bf5b27e2f0cf272f2c6619e97882a2d16f9697a302ad821e6cfdd9c32795cc3

Observation d03e39e0-be2c-4a5f-9454-7870c72894c8 · outbound

This paper cites Deep residual learning for image recognition.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.393319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.393319Z digest=sha256:435f077460605a69ec45f9ef3c28a547e007ff9d71b930e59f2b6afe76adde27

Observation 55f6247c-120e-4e43-8e06-93bea41fbc63 · outbound

This paper cites Gaussian error linear units (GELUs).

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Gaussian error linear units (GELUs)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.862310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.398742Z digest=sha256:27ede4350862d0c7f03da60773780379e9355aa226c18f809c3c70b0f7118721

Observation a225fc0b-6077-4c5f-9413-b594abc6ca26 · outbound

This paper cites Hubert: Self-supervised speech repre- sentation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 29:3451–3460, 2021.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Hubert: Self-supervised speech repre- sentation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 29:3451–3460, 2021

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.845446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.403526Z digest=sha256:a10779e24bb179baae1f363b0eebbbc1731b3dda5b867a5011d3cf819904c526

Observation 220e17b8-b852-496c-9776-fb4f4a0d4001 · outbound

This paper cites RepCodec: A Speech Representation Codec for Speech Tokenization.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram RepCodec: A Speech Representation Codec for Speech Tokenization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.409362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.409362Z digest=sha256:5c9f2099754ebc81052993acc2271c6f24465a5f7332863ebbb5849707eca9f1

Observation 194898d2-e1b8-4824-9575-de9a4ccab147 · outbound

This paper cites The LJ speech dataset.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram The LJ speech dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.829906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.414859Z digest=sha256:2300697e72faea257f81fea3a21b242cf97b08bfabc762ad22a15c843067ad97

Observation 5843b434-5ecb-4fc7-91d5-ce407d34d7eb · outbound

This paper cites Univnet:Aneuralvocoder with multi-resolution spectrogram discriminators for high-fidelity waveform gener- ation.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Univnet:Aneuralvocoder with multi-resolution spectrogram discriminators for high-fidelity waveform gener- ation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.814549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.419681Z digest=sha256:d1a666e09bb28916d3deee19f6185b091fb0af286df1925530dfbeb7a4509a20

Observation 0f21d5a1-6e49-49fc-a94a-538caa4c5177 · outbound

This paper cites GlotNet—a raw waveform model for the glottal excitation in statistical parametric speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram GlotNet—a raw waveform model for the glottal excitation in statistical parametric speech synthesis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.798943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.425218Z digest=sha256:dc630d41f6ebcaf3319612cd136ffea8e16711c8676a2ff43c03451b18d4c2af

Observation b4063528-f205-44e2-b288-744ca791cc66 · outbound

This paper cites iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.783108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.430599Z digest=sha256:724d30d684fdad30f9ff829c0b258ce8d92610f668a3473b5c2d0d199640951a

Observation ccc19170-6be2-4ceb-8727-b9059f02950c · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.764666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.437083Z digest=sha256:fa970db7cc3e7595d71d9cdb8fb414fc030598443243d1d87b1e11360783b70d

Observation 499348de-367a-4e3d-b0c0-1b7306d3e7cb · outbound

This paper cites HiFi-GAN: Generative adver- sarial networks for efficient and high fidelity speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram HiFi-GAN: Generative adver- sarial networks for efficient and high fidelity speech synthesis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.748534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.442900Z digest=sha256:3f45e8ad8feffe8d23c142a3aed0ad9ef351cf340d3cc24eb1efefb8220b26a8

Observation a5ab773e-2581-48cb-a717-aa434258e7e0 · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.731393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.448372Z digest=sha256:4e9f2dfcb4406db5c496f75112379b68efa2310ac2f02221b36d41cb87a66d8c

Observation 4fdfc843-8109-4659-ab48-0fae4904b63d · outbound

This paper cites MelGAN: generative adversarial networks for conditional waveform synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram MelGAN: generative adversarial networks for conditional waveform synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.715308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.453836Z digest=sha256:f6fb32755501db92b5ee29241ab63daca507a8d21b731188a53621935282d0f7

Observation cce4f8fb-943f-431b-9f79-c3c1df243fa0 · outbound

This paper cites PHASEAUG: A differentiable augmentation for speech synthesis to simulate one-to-many mapping.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram PHASEAUG: A differentiable augmentation for speech synthesis to simulate one-to-many mapping

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.699637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.459021Z digest=sha256:9e10b591606b3a99b01caa53b1c4b01ac09add906aacb440de03a4012ed371aa

Observation 430f0ddd-fbfb-465c-921b-6aaaf1760bc2 · outbound

This paper cites Neural vocoder is all you need for speech super-resolution, 2022.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural vocoder is all you need for speech super-resolution, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.683429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.464048Z digest=sha256:97279f0085ec34c3858ed665de8e1294b6a705afbb7a5b57a0eab4294902ee51

Observation 7cc61c9e-f6aa-4b7b-bfcf-c4d1f4a034e3 · outbound

This paper cites DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.471319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.471319Z digest=sha256:001f93f8730faad2758a3064c7c6d97a88432789deb69361abfef764443dd89e

Observation bf4bcccb-b280-4870-ac60-957f1ff73308 · outbound

This paper cites A convnet for the 2020s.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A convnet for the 2020s

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.666050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.477762Z digest=sha256:339d890c016c3cef20df31b41370e2d419017319124a8db77c79e94edbf9eb57

Observation 2653fa72-aaf2-41f2-9b47-7d3f33b439a8 · outbound

This paper cites Decoupled weight decay regularization.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Decoupled weight decay regularization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.649809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.483846Z digest=sha256:77965bb1580dbf2f8e3b3836c34bb3796930d06dc72cf29dc87417c1b7d6f609

Observation ec25e984-31ea-416e-a5ea-549839a9ae5b · outbound

This paper cites Source-filter-based generative adversarial neural vocoder for high fidelity speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Source-filter-based generative adversarial neural vocoder for high fidelity speech synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.632295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.489915Z digest=sha256:00dfe8a389b922ab49608b60534d9d95f93c5ed159bbb4cd84c6519da392cbdc

Observation 8b4c81a1-b1bc-4688-99b2-95a34b50ec34 · outbound

This paper cites MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.494938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.494938Z digest=sha256:d3497b18fc777b1e246177f16b07fa6f11a35cfd25e0afc197bc0efe86f3a485

Observation c2e6fa0c-9703-4bbc-93e8-b54035ab2da8 · outbound

This paper cites Rectifier nonlinearities improve neural network acoustic models.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Rectifier nonlinearities improve neural network acoustic models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.614586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.500806Z digest=sha256:839e90d199e7129d7978c5bf6c1a50520b0fa01fffe05bb68a6f4ebd4eee49ff

Observation 28320407-d618-4955-9d2c-260111eaf5c5 · outbound

This paper cites SampleRNN: An unconditional end-to-end neural audio generation model.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram SampleRNN: An unconditional end-to-end neural audio generation model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.598385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.505578Z digest=sha256:3a02936f37bb538583ebd238c2824543ba56a881ff4f2b43aa7dfae127c123c2

Observation a8c7ddc5-5a83-458b-81dd-b176c47734f2 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Finite scalar quantization: Vq-vae made simple

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.580994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.510632Z digest=sha256:cf9df8ac320ecaba56b87a73987831e42971ef9bb61e6504d07831654bbc0957

Observation c4024053-0c4c-4cc2-bc87-1602d3a5a2aa · outbound

This paper cites WORLD: A vocoder-based high-quality speech synthesis system for real-time applications.IEICE Transac- tions on Information and Systems , 99(7):1877–1884, 2016.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram WORLD: A vocoder-based high-quality speech synthesis system for real-time applications.IEICE Transac- tions on Information and Systems , 99(7):1877–1884, 2016

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.562869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.515933Z digest=sha256:f13e9c2623b4edfcbe2080c38663968c4570fe30b6ad5f417d09664cca0ee2ef

Observation c421bbc6-cd07-49eb-b690-2329909c9f50 · outbound

This paper cites Expediting tts synthesis with adversarial vocoding.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Expediting tts synthesis with adversarial vocoding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.542031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.521474Z digest=sha256:d8b92fe0a715bc4129c84f1317b46abc22092d99ca570190414863adbf0d923c

Observation a2a8e21e-e2b5-44d8-babd-9337ca96c34f · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.527155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.527155Z digest=sha256:ae5438b4ba3230f24bb02aa5121ad3f89eb7ed2821a7ca8ed7010d628cd1f239

Observation 446f2159-383d-447f-9972-00152859cf80 · outbound

This paper cites Parallel WaveNet: Fast high-fidelity speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Parallel WaveNet: Fast high-fidelity speech synthesis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.525246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.534346Z digest=sha256:eacdea50b0da26c9a1a8d48dc32f83edf3fafbddc2fcf89c30d05ca3b2bec4f8

Observation 846bffcc-0a1a-4ed1-8977-3dbd2776866c · outbound

This paper cites WaveNet: A generative model for raw audio.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram WaveNet: A generative model for raw audio

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.507049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.540109Z digest=sha256:2db348230c7169dc420069136924f2645385cccbbafc43f30fa232f9b64838fb

Observation f1ca97b6-c3cf-4cdf-9b5b-4aaeba487477 · outbound

This paper cites Linear predictive coding.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Linear predictive coding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.489617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.544982Z digest=sha256:87714a6565c16eb115ef4423a8fd1ce14f6e6712dc990a4df865a0b8c071342f

Observation ee2e052b-9e6d-4b35-b7e3-f2e102498c0f · outbound

This paper cites Generativeadversarialnetwork-basedapproachtosignal reconstruction from magnitude spectrogram.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Generativeadversarialnetwork-basedapproachtosignal reconstruction from magnitude spectrogram

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.472077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.549935Z digest=sha256:a2ede4c008812f3cfe7cd9f9c6b6143355d1bbdd3a4b8b3c98ea753c7b64e42f

Observation 88d59813-fc92-40cf-8001-f5b6c91ff54a · outbound

This paper cites Wave-gan: a deep learning approach for the prediction of nonlinear regular wave loads and run-up on a fixed cylinder.Coastal Engineering, 167:103902, 2021.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Wave-gan: a deep learning approach for the prediction of nonlinear regular wave loads and run-up on a fixed cylinder.Coastal Engineering, 167:103902, 2021

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.452183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.554746Z digest=sha256:0f2d3793157f70e0ec7351acdf29762715bd498dbdc992595bf39e73fe849f1f

Observation 1cefcd94-0c6d-4474-bd30-1d4e6dffccf9 · outbound

This paper cites ClariNet: Parallel wave generation in end-to-end text-to-speech.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram ClariNet: Parallel wave generation in end-to-end text-to-speech

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.433624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.559532Z digest=sha256:6f6cf933524d3e267a37d59f87c194be4d02dcdf4acb178ec8fd7605a0c61af8

Observation 8e4a7480-60ac-4177-892d-27b0e8a90bbc · outbound

This paper cites Waveflow: A compact flow- based model for raw audio.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Waveflow: A compact flow- based model for raw audio

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.406461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.564487Z digest=sha256:738dd499062d67712f100a5f9f29160a54b7a3aa5dc62c093de707dcf0e24388

Observation a57b22a4-2418-4bee-8c32-e4bb6897a897 · outbound

This paper cites Waveglow: A flow-based gen- erative network for speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Waveglow: A flow-based gen- erative network for speech synthesis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.389077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.570862Z digest=sha256:27f5471ec26e79b51c29e47b18d4bb4c39704391fe9d628723cdbe2d525809a2

Observation 0ab47423-2800-45be-9f89-3af3eb931652 · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.372286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.576285Z digest=sha256:ac5d8f532ce61a69cf2c3656dee3c1c443bb9974086c64fd044dfe63b948c814

Observation 4a9d6b56-0a72-4a98-acf0-3178a7d499ff · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Fastspeech 2: Fast and high-quality end-to-end text to speech

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.356217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.582351Z digest=sha256:66d8bcc7576862bfa9485b879b637a43fbe49cdd57abbbdd677e43f531b9c98a

Observation cce7f17d-e7df-428f-9ddd-745085f93a16 · outbound

This paper cites Fewer-token neural speech codec with time-invariant codes.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Fewer-token neural speech codec with time-invariant codes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.338323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.587468Z digest=sha256:653b82d82af720ac6b9de02c3060441d97ddd01832d2965d970e395849b231e4

Observation f35e040f-1fb3-4a21-bda6-ba0a552839fb · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Utmos: Utokyo-sarulab system for voicemos challenge 2022

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.592094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.592094Z digest=sha256:b9fe8185bab870dd2c62a82a16592ee470b211e68ee9e6acf502b9c977777878

Observation 6d485cd0-18c6-4f42-be48-006957a13a47 · outbound

This paper cites A toll quality 8 kb/s speech codec for the personal communications system (pcs).IEEE Transactions on Vehicular Technology, 43(3):808–816, 1994.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A toll quality 8 kb/s speech codec for the personal communications system (pcs).IEEE Transactions on Vehicular Technology, 43(3):808–816, 1994

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.311192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.597654Z digest=sha256:5217a10e73c7a89308baa0dd3a5326e30440b8948550e399c4ce53edd92f9714

Observation ad724549-5b05-493f-a20e-c0f221caf14f · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.294500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.603829Z digest=sha256:e773022c6395f08dbdedcc90dbe9ea89165ec1642f4f39c487731480840d5db4

Observation 223fd8ff-ccae-411d-ad3f-1723c05de0e4 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.609553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.609553Z digest=sha256:842bb55c662a03b234e357e5ee9b6ce1a04147db5a27a427f26949b60901c6af

Observation d25f7265-a625-4703-a6c3-af82bb254f83 · outbound

This paper cites Linear predictive coding systems.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Linear predictive coding systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.277542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.615331Z digest=sha256:ccb7df446b4749939ec79e3ef646d7373364f7af303c7f3e069e34ced325b984

Observation 536dd418-5b3b-4fec-8080-3cb6b22ec255 · outbound

This paper cites High- quality, low-delay music coding in the opus codec.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram High- quality, low-delay music coding in the opus codec

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.261629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.621588Z digest=sha256:e2ed3843e719e3b711ec1a36d8c5f44f6746d0f1934b06253deccc93546b4865

Observation 455395df-d923-45ba-a8fc-d65ea5869f46 · outbound

This paper cites LPCNet: Improving neural speech synthesis through linear prediction.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram LPCNet: Improving neural speech synthesis through linear prediction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.242851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.627258Z digest=sha256:c62427a2e8757d868eefea62e4770f8ed3072d6417b96b6d262f699dcefc0c21

Observation 5f834bf8-7453-4c3d-8c24-85198c114358 · outbound

This paper cites A review of vector quantization techniques.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A review of vector quantization techniques

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.226450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.632471Z digest=sha256:cdb68dde9e1a7446d0aa37affe55553465533ed3ab7332222d122c2083378d15

Observation 165e37aa-c693-4d38-a038-f46bb45f28ad · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.637690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.637690Z digest=sha256:092dec4165cae759928c6e2349c6adbb6ca96033826cabbd4c1246186db9fa1d

Observation 6da65fae-c7cc-433d-92dd-8627fff1668a · outbound

This paper cites Neural source-filter-based wave- form model for statistical parametric speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural source-filter-based wave- form model for statistical parametric speech synthesis

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.209143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.644418Z digest=sha256:0ef538e130f59fe73d301607d18cf9f478b576ca2059a4e6b09834cc5fa2d759

Observation 88c238f5-05fd-4d7b-b651-67814202c71e · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis.Interspeech 2017, 2017.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Tacotron: Towards end-to-end speech synthesis.Interspeech 2017, 2017

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.191451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.650301Z digest=sha256:b27f8d851396e7c3b9be8b65f416c32f450c88d61a0c8686b2e681c35e440265

Observation 1bdeb114-39a5-4218-8e64-b04765d03778 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.172141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.655870Z digest=sha256:15fdb159a039db23d3e5b6a304da7168a08889639de3a6235b6d16269fc530ef

Observation dac23bcf-ef23-4e67-8590-0c402b8028b5 · outbound

This paper cites Audiodec: An open-source streaming high-fidelity neural audio codec.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Audiodec: An open-source streaming high-fidelity neural audio codec

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.154250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.662140Z digest=sha256:f42c1bdfe8311b7e8d46cd49bbc1b56f7f0c1230d2c6fd5951ea9aa19fff4561

Observation 48aca32e-1094-499e-ad7a-dccea276c529 · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92).

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.136766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.668031Z digest=sha256:75a72ed60e219ece4371d62f1cddf6382fe16bfae1023468c5dc69b1fcf8fee0

Observation eccdc361-0a0e-40e0-aae4-1271ce9acdc2 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.673711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.673711Z digest=sha256:7b938813fe1ff6b5b5f868fa7d508a8a2d80dfeef2174d618a2b90483c51f6ec

Observation db58b2bc-4a36-4a1d-a59c-9283c5753b11 · outbound

This paper cites Source-filter hifi-gan: Fast and pitch controllable high-fidelity neural vocoder.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Source-filter hifi-gan: Fast and pitch controllable high-fidelity neural vocoder

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:51.989129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.678726Z digest=sha256:93616dbea7c99efe5fb4226109aca8a195877d0e27f82833364008d385152e31

Observation e86c4863-6802-4197-ad1c-3619e1441c71 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , 30:495–507, 2021.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Soundstream: An end-to-end neural audio codec.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , 30:495–507, 2021

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:51.970628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:47:51.683963Z digest=sha256:8ebeeb85be8907f705cb1fecd144dcf60ccefb626743cd39bf331ec91c45a9f7

Observation 0a7a1dab-4a31-4b46-9b73-f62714c0ae11 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.689848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.689848Z digest=sha256:7ce049dca5e2bd4a0e382af68a752ea198cec626a040c06000c3eec2ed18f94a

Pith citing papers

No inbound Pith citation observations are available.