Pith. sign in

Paper Citation Record · LEDGER

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

As of 12 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2501.05332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05332 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:15.195399Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78316f17-55a2-4429-872a-b6473ab948d5 · outbound

This paper cites an unresolved cited work.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:18:16.056272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.921168Z digest=sha256:498581ee70aece9f5b6e67c744e3f18f8390de417411049c7a050ad57d2f8598

Observation cf939511-44de-4644-af16-085bce42c0ac · outbound

This paper cites Speech analysis/synthesis based on a sinusoidal representation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech analysis/synthesis based on a sinusoidal representation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.039414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.927861Z digest=sha256:8f6bdf1dd16f1978439cc2781e073e5d8c313b0b1ef2c6f655d6c57c3fed2a4b

Observation c97c97b4-7d2d-4cca-8193-6c69c27c2290 · outbound

This paper cites Speech analysis/synthesis and modifica- tion using an analysis-by-synthesis/overlap-add sinusoidal model,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech analysis/synthesis and modifica- tion using an analysis-by-synthesis/overlap-add sinusoidal model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.021198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.933873Z digest=sha256:85fce059c677a577f3084ef27d8d0b961ee1c11ace74b3b58f00a3dd6d9e86cb

Observation 8a77f054-5119-4a95-a723-b81134781679 · outbound

This paper cites Spectral modeling synthesis: A sound analy- sis/synthesis system based on a deterministic plus stochastic decomposi- tion,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Spectral modeling synthesis: A sound analy- sis/synthesis system based on a deterministic plus stochastic decomposi- tion,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.001357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.939389Z digest=sha256:cd99b0b43e5e76c98f28f7b74e78ba97ef6b2425d4d9821808752878e373dbee

Observation ab84d7e6-ad06-4345-87f5-2bb589433605 · outbound

This paper cites HNS: Speech modification based on a harmonic + noise model,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder HNS: Speech modification based on a harmonic + noise model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.977931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.945301Z digest=sha256:f31ec87b27949d7d331e2a3404826d20033448782a5fbc10a19d6154309d9e97

Observation 188a3b83-9118-433d-a814-fc430dc3649b · outbound

This paper cites Analysis/synthesis and modification of the speech aperiodic component,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Analysis/synthesis and modification of the speech aperiodic component,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.957651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.950661Z digest=sha256:85a7a0944b399307f2cb17f5327c10117dfdc8e6891baad7ada4a4adce3d7e89

Observation 9455ab30-511c-45e4-9d69-20246d1ccfd7 · outbound

This paper cites Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.937209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.956552Z digest=sha256:c652c470853d16ff38fcff5deff616e87809f545c5c7995e6a3af1d0dbc85805

Observation e7f06493-4983-4df9-bd53-89ec86b944a3 · outbound

This paper cites Improved phase vocoder time-scale modi- fication of audio,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Improved phase vocoder time-scale modi- fication of audio,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.911352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.961508Z digest=sha256:a557d1fbf839c07206873953532275d4a6b12f988848d85c214be175070af7ee

Observation e74256c8-99bb-49da-89e2-0240c9d9bdda · outbound

This paper cites STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.893689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.966599Z digest=sha256:7f792904f1161a1c4e80ceab18135903e96217a554845203a8c42892fd0cf564

Observation 09759bbd-2886-4718-be68-de420a8b5dc9 · outbound

This paper cites World: a vocoder-based high- quality speech synthesis system for real-time applications,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder World: a vocoder-based high- quality speech synthesis system for real-time applications,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.874649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.971913Z digest=sha256:64670028fdef2d5bf19bd662176639294c07f938b600365fa65d720d0fdc59ce

Observation ef9f6a6b-fccc-4a3c-a941-b8d69dc941d3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Soundstream: An end-to-end neural audio codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.854015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.976248Z digest=sha256:c1fbcbe779d3345ed2cdf95e0f70be213c3d1ca84814e95b10e992cc26a1f431

Observation cbeca490-e130-432e-89c3-f35dc3aa61a0 · outbound

This paper cites High fidelity neural audio compression,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder High fidelity neural audio compression,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:14.980882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:14.980882Z digest=sha256:617c1a26929f2b9870bc6215312ec37d2751d39a0fdf359b279af440f05ffb44

Observation caa1de04-6e11-47e2-a2ca-af8b7e9b465e · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.822520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.985768Z digest=sha256:14abda9c112777ac33f954dab0e049c6ba46a7eff226cd734da2c6d3b7d27d02

Observation c9521e5c-be4e-42b5-b8a2-e72bb11f31ff · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:14.990854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:14.990854Z digest=sha256:c00bad24487fd1a288549e8cdd126cc2e89adeb26f9e6d120838e5c9ba086545

Observation 87ea4ccf-987b-4777-9fd3-b96796e9e6a6 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Masked autoencoders are scalable vision learners,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.803457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:14.997002Z digest=sha256:1bd3b4b8283b321755bee0e2e3f70dd4c5fe245e915680025f85239ccf3ddf0a

Observation cfa6c315-f0c5-4f3c-ad6b-037b4117194d · outbound

This paper cites Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.786693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.011918Z digest=sha256:587eb70b0e92d7307aa10f1d80662d9b32e0306384aed1491341aeba81ecbfa8

Observation 764b5b8e-019d-448b-b4a8-547843a84ddd · outbound

This paper cites CREPE: A convolutional representation for pitch estimation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder CREPE: A convolutional representation for pitch estimation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.763028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.020133Z digest=sha256:14157d6c210e97c0defd147d42a378377bc827d109941233258b3645bc6b50c0

Observation f2252bfe-383e-4dcd-8ce8-a514fd94867f · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.025829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.025829Z digest=sha256:cece7466ebb14dc3f6a7139424ebf5551674308d6cc7168ceba58cc54c903024

Observation be57210f-98ea-44cc-8da8-2b47308fb7bc · outbound

This paper cites Brouhaha: Multi- task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Brouhaha: Multi- task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.747351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.032234Z digest=sha256:176cc70b7c6236812d1438ad653f9aa13470f39d9df373f2b722643a5fc1ae40

Observation cd3d3ccd-d69f-42f2-86e1-a075dd9b3454 · outbound

This paper cites Neural discrete representation learning,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Neural discrete representation learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.731382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.039976Z digest=sha256:bc85b8ec4159ca1a41083a2fcf0f2a6e9fccc0db362a10dedc2e9aa00ba9f81d

Observation 0131dad0-d0cc-41a3-9ce5-e2bf3e9f6390 · outbound

This paper cites A vector quantized masked autoencoder for audiovisual speech emotion recognition.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A vector quantized masked autoencoder for audiovisual speech emotion recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.048903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.048903Z digest=sha256:b6953ca869715798e80d3550392c5bf73e25762d1c58d9875b23ad346d97d439

Observation 5b12b587-a5e6-4c31-a7e9-00eead8f41b0 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A comparison of discrete and soft speech units for improved voice conversion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.715030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.055931Z digest=sha256:6ddef81baabb08c4f5b06278c2e5180842a3e3aa245b3c00fda11ae10a68c58e

Observation 9c9d8261-2a90-4ef7-9df1-48e127b2fe7b · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.697722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.061930Z digest=sha256:3ddd73ffac32cf13e2e318d7f7c687d31ed6521d61db07816d60f04cac432f8c

Observation 3c0d0dc3-acb0-4387-a660-9bafcd021556 · outbound

This paper cites Attention is all you need,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Attention is all you need,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.682108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.069225Z digest=sha256:cc1554a49e3a9f42849e26f04d422ceb03af9badbcec3b1b586bba9de1b45582

Observation ead23dc7-9872-488e-bf7f-f032c23489c4 · outbound

This paper cites DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.078041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.078041Z digest=sha256:603f03b3cff475f6d9d7c3d759441b80cb246d9e7eab3287c0ecac80b24c9d85

Observation 512054d5-6ed5-4586-83dc-012f29817b1a · outbound

This paper cites Phase- aware speech enhancement with deep complex U-net,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Phase- aware speech enhancement with deep complex U-net,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.666451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.084578Z digest=sha256:25def4ffbe6290b74102ffd4e17e51e0c4c64a2178a98bfc66ecf9d3d2cda9e5

Observation 0de9e15e-57ae-4738-aec3-aba81e331c3c · outbound

This paper cites Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.089856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.089856Z digest=sha256:9401f4787f98ba701376e9680ebfc132625f8334102835f47946eed41426df73

Observation c94154bf-a919-4a22-848b-827755ed68b3 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.649768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.095540Z digest=sha256:c3cc20d3cb8edba7d5eea758bf413abb8e44e299310e99a70434ab6b70f72a28

Observation a1c9c67c-0775-47b4-b998-8cc8c4419e08 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Librispeech: an asr corpus based on public domain audio books,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.632143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.101155Z digest=sha256:c46a0780eeaf82dd3a2783907a90b1d17ab57372b3b9eec8d950036736f14f48

Observation 1b54d229-4a4f-46f9-8fd5-7a8766f7cacc · outbound

This paper cites DEMAND: a collection of multi- channel recordings of acoustic noise in diverse environments,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DEMAND: a collection of multi- channel recordings of acoustic noise in diverse environments,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.616365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.106247Z digest=sha256:c7f3cb761f74bedddf12a93c54d8a79409a1956144c5b9e70adfcefd0817f90e

Observation fecfb541-5dc9-4496-b178-08330e6a079a · outbound

This paper cites Pedalboard,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Pedalboard,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.110816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.110816Z digest=sha256:b58e61ebf58374ecebba82df7f1b514c8842843d07e52ef6f77759737a3394e7

Observation 2f270ae9-7702-4761-bae5-017a66649232 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.115417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.115417Z digest=sha256:15a564308533aa93303f9481fc49d9a602be07fa1c56916fa51c9945a2a4bc98

Observation a8e88527-72e1-4c09-890a-5cc3a0765c84 · outbound

This paper cites Wham!: Extending speech separation to noisy environments,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Wham!: Extending speech separation to noisy environments,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.599067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.120486Z digest=sha256:c117d32ee02681a9f89b62a95c5a02315421b650bf98e2086c6e7ea850ec0c8f

Observation d47231e4-35b9-4781-a762-26a42955d8a6 · outbound

This paper cites Decoupled Weight Decay Regularization.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Decoupled Weight Decay Regularization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.129971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.129971Z digest=sha256:273830da5d5e8fcf528c4dc7c4100e549cf319b11aa2aba546dd1e615e5a6b15

Observation 2f8f3491-dbb2-4707-a545-67bf629a53cd · outbound

This paper cites SDR–half-baked or well done?.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder SDR–half-baked or well done?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.581834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.136261Z digest=sha256:4104f394907aff4d0c2a88f970ad9d8c566eb492ec1f0461d5b83b2fa886cf3b

Observation 3b48262a-ffeb-4c02-bfa6-c10ab9d26b63 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ) – a new method for speech quality assessment of telephone networks and codecs,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Perceptual evaluation of speech quality (PESQ) – a new method for speech quality assessment of telephone networks and codecs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.566373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.142181Z digest=sha256:ac06131fb2b77b937b66326967fb0f4236922c93423479ec0d24042b7f5462da

Observation 8b20d736-d8d6-40eb-b500-868462a92682 · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.550740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.147179Z digest=sha256:94408eb5b317fb5e64b69617af31ea666bf89eff5997a3eceb5f2f3ffebf1c67

Observation d7e7dd40-04d7-4d69-a8f0-bca3b4530130 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Generalized end-to-end loss for speaker verification,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.534553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.152200Z digest=sha256:37a4aade27a76399f28653756a49f9f639055666515ffed99a32f545a1dbf2c6

Observation 45192b41-6934-44e2-8e63-2e92347c00bb · outbound

This paper cites Speech quality assessment through MOS using non-matching references,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech quality assessment through MOS using non-matching references,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.517122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.157220Z digest=sha256:aabd2824860ddf525f8f6caf075e8392256217e8221a7faee05195095bfe342e

Observation 1f76d720-2f26-40d7-9d5d-f9e853a80288 · outbound

This paper cites Torchaudio-Squim: Reference-less speech quality and intelligibility measures in Torchaudio,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Torchaudio-Squim: Reference-less speech quality and intelligibility measures in Torchaudio,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.500396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.162432Z digest=sha256:07e424c63a633c789365a821f46e73dbfa2bcf5097a1661d521763adb295f637

Observation ce955239-e364-4fda-991a-f0bf358e0c7a · outbound

This paper cites DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.482873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.167258Z digest=sha256:a987df58c40002bdd9eebdb7bc58bacb65ab134df7a232f579bdb681e8fc5e6c

Observation 7ed81e39-7248-4c77-a725-851eead3c682 · outbound

This paper cites A pitch tracking corpus with evaluation on multipitch tracking scenario,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A pitch tracking corpus with evaluation on multipitch tracking scenario,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.465591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.172809Z digest=sha256:2bcfc65a812a42b61b5040f63645ad9a0b4fa6d73ef3d73f1708b2cfa5859d55

Observation 5b793503-8cbb-4b0c-8407-9945181c9317 · outbound

This paper cites pYIN: A fundamental frequency estimator using probabilistic threshold distributions,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder pYIN: A fundamental frequency estimator using probabilistic threshold distributions,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.445626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.178371Z digest=sha256:b9eae46aceab0f8e1d11e9bac23f7b25cc47c97507904c2003839517b295ff6f

Observation e31f8484-8a53-46c8-b473-90725fca49ec · outbound

This paper cites A sawtooth waveform inspired pitch estimator for speech and music,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A sawtooth waveform inspired pitch estimator for speech and music,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.184183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.184183Z digest=sha256:882d22f39ce5b72c5a722b53c57f1f948fd9382ab507d161bee41129df66487f

Observation 3016e0d9-78bb-43c2-b5bf-7450912cd148 · outbound

This paper cites Revise: Self- supervised speech resynthesis with visual input for universal and generalized speech regeneration,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Revise: Self- supervised speech resynthesis with visual input for universal and generalized speech regeneration,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.412512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.190587Z digest=sha256:17c04862131630fdd79539217678771d9c8007d19894dab6588609269f437b9b

Observation 6cda7d60-be83-49b5-84b9-70df58dceec9 · outbound

This paper cites Generating diverse high- fidelity images with VQ-V AE-2,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Generating diverse high- fidelity images with VQ-V AE-2,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.395111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:18:15.195399Z digest=sha256:dca38dc0065cdf826f5e416202bdb0781e79cc40c80ef69d8bcd010785f38a6a

Pith citing papers

No inbound Pith citation observations are available.