Pith. sign in

Paper Citation Record · LEDGER

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2501.05332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05332 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:15.195399Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78316f17-55a2-4429-872a-b6473ab948d5 · outbound

This paper cites an unresolved cited work.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:18:16.056272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.921168Z digest=sha256:77ef321614809e72da1ef70d91f3740fabe51f69ce63821c487c09ebbf087b88

Observation cf939511-44de-4644-af16-085bce42c0ac · outbound

This paper cites Speech analysis/synthesis based on a sinusoidal representation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech analysis/synthesis based on a sinusoidal representation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.039414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.927861Z digest=sha256:68f776bd507678140ee6c5a3efc5d9c09d2711a19a67f4b96231ab0f143bff88

Observation c97c97b4-7d2d-4cca-8193-6c69c27c2290 · outbound

This paper cites Speech analysis/synthesis and modifica- tion using an analysis-by-synthesis/overlap-add sinusoidal model,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech analysis/synthesis and modifica- tion using an analysis-by-synthesis/overlap-add sinusoidal model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.021198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.933873Z digest=sha256:bba7491474dba8d6f469710c8ed24ff738f92d353d5af828f6f654d825288fb0

Observation 8a77f054-5119-4a95-a723-b81134781679 · outbound

This paper cites Spectral modeling synthesis: A sound analy- sis/synthesis system based on a deterministic plus stochastic decomposi- tion,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Spectral modeling synthesis: A sound analy- sis/synthesis system based on a deterministic plus stochastic decomposi- tion,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.001357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.939389Z digest=sha256:539e67f948cdd36b6de1a3c847d811316feec7b2ebdb06e87f0e353cf7cf5c58

Observation ab84d7e6-ad06-4345-87f5-2bb589433605 · outbound

This paper cites HNS: Speech modification based on a harmonic + noise model,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder HNS: Speech modification based on a harmonic + noise model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.977931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.945301Z digest=sha256:3725a9a3b3686cff65b9aa973fab6daae1c347f9057eb80ebac591cfe4beb325

Observation 188a3b83-9118-433d-a814-fc430dc3649b · outbound

This paper cites Analysis/synthesis and modification of the speech aperiodic component,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Analysis/synthesis and modification of the speech aperiodic component,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.957651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.950661Z digest=sha256:d8c8d9ccd3ee2acbb0572e3d1ddb7e52436403a9897c2116b746a65e59d4fb62

Observation 9455ab30-511c-45e4-9d69-20246d1ccfd7 · outbound

This paper cites Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.937209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.956552Z digest=sha256:3339a38507ad3d7396a9e57b1693e8e2b07af1e0c252b75729de727ee9ae08db

Observation e7f06493-4983-4df9-bd53-89ec86b944a3 · outbound

This paper cites Improved phase vocoder time-scale modi- fication of audio,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Improved phase vocoder time-scale modi- fication of audio,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.911352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.961508Z digest=sha256:157d51e65abb362d3ceb486efbda0e3f04ea282a15a4632ca1fd7f052b3a6a10

Observation e74256c8-99bb-49da-89e2-0240c9d9bdda · outbound

This paper cites STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.893689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.966599Z digest=sha256:2b0bd0cd781a8e06ab3bb9c81a172896e25450e477d7a366cf8dfcee77dc9446

Observation 09759bbd-2886-4718-be68-de420a8b5dc9 · outbound

This paper cites World: a vocoder-based high- quality speech synthesis system for real-time applications,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder World: a vocoder-based high- quality speech synthesis system for real-time applications,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.874649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.971913Z digest=sha256:5d0c3c08569852706e1ce29c84e3f27850a3d3d64255ec1d6368918c6c192b96

Observation ef9f6a6b-fccc-4a3c-a941-b8d69dc941d3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Soundstream: An end-to-end neural audio codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.854015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.976248Z digest=sha256:aa39d7349dc0278567771429ed6ca9de3b39726b7df96508fb4ed5d1c563a884

Observation cbeca490-e130-432e-89c3-f35dc3aa61a0 · outbound

This paper cites High fidelity neural audio compression,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder High fidelity neural audio compression,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:14.980882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:14.980882Z digest=sha256:7c8491ff8bc7a708cb04534f36e8292ceb4f9236e025070ac1c79dda5b082aec

Observation caa1de04-6e11-47e2-a2ca-af8b7e9b465e · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.822520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.985768Z digest=sha256:01fa75b9e41f010a20e5f468444ef5704b64aa3becb7b4bdc36513bfbba9fa84

Observation c9521e5c-be4e-42b5-b8a2-e72bb11f31ff · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:14.990854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:14.990854Z digest=sha256:daa7efe1036461e650147decaef5b8f18650d2a0b4729a063e22129e16db0c45

Observation 87ea4ccf-987b-4777-9fd3-b96796e9e6a6 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Masked autoencoders are scalable vision learners,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.803457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.997002Z digest=sha256:909295e5a6e812ed1b64bb0bfc30b0c24d928666d87b8a677bd99d6a1b2c8685

Observation cfa6c315-f0c5-4f3c-ad6b-037b4117194d · outbound

This paper cites Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.786693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.011918Z digest=sha256:842cf017abf0ea9487b1b0696675187cf38f528fbf12a3dccdaf5e09c9d9028a

Observation 764b5b8e-019d-448b-b4a8-547843a84ddd · outbound

This paper cites CREPE: A convolutional representation for pitch estimation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder CREPE: A convolutional representation for pitch estimation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.763028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.020133Z digest=sha256:af6d413cad5ef8badac0cdd5466ca9ae98abc915f8a610e5fa4667bd89c5877e

Observation f2252bfe-383e-4dcd-8ce8-a514fd94867f · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.025829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.025829Z digest=sha256:3b9ac2a15032cd3f5cb957aebf002aeb679b949f5951b87caa3d651bad6763cc

Observation be57210f-98ea-44cc-8da8-2b47308fb7bc · outbound

This paper cites Brouhaha: Multi- task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Brouhaha: Multi- task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.747351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.032234Z digest=sha256:98031d877f7cfc07f6243e9e00f2fc2a9382a1f660c2176ab20ead6448c79b72

Observation cd3d3ccd-d69f-42f2-86e1-a075dd9b3454 · outbound

This paper cites Neural discrete representation learning,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Neural discrete representation learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.731382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.039976Z digest=sha256:e1883215afe132dbe16eb016432ad5ba860cb3d7357fbee658cb2bde60f862cc

Observation 0131dad0-d0cc-41a3-9ce5-e2bf3e9f6390 · outbound

This paper cites A vector quantized masked autoencoder for audiovisual speech emotion recognition.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A vector quantized masked autoencoder for audiovisual speech emotion recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.048903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.048903Z digest=sha256:ae1ac61536ebeb3201e3d8c0e4d0679130aa0bccc7d43d860ea6da99dc2a641d

Observation 5b12b587-a5e6-4c31-a7e9-00eead8f41b0 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A comparison of discrete and soft speech units for improved voice conversion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.715030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.055931Z digest=sha256:f4d6d2d0008a6211cba43d48f4da5ecf1ddcdc41f7bacc84a6c8821e5bdaa4bb

Observation 9c9d8261-2a90-4ef7-9df1-48e127b2fe7b · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.697722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.061930Z digest=sha256:40f1096c47cc992bd7b667f119bd83d69c19e67da39dfe3fbc59adfa87f1806f

Observation 3c0d0dc3-acb0-4387-a660-9bafcd021556 · outbound

This paper cites Attention is all you need,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Attention is all you need,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.682108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.069225Z digest=sha256:7527b107644f61a38d644cf698377ab9060f2dc73d2e1d6532873655d8dbd819

Observation ead23dc7-9872-488e-bf7f-f032c23489c4 · outbound

This paper cites DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.078041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.078041Z digest=sha256:7a4f344789852a028775d9f27ea813fb453bb853fc5dfbb616e1776dbf5954bd

Observation 512054d5-6ed5-4586-83dc-012f29817b1a · outbound

This paper cites Phase- aware speech enhancement with deep complex U-net,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Phase- aware speech enhancement with deep complex U-net,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.666451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.084578Z digest=sha256:04f97b37f1fc2a3f160053f287377ea2318403f67d7ce3d3dcc9ef52990e4708

Observation 0de9e15e-57ae-4738-aec3-aba81e331c3c · outbound

This paper cites Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.089856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.089856Z digest=sha256:2f7af7144416cd75a1b3743fe94a3ac2cf0f697d4b9a8e8a038e1c605062338b

Observation c94154bf-a919-4a22-848b-827755ed68b3 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.649768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.095540Z digest=sha256:ebf07d844aaf7012ad8fafeb07d5c383d23218e1820f10ecdf13afca3e0e2b27

Observation a1c9c67c-0775-47b4-b998-8cc8c4419e08 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Librispeech: an asr corpus based on public domain audio books,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.632143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.101155Z digest=sha256:a8cc8c0c7a1ef617735c1d7a27f331c6bcd5d72e923d404c65a7ba329f8cb0eb

Observation 1b54d229-4a4f-46f9-8fd5-7a8766f7cacc · outbound

This paper cites DEMAND: a collection of multi- channel recordings of acoustic noise in diverse environments,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DEMAND: a collection of multi- channel recordings of acoustic noise in diverse environments,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.616365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.106247Z digest=sha256:e3ec68a003b718292e5dc557f9b141a838685bb1cdb8d58626e2503a2e5884db

Observation fecfb541-5dc9-4496-b178-08330e6a079a · outbound

This paper cites Pedalboard,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Pedalboard,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.110816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.110816Z digest=sha256:8f87f6ac2e886f30d713e1d95e00587e41d19183ddebe6baeef40599ccd5a9d9

Observation 2f270ae9-7702-4761-bae5-017a66649232 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.115417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.115417Z digest=sha256:87d3d4133c87605571ccfd27a70ad0ab35b0ab029654bdd42b418e93f8ece83b

Observation a8e88527-72e1-4c09-890a-5cc3a0765c84 · outbound

This paper cites Wham!: Extending speech separation to noisy environments,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Wham!: Extending speech separation to noisy environments,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.599067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.120486Z digest=sha256:60b4f5b69d850a18852224df0bf5dace938e514a037b22a7a04747dd4799f49f

Observation d47231e4-35b9-4781-a762-26a42955d8a6 · outbound

This paper cites Decoupled Weight Decay Regularization.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Decoupled Weight Decay Regularization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.129971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.129971Z digest=sha256:892aad23adadfc6e55859a33e7c6f067b98abb59e554d9a6ce77ea1f5efde9db

Observation 2f8f3491-dbb2-4707-a545-67bf629a53cd · outbound

This paper cites SDR–half-baked or well done?.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder SDR–half-baked or well done?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.581834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.136261Z digest=sha256:b72288f8209261f6ad293da51d7889cde3a254622a655de558ede792f9e29280

Observation 3b48262a-ffeb-4c02-bfa6-c10ab9d26b63 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ) – a new method for speech quality assessment of telephone networks and codecs,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Perceptual evaluation of speech quality (PESQ) – a new method for speech quality assessment of telephone networks and codecs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.566373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.142181Z digest=sha256:cfd208dae70cea87974b20f9bc8ec6911d6736352df42b613c083a141aa786a4

Observation 8b20d736-d8d6-40eb-b500-868462a92682 · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.550740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.147179Z digest=sha256:a3b1ea0aa9c0edd77a9519592f2e9a5c7321c6cafb30ee7515e755bad7eee5b5

Observation d7e7dd40-04d7-4d69-a8f0-bca3b4530130 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Generalized end-to-end loss for speaker verification,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.534553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.152200Z digest=sha256:043c745efc75ae56138ec6a1bcd470d352176b19add8da23bb8b82326b61530e

Observation 45192b41-6934-44e2-8e63-2e92347c00bb · outbound

This paper cites Speech quality assessment through MOS using non-matching references,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech quality assessment through MOS using non-matching references,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.517122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.157220Z digest=sha256:789357338e965f0572bc2773c2eca1ec2b03062b5c33df38501d5beb3adecb4b

Observation 1f76d720-2f26-40d7-9d5d-f9e853a80288 · outbound

This paper cites Torchaudio-Squim: Reference-less speech quality and intelligibility measures in Torchaudio,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Torchaudio-Squim: Reference-less speech quality and intelligibility measures in Torchaudio,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.500396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.162432Z digest=sha256:182013296b98d8d28195f2b72be557d3adf50aa27c773a7c7285973ccaa67cfb

Observation ce955239-e364-4fda-991a-f0bf358e0c7a · outbound

This paper cites DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.482873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.167258Z digest=sha256:a82cddcf3a41a193abceef21be37c01fbe799e3303f56b859330df04cd752c9c

Observation 7ed81e39-7248-4c77-a725-851eead3c682 · outbound

This paper cites A pitch tracking corpus with evaluation on multipitch tracking scenario,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A pitch tracking corpus with evaluation on multipitch tracking scenario,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.465591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.172809Z digest=sha256:132df7a5cee2198383d1660a36c2bfa448daa2c55a09a17c56bf67d641b20381

Observation 5b793503-8cbb-4b0c-8407-9945181c9317 · outbound

This paper cites pYIN: A fundamental frequency estimator using probabilistic threshold distributions,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder pYIN: A fundamental frequency estimator using probabilistic threshold distributions,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.445626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.178371Z digest=sha256:04bd36bd383c23ee21ae90143bca0d4d002e28b32aae8242f403bbd555ad426b

Observation e31f8484-8a53-46c8-b473-90725fca49ec · outbound

This paper cites A sawtooth waveform inspired pitch estimator for speech and music,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A sawtooth waveform inspired pitch estimator for speech and music,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.184183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.184183Z digest=sha256:ed874a9c0240673610642b00ccfb01ed5bd92b57b44a4e960dea99ef918aacfd

Observation 3016e0d9-78bb-43c2-b5bf-7450912cd148 · outbound

This paper cites Revise: Self- supervised speech resynthesis with visual input for universal and generalized speech regeneration,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Revise: Self- supervised speech resynthesis with visual input for universal and generalized speech regeneration,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.412512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.190587Z digest=sha256:e8ce9c67283649a2cbc32f26cfd429e38fbe38da8ea5f75b4525da6b09b554dd

Observation 6cda7d60-be83-49b5-84b9-70df58dceec9 · outbound

This paper cites Generating diverse high- fidelity images with VQ-V AE-2,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Generating diverse high- fidelity images with VQ-V AE-2,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.395111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.195399Z digest=sha256:8c7fdd43d920f1a93d1367cab80bbef53092162fa4cd7605d36b5d36318b4f53

Pith citing papers

No inbound Pith citation observations are available.