Pith. sign in

Paper Citation Record · LEDGER

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.14767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14767 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:45.148144Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T22:36:21.635792Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.524737Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba1e0b1e-8c72-470e-883e-3c2d34136503 · outbound

This paper cites The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.678703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.678703Z digest=sha256:ba3359b9eae816ffc0fae87b0855416c4959b4d06f7aaea574f006934860baa4

Observation f2a6e95a-cd6c-44f3-852b-6107383c8b37 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Audiolm: A language modeling approach to audio generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.686691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.686691Z digest=sha256:6bbb404d3ead7672b7f93b7266ce70706d03b4fedc5687d91f59b4b076d5c792

Observation 7081adfd-e962-4bae-8224-f7250c0f310b · outbound

This paper cites R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.694626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.694626Z digest=sha256:3dae70560c21a2f0d3b957499ee3a82248105d1d7a31780e23245502d3a0bb1d

Observation 9f61c2dd-2a54-4237-a64a-74f2d07545c2 · outbound

This paper cites A vector quantized approach for text to speech synthesis on real-world spontaneous speech.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models A vector quantized approach for text to speech synthesis on real-world spontaneous speech

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T00:15:45.456626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.706752Z digest=sha256:2070391d6f5d05310ae1f37635f260c15b9fe63168359b8d4900e82ce02d79b4

Observation b77a3c44-563f-4113-9052-de17d230f94b · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.711037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.711037Z digest=sha256:7ae2c4ec64d2fbfab75abc565a9711bcb33d7a106b629894b47941ea03a7ff44

Observation 564ecf6b-fd82-4b41-b60e-b9bf15fb71bc · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.715062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.715062Z digest=sha256:26ebbd29e90bdcbc07a83574d5472ff2f0d12f0185f4969699f04348968b8d85

Observation 89de81df-8587-459f-866e-90107bc36d87 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Neural codec language models are zero-shot text to speech synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.719272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.719272Z digest=sha256:487d767e5829712b31269b98ace9ada91042c31538fac103242c4c80bccd4e47

Observation 1ebd119f-6a03-4b11-ad4e-e0f3d1dc3ba3 · outbound

This paper cites High fidelity neural audio compression.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models High fidelity neural audio compression

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.246973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.722483Z digest=sha256:82ce0b1c3f081d2db81014c8da1b5008c41d97378b728c4d91afbda3c97d8db1

Observation cde48c6a-a147-4d50-b57b-8ced722801bb · outbound

This paper cites Density estimation using real NVP.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Density estimation using real NVP

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.230593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.727952Z digest=sha256:ecdb15a83adf66bae847a0dfeee0554feac39ba79b38d20a0f36ba36535496ef

Observation 76296d15-6b1d-4f77-a815-ea0fe44a5d52 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.732151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.732151Z digest=sha256:6436844faf2c05b86e493362c18e6016fc6a31791c524da633351362c649df6b

Observation 0d48f864-0325-4e74-8888-1f56b95b740d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.737743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.737743Z digest=sha256:2a6c2cffe057fb068f148188ccbb27476bb6d8f0a82539af99b1346482da5152

Observation ca1acb72-cb5d-46c0-96ac-d265c4e1dacb · outbound

This paper cites R., Schuller, B.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models R., Schuller, B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.218086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.743777Z digest=sha256:adaab7ebfcc4b266cfb7785bcd263cb99c47287c975280ea72571e3bb2f8dc24

Observation d193aad3-7f07-40cc-8916-e05d2be7e331 · outbound

This paper cites Cyclical annealing schedule: A simple approach to mitigating KL vanishing.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Cyclical annealing schedule: A simple approach to mitigating KL vanishing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.747794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.747794Z digest=sha256:002684b47c0172e158cb6ba675f6ad73b0d78733f7beca1295d7b768d4e357a1

Observation 0d6769a3-93da-4b33-bb26-ed499e4c2394 · outbound

This paper cites A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Defossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Defossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.203129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.751608Z digest=sha256:6af18599a6a7145cc4623f328008a61dd6633bf6f454cc35f16e223b01b71e02

Observation b9c97a94-90df-4789-96a5-e8211f254aa4 · outbound

This paper cites and Gimpel, K.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Gimpel, K

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.191551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.755063Z digest=sha256:1427ac02a241887f36cecb5ada606a4a6432ea1a5b35069da1b67a35f908c2c0

Observation 698c1489-ce5d-4a3d-ae2a-37e220a2dbb3 · outbound

This paper cites beta- VAE : Learning basic visual concepts with a constrained variational framework.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models beta- VAE : Learning basic visual concepts with a constrained variational framework

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.178622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.759218Z digest=sha256:98195beb90345cea93575be894b7be54ed23d4a085ef42b66cec666ded4f7657

Observation 93b6eb26-fadb-4ee1-86ed-11e7a50b624f · outbound

This paper cites Denoising diffusion probabilistic models.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Denoising diffusion probabilistic models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.163238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.763093Z digest=sha256:a0b86de35c64acac87556890df3352d3e339333f0e8e06def1c5de2a848ab107

Observation 5399a6ff-35fc-46a3-8a72-4035d2294f04 · outbound

This paper cites H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.767005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.767005Z digest=sha256:a0e4cfbc360871c51f2b74b2bbaf28b5d4b51730cda7791d19e08810cfd9f8b3

Observation b4ff4c7b-9797-48ac-9d62-c757a7a281e6 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Libri-light: A benchmark for ASR with limited or no supervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.148703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.771294Z digest=sha256:a62fd470d64d5660ce94468f29030efa3889fb07cd18f2709e220a1276b77a40

Observation 3e0fef10-d4a0-4037-a8d8-1791feefc7c7 · outbound

This paper cites A., Riviere, M., Mohamed, A., Dupoux, E., and Hsu, W.-N.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models A., Riviere, M., Mohamed, A., Dupoux, E., and Hsu, W.-N

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.778686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.778686Z digest=sha256:dc4328926ed3f05a49447483593fe5fdfc98962ae9c6bd233ab37dfad611a212

Observation b25007a2-782c-44e2-aa25-61a3ed16ad22 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Glow-tts: A generative flow for text-to-speech via monotonic alignment search

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.126895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.786060Z digest=sha256:5f0deed7fd262ee8a8323240d2bdb3d1446cb98a8e7a5147ef78f7ef6306339a

Observation c3534bbe-3196-4cbc-a46a-2700ccb22eaf · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.102107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.807307Z digest=sha256:224cf866403417f4059b21081a6dfbec686cf119aa8b2f57624d8741d0ddf3c6

Observation ecd4c332-ee4b-450a-9065-e6a0a012edb8 · outbound

This paper cites W., Salamon, J., Li, P., and Bello, J.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models W., Salamon, J., Li, P., and Bello, J

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.811638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.811638Z digest=sha256:890d78fae48b4cd3cee9251129a5c8f2757f2445c89f0aaf7a93e672666bef60

Observation 759585a2-6a73-4d2f-8e28-20e96b1913aa · outbound

This paper cites an unresolved cited work.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.818856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.818856Z digest=sha256:5e60a11df396369a66fbfc9f99eacc0f599b0a86f4ab8c008c5667381f1a2529

Observation 9939882c-0785-4730-93e0-b7aec7842f83 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.088472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.832434Z digest=sha256:84da1cd53db26bdd40120b439b471e037ffa27977d62b428f250eb5e5dbe2bb3

Observation 0c5aef27-0c33-4419-b661-b77d5c82c445 · outbound

This paper cites On generative spoken language modeling from raw audio.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models On generative spoken language modeling from raw audio

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.839065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.839065Z digest=sha256:275e7242f288d1913830ba2f4faed75a101feddd2b6baf1d35d6c5f0cd55b1a5

Observation 788a30d6-ce08-4c10-98e1-02cbc78266dc · outbound

This paper cites and Hutter, F.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Hutter, F

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.847821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.847821Z digest=sha256:35850069cc844b9a38f4ee516692b348a9d24fa39a76b050dca50e50a0cf7ff4

Observation 09122a09-c6f1-4d73-bc4e-ac17bd61a597 · outbound

This paper cites Voxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Voxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.857666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.857666Z digest=sha256:f4433d1d35c8a3dfbed4be3443183af56c6e7931d5d6c57765492ea864ff64e4

Observation b3dfe765-ef91-4123-97e4-f790772ba959 · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.865712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.865712Z digest=sha256:03483263dff6e846dc3b00124507e7e96bbd1d70d1eb7ac77e16c86371c2d231

Observation b5b906e3-3692-4993-a309-9a45dc5f3820 · outbound

This paper cites GPT-4 Technical Report.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.894816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.894816Z digest=sha256:b066316ac90e62c6fba4df603fe5c735d4f19315323ff970960b9aaa2c81b416

Observation a40358ec-9d19-4c10-8439-2ee44a0c04d2 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Librispeech: An ASR corpus based on public domain audio books

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.913607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.913607Z digest=sha256:45533bbe05f00cb33fd6cb26c1da86e9fa810f7b9e1afbe724892d125359807e

Observation f4f05302-5d31-4840-bb16-4c470fb3169e · outbound

This paper cites FiLM : Visual reasoning with a general conditioning layer.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models FiLM : Visual reasoning with a general conditioning layer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.917330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.917330Z digest=sha256:1e07aef3b7f92e4c1f59dd8aab64ac308088319634a5501b3104207130926bb4

Observation 6161288b-832b-4021-8b9f-338a826f5ee1 · outbound

This paper cites Train short, test long: Attention with linear biases enables input length extrapolation.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Train short, test long: Attention with linear biases enables input length extrapolation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.921466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.921466Z digest=sha256:36eb9d808246ed9002d50e37fc2a205b5815fffa9229788c3138765d6e4590ba

Observation 68b61674-7217-4229-8eb5-a2286a4006fa · outbound

This paper cites W., Xu, T., Brockman, G., Mcleavey, C., and Sutskever, I.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models W., Xu, T., Brockman, G., Mcleavey, C., and Sutskever, I

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.925313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.925313Z digest=sha256:f06b2431a830d81b3944b52c4df5871e2a1ab6b483e6c2f0191c2d5e7f387df3

Observation 8d0a9955-4f6c-4e3b-b365-c57975b17953 · outbound

This paper cites and Torre, R.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Torre, R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.026292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.931413Z digest=sha256:e2d93bbcb6b602a175412e629b4237179cd2cd7f7c3b9d93ab53b880bc7a8c93

Observation 1f065bbb-92a8-42b2-af87-0012358b00fb · outbound

This paper cites and Mohamed, S.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Mohamed, S

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.004749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.936155Z digest=sha256:d0460811e746563a8560e90a18b8e064f1be2844e7f47d6f40e472ae6b06b8a6

Observation b6086208-4561-4e62-a19c-57ce05649417 · outbound

This paper cites U-Net : Convolutional networks for biomedical image segmentation.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models U-Net : Convolutional networks for biomedical image segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.981677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.940000Z digest=sha256:4b81a354952d11ffffab76a44d484cf412416493fc6062f17f9b188b9e68a1e3

Observation 250938e1-419f-42c5-8657-6eeebc8e36a0 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.944189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.944189Z digest=sha256:74ff19250581db5b8b13c51ae6330198938614dc0cb75e0b5a7a7f5daa33f118

Observation ca725f33-4967-4d55-8d71-7663d8635d77 · outbound

This paper cites The interspeech 2009 emotion challenge.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The interspeech 2009 emotion challenge

Reference 39

Resolution
verified exact
doi, observed 2026-08-07T00:15:45.321783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.948242Z digest=sha256:732dd4cfa63835c45a47b666de6357425eaa84b451e9795ea8b3f5bc1b4785f7

Observation cf588aef-9ebe-4026-b8c7-f318b7b238e1 · outbound

This paper cites The interspeech 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The interspeech 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.952054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.952054Z digest=sha256:466fdc0e5eb03fb36224dd009dd5ff3098416eb47af75628bf99ae67777e17bf

Observation 50183299-26c3-43e0-956c-9f0784219d60 · outbound

This paper cites Denoising diffusion implicit models.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Denoising diffusion implicit models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.948606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.957310Z digest=sha256:240e28f367281c759ba7916272c1766d84d61231ac696e4b639d9f510a15f093

Observation 5d9a725c-12f5-428e-a78a-29c0ef522c8a · outbound

This paper cites J., Cao, Y., Zen, H., Rosenberg, A., Ramabhadran, B., and Wu, Y.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models J., Cao, Y., Zen, H., Rosenberg, A., Ramabhadran, B., and Wu, Y

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.963041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.963041Z digest=sha256:44341df3884823413719507fc3bae919813a64b64c51e205f883891972dd8fd8

Observation 50697df9-0e5b-420f-993f-35ef51bc8b7e · outbound

This paper cites Instance Normalization: The Missing Ingredient for Fast Stylization.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Instance Normalization: The Missing Ingredient for Fast Stylization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.968933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.968933Z digest=sha256:8f3fc9586346146ad7e930e25b0b851b9934dd6157609f6d5f8da98ffc10dbe8

Observation a02db15b-48f2-4f0c-9a22-08874b0ee3a5 · outbound

This paper cites and Kautz, J.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Kautz, J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.931145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.973522Z digest=sha256:262188c216e158ef33ef4261faa18c6791b005a5dc6b220224ad06ebf719988d

Observation b23bd89d-c775-4353-b9d6-bfd3ddf53469 · outbound

This paper cites Learning de-identified representations of prosody from raw audio.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Learning de-identified representations of prosody from raw audio

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.915433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.980070Z digest=sha256:01a3d372c394aa10d762e03f828e34b661f396b07c72d1d8dfc11e1c4c03bf71

Observation 0a438fe2-7ef9-48ee-9747-3270430b9103 · outbound

This paper cites On layer normalization in the transformer architecture.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models On layer normalization in the transformer architecture

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.897756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.990676Z digest=sha256:4522c0a8d868899f07f3b8d878dc5c25ca22ebc6e08860a92a8165cd513812b4

Observation af5c8ce8-8abd-4d61-a155-cb6747b226f9 · outbound

This paper cites CSTR VCTK Corpus : English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models CSTR VCTK Corpus : English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.997848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.997848Z digest=sha256:e302fe7781c08b5f18498397258b741c14e4dd0cd1a1bc3aa896f206d6fc999a

Observation 23b41b76-da0d-4cb5-84e0-be94071ce7f3 · outbound

This paper cites an unresolved cited work.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:15:46.880905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.002990Z digest=sha256:04fc109a2812842bb5288169fc617a34b1c3ad983a509204a87cb2af635e514f

Observation 83211cda-db71-4d3b-9aac-3d338593031b · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Soundstream: An end-to-end neural audio codec

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:45.017442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.017442Z digest=sha256:4e67ed827c019f8e8bbb0d8a01cbcceffde5fd60c8c4371e6b0360de298c5987

Observation ff1b8f43-70e1-47ae-9c2b-1067a04d88a1 · outbound

This paper cites and Sennrich, R.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Sennrich, R

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.864209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.044841Z digest=sha256:82ddd68e0b719777300f6473ccc5d442531fb208098df4e9b45210b692ec19f1

Observation c1c6ae91-4d1c-44e4-bfdb-9ac0a89cc508 · outbound

This paper cites SpeechTokenizer : Unified speech tokenizer for speech language models.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models SpeechTokenizer : Unified speech tokenizer for speech language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.849087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.085676Z digest=sha256:8ce1d0e53d5b4cdd5285f986a114c07c66f340a871e81bb50cbc1e3597dc8f9a

Observation ccce29af-1127-41b4-a35f-07f657f13abb · outbound

This paper cites R., Kadav, A., and Graf, H.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models R., Kadav, A., and Graf, H

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.828978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.092264Z digest=sha256:6bd89408c846ddcf672939f49fef63ba63a9a3c1605b6f3b462ffb22ddfe409e

Observation 82f83d34-0588-4528-8c2b-b1d7d84b1214 · outbound

This paper cites @esa (Ref.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models @esa (Ref

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:45.106503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.106503Z digest=sha256:d6ab3fbe2a048933d785279505c389f5dfbc78fff6ca4333dc85362c22257595

Observation 17f572f4-8e35-4aa1-9047-dbed4a13db6d · outbound

This paper cites an unresolved cited work.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:45.135549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.135549Z digest=sha256:f29a16b57ad7a8d8419adbf504e914318f32143a88f14a7f2487728d4e5709cc

Observation 0bee6347-6538-413a-968a-799224ee8119 · outbound

This paper cites The evaluation metrics are detailed in Section ssec:eval-all.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The evaluation metrics are detailed in Section ssec:eval-all

Reference 55

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:15:45.148144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.148144Z digest=sha256:db5afb7de598978221997c255daf7af7c2ec562e5d133ee2987467f731eb5d7b

Pith citing papers

Observation 8206ffa6-12e9-41c5-88f8-80bacbdf1fd3 · inbound

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias cites this paper.

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias A Variational Framework for Improving Naturalness in Generative Spoken Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.526798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T22:36:21.635792Z digest=sha256:e7ca08611c8799b96996cb0058887d931162e05d2b68a34bc1e37426691552c5