Pith. sign in

Paper Citation Record · LEDGER

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.14767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14767 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:45.148144Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T22:36:21.635792Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.524737Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba1e0b1e-8c72-470e-883e-3c2d34136503 · outbound

This paper cites The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.678703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.678703Z digest=sha256:cd71b9cf462181219013dd50401a4bd9f042ba208790c6f45afb0ae45f8f2ee5

Observation f2a6e95a-cd6c-44f3-852b-6107383c8b37 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Audiolm: A language modeling approach to audio generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.686691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.686691Z digest=sha256:6bbb404d3ead7672b7f93b7266ce70706d03b4fedc5687d91f59b4b076d5c792

Observation 7081adfd-e962-4bae-8224-f7250c0f310b · outbound

This paper cites R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.694626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.694626Z digest=sha256:3dae70560c21a2f0d3b957499ee3a82248105d1d7a31780e23245502d3a0bb1d

Observation 9f61c2dd-2a54-4237-a64a-74f2d07545c2 · outbound

This paper cites A vector quantized approach for text to speech synthesis on real-world spontaneous speech.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models A vector quantized approach for text to speech synthesis on real-world spontaneous speech

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T00:15:45.456626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.706752Z digest=sha256:4afad11fec55887b824bdbc31a52b0b2b6677028acf1180b55c9308b215f5846

Observation b77a3c44-563f-4113-9052-de17d230f94b · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.711037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.711037Z digest=sha256:7ae2c4ec64d2fbfab75abc565a9711bcb33d7a106b629894b47941ea03a7ff44

Observation 564ecf6b-fd82-4b41-b60e-b9bf15fb71bc · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.715062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.715062Z digest=sha256:26ebbd29e90bdcbc07a83574d5472ff2f0d12f0185f4969699f04348968b8d85

Observation 89de81df-8587-459f-866e-90107bc36d87 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Neural codec language models are zero-shot text to speech synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.719272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.719272Z digest=sha256:487d767e5829712b31269b98ace9ada91042c31538fac103242c4c80bccd4e47

Observation 1ebd119f-6a03-4b11-ad4e-e0f3d1dc3ba3 · outbound

This paper cites High fidelity neural audio compression.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models High fidelity neural audio compression

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.246973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.722483Z digest=sha256:662f8de9ee575ddcaa9978376471a6beae95ae032bc0ca5d419ea8d9033b7b3b

Observation cde48c6a-a147-4d50-b57b-8ced722801bb · outbound

This paper cites Density estimation using real NVP.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Density estimation using real NVP

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.230593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.727952Z digest=sha256:33eba55a0ded1d3c48821b7732e2f29694bf96fdb6460eec7cfc982e3bcf4829

Observation 76296d15-6b1d-4f77-a815-ea0fe44a5d52 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.732151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.732151Z digest=sha256:6436844faf2c05b86e493362c18e6016fc6a31791c524da633351362c649df6b

Observation 0d48f864-0325-4e74-8888-1f56b95b740d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.737743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.737743Z digest=sha256:2a6c2cffe057fb068f148188ccbb27476bb6d8f0a82539af99b1346482da5152

Observation ca1acb72-cb5d-46c0-96ac-d265c4e1dacb · outbound

This paper cites R., Schuller, B.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models R., Schuller, B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.218086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.743777Z digest=sha256:a88450d14c3a858d4e7ed588611025c15379ddc7abbe6acbf86607c098da0ff6

Observation d193aad3-7f07-40cc-8916-e05d2be7e331 · outbound

This paper cites Cyclical annealing schedule: A simple approach to mitigating KL vanishing.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Cyclical annealing schedule: A simple approach to mitigating KL vanishing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.747794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.747794Z digest=sha256:002684b47c0172e158cb6ba675f6ad73b0d78733f7beca1295d7b768d4e357a1

Observation 0d6769a3-93da-4b33-bb26-ed499e4c2394 · outbound

This paper cites A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Defossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Defossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.203129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.751608Z digest=sha256:8d3f7a0d4dcdd26941b6f29a8e1880cc28d5a30a73652c4f748dfd02964a16fe

Observation b9c97a94-90df-4789-96a5-e8211f254aa4 · outbound

This paper cites and Gimpel, K.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Gimpel, K

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.191551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.755063Z digest=sha256:240c9a4ea4dfc7e323f97b63a9730784b539b82560b3e2a5784d357b3d9c8a66

Observation 698c1489-ce5d-4a3d-ae2a-37e220a2dbb3 · outbound

This paper cites beta- VAE : Learning basic visual concepts with a constrained variational framework.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models beta- VAE : Learning basic visual concepts with a constrained variational framework

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.178622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.759218Z digest=sha256:db6318d50a07137234bcd1bd8497109564948c6ea94674d8834b4dcd5803e921

Observation 93b6eb26-fadb-4ee1-86ed-11e7a50b624f · outbound

This paper cites Denoising diffusion probabilistic models.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Denoising diffusion probabilistic models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.163238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.763093Z digest=sha256:d75fb8847b88fb8e43f0dc2c6bc62c5cd2f4f28662d686ef5ea8df399b5ba678

Observation 5399a6ff-35fc-46a3-8a72-4035d2294f04 · outbound

This paper cites H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.767005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.767005Z digest=sha256:a0e4cfbc360871c51f2b74b2bbaf28b5d4b51730cda7791d19e08810cfd9f8b3

Observation b4ff4c7b-9797-48ac-9d62-c757a7a281e6 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Libri-light: A benchmark for ASR with limited or no supervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.148703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.771294Z digest=sha256:e97c0803a8f9737c66824b9f618aa8bb1bbf08192c8670b916f8b6b4931f3001

Observation 3e0fef10-d4a0-4037-a8d8-1791feefc7c7 · outbound

This paper cites A., Riviere, M., Mohamed, A., Dupoux, E., and Hsu, W.-N.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models A., Riviere, M., Mohamed, A., Dupoux, E., and Hsu, W.-N

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.778686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.778686Z digest=sha256:dc4328926ed3f05a49447483593fe5fdfc98962ae9c6bd233ab37dfad611a212

Observation b25007a2-782c-44e2-aa25-61a3ed16ad22 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Glow-tts: A generative flow for text-to-speech via monotonic alignment search

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.126895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.786060Z digest=sha256:1bbb84a42c4ba835a8d40ade90919e4913d9999a24fff9c65d87184ce037b58e

Observation c3534bbe-3196-4cbc-a46a-2700ccb22eaf · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.102107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.807307Z digest=sha256:e5e7eaae45587711e615170a44ed9eeadcb7c5da1f81024e76ed58a2dc6f94cd

Observation ecd4c332-ee4b-450a-9065-e6a0a012edb8 · outbound

This paper cites W., Salamon, J., Li, P., and Bello, J.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models W., Salamon, J., Li, P., and Bello, J

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.811638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.811638Z digest=sha256:890d78fae48b4cd3cee9251129a5c8f2757f2445c89f0aaf7a93e672666bef60

Observation 759585a2-6a73-4d2f-8e28-20e96b1913aa · outbound

This paper cites an unresolved cited work.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.818856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.818856Z digest=sha256:5e60a11df396369a66fbfc9f99eacc0f599b0a86f4ab8c008c5667381f1a2529

Observation 9939882c-0785-4730-93e0-b7aec7842f83 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.088472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.832434Z digest=sha256:812f90e1fbb774f975940bd82cc76888fb1a7211174367e32effffa10b7bc40f

Observation 0c5aef27-0c33-4419-b661-b77d5c82c445 · outbound

This paper cites On generative spoken language modeling from raw audio.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models On generative spoken language modeling from raw audio

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.839065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.839065Z digest=sha256:275e7242f288d1913830ba2f4faed75a101feddd2b6baf1d35d6c5f0cd55b1a5

Observation 788a30d6-ce08-4c10-98e1-02cbc78266dc · outbound

This paper cites and Hutter, F.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Hutter, F

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.847821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.847821Z digest=sha256:35850069cc844b9a38f4ee516692b348a9d24fa39a76b050dca50e50a0cf7ff4

Observation 09122a09-c6f1-4d73-bc4e-ac17bd61a597 · outbound

This paper cites Voxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Voxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.857666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.857666Z digest=sha256:f4433d1d35c8a3dfbed4be3443183af56c6e7931d5d6c57765492ea864ff64e4

Observation b3dfe765-ef91-4123-97e4-f790772ba959 · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.865712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.865712Z digest=sha256:e51fe0bd0356e1db6c2e4b3b00ac58daf15223a584787420f8f5b87cfaf8e377

Observation b5b906e3-3692-4993-a309-9a45dc5f3820 · outbound

This paper cites GPT-4 Technical Report.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.894816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.894816Z digest=sha256:b066316ac90e62c6fba4df603fe5c735d4f19315323ff970960b9aaa2c81b416

Observation a40358ec-9d19-4c10-8439-2ee44a0c04d2 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Librispeech: An ASR corpus based on public domain audio books

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.913607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.913607Z digest=sha256:45533bbe05f00cb33fd6cb26c1da86e9fa810f7b9e1afbe724892d125359807e

Observation f4f05302-5d31-4840-bb16-4c470fb3169e · outbound

This paper cites FiLM : Visual reasoning with a general conditioning layer.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models FiLM : Visual reasoning with a general conditioning layer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.917330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.917330Z digest=sha256:1e07aef3b7f92e4c1f59dd8aab64ac308088319634a5501b3104207130926bb4

Observation 6161288b-832b-4021-8b9f-338a826f5ee1 · outbound

This paper cites Train short, test long: Attention with linear biases enables input length extrapolation.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Train short, test long: Attention with linear biases enables input length extrapolation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.921466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.921466Z digest=sha256:36eb9d808246ed9002d50e37fc2a205b5815fffa9229788c3138765d6e4590ba

Observation 68b61674-7217-4229-8eb5-a2286a4006fa · outbound

This paper cites W., Xu, T., Brockman, G., Mcleavey, C., and Sutskever, I.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models W., Xu, T., Brockman, G., Mcleavey, C., and Sutskever, I

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.925313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.925313Z digest=sha256:f06b2431a830d81b3944b52c4df5871e2a1ab6b483e6c2f0191c2d5e7f387df3

Observation 8d0a9955-4f6c-4e3b-b365-c57975b17953 · outbound

This paper cites and Torre, R.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Torre, R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.026292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.931413Z digest=sha256:8354d3bf1d1a96df8b7ee452b4e15724506f27af358d168ed676d2f7c77df003

Observation 1f065bbb-92a8-42b2-af87-0012358b00fb · outbound

This paper cites and Mohamed, S.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Mohamed, S

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:47.004749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.936155Z digest=sha256:2c49465709bf8fa76b391cad3572b8c3c97d8f9bd9e2c91d06230725b6352c81

Observation b6086208-4561-4e62-a19c-57ce05649417 · outbound

This paper cites U-Net : Convolutional networks for biomedical image segmentation.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models U-Net : Convolutional networks for biomedical image segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.981677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.940000Z digest=sha256:9968e20a03c86e11672f2caef851f6c0d363f9f242b166ff64d4a0f275dae958

Observation 250938e1-419f-42c5-8657-6eeebc8e36a0 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.944189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.944189Z digest=sha256:74ff19250581db5b8b13c51ae6330198938614dc0cb75e0b5a7a7f5daa33f118

Observation ca725f33-4967-4d55-8d71-7663d8635d77 · outbound

This paper cites The interspeech 2009 emotion challenge.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The interspeech 2009 emotion challenge

Reference 39

Resolution
verified exact
doi, observed 2026-08-07T00:15:45.321783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.948242Z digest=sha256:adfb171cfbb1ab4d1a50b00e1ea8c1e0f9e596ee8191114217bb9e9a86de01c2

Observation cf588aef-9ebe-4026-b8c7-f318b7b238e1 · outbound

This paper cites The interspeech 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The interspeech 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.952054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.952054Z digest=sha256:466fdc0e5eb03fb36224dd009dd5ff3098416eb47af75628bf99ae67777e17bf

Observation 50183299-26c3-43e0-956c-9f0784219d60 · outbound

This paper cites Denoising diffusion implicit models.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Denoising diffusion implicit models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.948606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.957310Z digest=sha256:52e1cdcc126034c014120eaa0cdc0657a1e627675740f62fd7972cd3882b11cc

Observation 5d9a725c-12f5-428e-a78a-29c0ef522c8a · outbound

This paper cites J., Cao, Y., Zen, H., Rosenberg, A., Ramabhadran, B., and Wu, Y.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models J., Cao, Y., Zen, H., Rosenberg, A., Ramabhadran, B., and Wu, Y

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.963041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.963041Z digest=sha256:44341df3884823413719507fc3bae919813a64b64c51e205f883891972dd8fd8

Observation 50697df9-0e5b-420f-993f-35ef51bc8b7e · outbound

This paper cites Instance Normalization: The Missing Ingredient for Fast Stylization.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Instance Normalization: The Missing Ingredient for Fast Stylization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.968933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.968933Z digest=sha256:8f3fc9586346146ad7e930e25b0b851b9934dd6157609f6d5f8da98ffc10dbe8

Observation a02db15b-48f2-4f0c-9a22-08874b0ee3a5 · outbound

This paper cites and Kautz, J.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Kautz, J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.931145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.973522Z digest=sha256:9eb1b68c59ee6b6ee987513cabf58dcfe4e2db83ea915fa9b2457e4b431192e5

Observation b23bd89d-c775-4353-b9d6-bfd3ddf53469 · outbound

This paper cites Learning de-identified representations of prosody from raw audio.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Learning de-identified representations of prosody from raw audio

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.915433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.980070Z digest=sha256:73a3a063cfab62888e47e4b6344d6f1d2bbf87e2638ac4cff0dd905bf4b251ae

Observation 0a438fe2-7ef9-48ee-9747-3270430b9103 · outbound

This paper cites On layer normalization in the transformer architecture.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models On layer normalization in the transformer architecture

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.897756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:44.990676Z digest=sha256:1cdde388da3f867c292ab47d2793675fdaf1c700d98a7ebeadd50c46a45c624a

Observation af5c8ce8-8abd-4d61-a155-cb6747b226f9 · outbound

This paper cites CSTR VCTK Corpus : English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models CSTR VCTK Corpus : English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.997848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.997848Z digest=sha256:e302fe7781c08b5f18498397258b741c14e4dd0cd1a1bc3aa896f206d6fc999a

Observation 23b41b76-da0d-4cb5-84e0-be94071ce7f3 · outbound

This paper cites an unresolved cited work.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:15:46.880905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.002990Z digest=sha256:cfba66f77fab078323a6ea3444e0cce7ddf5effcf9fba06600adbf759f1a35af

Observation 83211cda-db71-4d3b-9aac-3d338593031b · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Soundstream: An end-to-end neural audio codec

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:45.017442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.017442Z digest=sha256:4e67ed827c019f8e8bbb0d8a01cbcceffde5fd60c8c4371e6b0360de298c5987

Observation ff1b8f43-70e1-47ae-9c2b-1067a04d88a1 · outbound

This paper cites and Sennrich, R.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models and Sennrich, R

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.864209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.044841Z digest=sha256:dc7805fcacf697d0b678973d9da341413974cf6139a3a148295d35d325294174

Observation c1c6ae91-4d1c-44e4-bfdb-9ac0a89cc508 · outbound

This paper cites SpeechTokenizer : Unified speech tokenizer for speech language models.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models SpeechTokenizer : Unified speech tokenizer for speech language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.849087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.085676Z digest=sha256:10788cf9df395614c110548b0b08b214efb42d5ff9cf4abd34fd8b6390080fef

Observation ccce29af-1127-41b4-a35f-07f657f13abb · outbound

This paper cites R., Kadav, A., and Graf, H.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models R., Kadav, A., and Graf, H

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:15:46.828978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T00:15:45.092264Z digest=sha256:87237b16ff4de0e1f9316d6f4a98810ecb2267a59ecef944a70f133253b749d1

Observation 82f83d34-0588-4528-8c2b-b1d7d84b1214 · outbound

This paper cites @esa (Ref.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models @esa (Ref

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:45.106503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.106503Z digest=sha256:d6ab3fbe2a048933d785279505c389f5dfbc78fff6ca4333dc85362c22257595

Observation 17f572f4-8e35-4aa1-9047-dbed4a13db6d · outbound

This paper cites an unresolved cited work.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:45.135549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.135549Z digest=sha256:f29a16b57ad7a8d8419adbf504e914318f32143a88f14a7f2487728d4e5709cc

Observation 0bee6347-6538-413a-968a-799224ee8119 · outbound

This paper cites The evaluation metrics are detailed in Section ssec:eval-all.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models The evaluation metrics are detailed in Section ssec:eval-all

Reference 55

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:15:45.148144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:45.148144Z digest=sha256:db5afb7de598978221997c255daf7af7c2ec562e5d133ee2987467f731eb5d7b

Pith citing papers

Observation 8206ffa6-12e9-41c5-88f8-80bacbdf1fd3 · inbound

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias cites this paper.

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias A Variational Framework for Improving Naturalness in Generative Spoken Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.526798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:36:21.635792Z digest=sha256:644385514a62b9b98fb86615c9ad1f7b83990cd4f61aca638d2b9d6176988f2d