Pith. sign in

Paper Citation Record · LEDGER

EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2308.05725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.05725 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:40.366858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:28.620667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d5f7645-2d04-4d74-be77-c2ff20279371 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:27:25.612037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:48569223e7e749b1eb9a748ce7600fd29cdf15feb397929f9ed93ba3d6b5fbc2

Observation 742884e6-2cf3-47cd-8ac0-ae6d3c53620f · inbound

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection cites this paper.

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:40.366858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:40.366858Z digest=sha256:74341e51e2207430b5a9b974fa0ac4087ba87be70bd1b640a4bb7f72bcbdc5ac

Observation 5ab751e6-8225-4415-82ca-520f8cb83fcc · inbound

StressTest: Can YOUR Speech LM Handle the Stress? cites this paper.

StressTest: Can YOUR Speech LM Handle the Stress? EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.259217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T13:00:23.002962Z digest=sha256:6f28f80a788670ca36e5f00d34a82d5f160d6a53154e75978d7a42226ec9a9ea

Observation a5e2b151-c0e1-4e2c-beb8-922275a56e54 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.345371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.345371Z digest=sha256:8e929c9fa1010c9ea44638b9bbf7d096f32e67e73f50fea976d3cdad1416e4c7

Observation 79602e4a-bfe0-4eec-9975-1136cfc0f26c · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.438991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.438991Z digest=sha256:da9e11b728587ea20314a4e17452c674a5247b4a399956232259809be2f5d3a8

Observation 4402f487-40f7-461f-be2d-02e40eaa173d · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.393989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.393989Z digest=sha256:f5bf970dc8302c1fb1d7b1951d3f04a4a3f6c5afb08bef1e869f8db78e00bab5

Observation 2bbcfe48-1e89-402a-ad67-8a90aa46b3da · inbound

Computational Narrative Understanding for Expressive Text-to-Speech cites this paper.

Computational Narrative Understanding for Expressive Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:31:47.053839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T19:30:53.250247Z digest=sha256:696408cee8006724bc8f38ed8ccd8a6fd10da84ee9fb26b9258d5a275a4dc024

Observation 8a8b3a10-c5d3-42aa-bd5c-3168ab4bdcfe · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.344834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:8ab636f9ea9dcb48b692991df714cabc7ae3f54f4544acfd47cf306153bfeefa

Observation 0252de93-6edd-4e86-b63c-26a9a7e93e21 · inbound

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders cites this paper.

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:28:01.655676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:28:01.655676Z digest=sha256:478a2c6a2615b31c5d12061e2cbce4de6dacb9402c64a1138f8b7c3fee29e1e3

Observation d8ea44ac-40cb-4828-8173-552871608972 · inbound

Voxtral TTS cites this paper.

Voxtral TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.896911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:ef67de04a7d81c8d68ded834f18d0bedd81584eee9ced985fbe0166c25f1bbba

Observation 549b17d1-640d-431b-8730-8f5d3ea72df8 · inbound

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions cites this paper.

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:08.380039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:43:33.397158Z digest=sha256:4aed9128fbbe83f8d7fc4b284acfd7ef9d5f825fb197ced47821624cab3b68fb

Observation 0661bbb5-e44f-4264-8a08-e077c33939ac · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.471353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:f540067236365b09a6b932ca11897cfb33bf8cf00c9c10f2e8cfd3f7c862a1c0

Observation ff5f36a2-5395-4982-b1d5-1428c3a7f1f3 · inbound

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control cites this paper.

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T20:53:57.753462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T20:53:47.252916Z digest=sha256:05d7fb8596ad9031d9777d74c7fa53f39016830849a036de01c0cb83c69ad17e

Observation b73b78d3-4db9-431a-b7d5-8d14dca35cc9 · inbound

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS cites this paper.

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:16:11.068632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T21:31:09.399885Z digest=sha256:3943e292f1b0c1c045683b1c1b82570a4880fe4bd1d601c80bacf11ff1b8884d

Observation 0ef35b23-59a6-4ae5-b80e-62ec4283707a · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.556094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:e0624fd4067c656e268949f1fbbb2d6d0a748306e9c8201c1e8a6fc3170176af

Observation 483f3ab7-4cda-43af-8e3c-9493413f6c44 · inbound

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning cites this paper.

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.622473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T14:35:17.004680Z digest=sha256:c402cb28f264ad1ff64f3216d1158fee8e54ced43ed1092047b578010856fb1e

Observation 612af17a-827e-4187-b885-0b712fcafe7d · inbound

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants cites this paper.

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:30:34.179334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:30:34.179334Z digest=sha256:53623020ffb0f00ee37344d0a31de7435bdc3a9993d6f3e41e0af35400b61b45

Observation 256e3517-1a63-4ca0-b573-407c72110a94 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.474782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.474782Z digest=sha256:c1c9dfd24c5c1b23b8dab510227af42258506fc98a9ba2b98bd1337c23510623

Observation fec8ed68-93ac-427b-a97e-d3dea908ede9 · inbound

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces cites this paper.

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:33.274250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:33.274250Z digest=sha256:f2cd6374bf12b92c731c732b1cd3c47b3cf965e0555f02cf9d75dd4b8f887ce4