Pith. sign in

Paper Citation Record · LEDGER

EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2308.05725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.05725 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:47:51.527155Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:28.620667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2a8e21e-e2b5-44d8-babd-9337ca96c34f · inbound

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram cites this paper.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.527155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.527155Z digest=sha256:ae5438b4ba3230f24bb02aa5121ad3f89eb7ed2821a7ca8ed7010d628cd1f239

Observation ee9a3be0-a176-45b5-bda9-b2bdf0e8327e · inbound

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios cites this paper.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.997721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.997721Z digest=sha256:81838522f4060efb72cfa5cf7293af23d35c29ac9a7e7ae506e0d18650b2970b

Observation 859b9f78-d0d5-48ab-a872-84b4d20ea4e9 · inbound

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles cites this paper.

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:36.249771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:41:36.249771Z digest=sha256:725e3f589c31580bfdf7abd2901e3e2163b4cafa4321aa88e02a0f422f7f66ee

Observation cbbfc07c-f8c8-4409-aeea-5b45fd9e2ecd · inbound

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution cites this paper.

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:35.369528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:27:35.369528Z digest=sha256:d8b1519718c5690a43e8ec5217ee3715a63c60ad4545b562a4e6420dfaf565bf

Observation 0d5f7645-2d04-4d74-be77-c2ff20279371 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:27:25.612037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:ff86158bb06982efac34aa2def7798a2ba0c4ef7cad9b48e1f24fb5960d6e731

Observation 742884e6-2cf3-47cd-8ac0-ae6d3c53620f · inbound

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection cites this paper.

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:40.366858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:40.366858Z digest=sha256:c035bbdb01e50949ee80104a6676f3f7105ab0e54ec87a8c40a501fc61657147

Observation 5ab751e6-8225-4415-82ca-520f8cb83fcc · inbound

StressTest: Can YOUR Speech LM Handle the Stress? cites this paper.

StressTest: Can YOUR Speech LM Handle the Stress? EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.259217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T13:00:23.002962Z digest=sha256:a1e0a00936c2e19149febefc3d37dfd5bcc33c3a41e44dcec86c2d9c48fc96e2

Observation a5e2b151-c0e1-4e2c-beb8-922275a56e54 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.345371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.345371Z digest=sha256:1318eb2c619c09a7fe5c88f1e350f715576a399396574e989b6a4deeaf21faac

Observation 79602e4a-bfe0-4eec-9975-1136cfc0f26c · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.438991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.438991Z digest=sha256:707030b46fce2e1a82a089c09060d6d1a6783bbd6b3c752e95006510db1547c2

Observation 4402f487-40f7-461f-be2d-02e40eaa173d · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.393989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.393989Z digest=sha256:10e7d8286c00d4f4c9bb8321a626cc346c0fff7a8cf7a20ccce713b1388fc605

Observation 2bbcfe48-1e89-402a-ad67-8a90aa46b3da · inbound

Computational Narrative Understanding for Expressive Text-to-Speech cites this paper.

Computational Narrative Understanding for Expressive Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:31:47.053839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T19:30:53.250247Z digest=sha256:623e3e8fe3806e5bac9f5d4b74e2f23b603e045ddf9e84b62579835f53669076

Observation 8a8b3a10-c5d3-42aa-bd5c-3168ab4bdcfe · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.344834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:1b6bfb9c55e38b0382b535f1dfd32050ba4b4b181b2aa8392a88d875d1fdf3f8

Observation 0252de93-6edd-4e86-b63c-26a9a7e93e21 · inbound

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders cites this paper.

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:28:01.655676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:28:01.655676Z digest=sha256:a1079bb26e7c73c44437de6a370af8821efc425a400a162b68913f61a5c69940

Observation d8ea44ac-40cb-4828-8173-552871608972 · inbound

Voxtral TTS cites this paper.

Voxtral TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.896911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:639a1833ae62bce6bd05ad213780a70f7788e3689bb5872968a4007bbb955cf1

Observation 549b17d1-640d-431b-8730-8f5d3ea72df8 · inbound

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions cites this paper.

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:08.380039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T16:43:33.397158Z digest=sha256:c28c41e8719b80490e6ae4186b14d5ab4c453fa742ebcb0bda4d4877af793261

Observation 0661bbb5-e44f-4264-8a08-e077c33939ac · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.471353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:099fe4b447ac40585d1dd06102571ad324cad6aff0f5895c7ff889e93cc337d2

Observation ff5f36a2-5395-4982-b1d5-1428c3a7f1f3 · inbound

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control cites this paper.

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T20:53:57.753462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T20:53:47.252916Z digest=sha256:4094e36b6eb33ebf55f5bb29dc61432d11ad77e75ed5e1505b1b42957c6c4e5e

Observation b73b78d3-4db9-431a-b7d5-8d14dca35cc9 · inbound

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS cites this paper.

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:16:11.068632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T21:31:09.399885Z digest=sha256:49b7912d3c1192563e02420eb87bd4e9b75cb9c51f9096900e6d0ec18cea8f0a

Observation 0ef35b23-59a6-4ae5-b80e-62ec4283707a · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.556094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:1d03d707bc7d2dfc866f11d8c4a3902653d366ed2233fd6c7e3916256a3a1943

Observation 483f3ab7-4cda-43af-8e3c-9493413f6c44 · inbound

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning cites this paper.

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.622473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T14:35:17.004680Z digest=sha256:a82b3c2aebe566aa9c0fd7a7ec5f2d0f1ae774e347775d86c40145f8a2a226b4

Observation 612af17a-827e-4187-b885-0b712fcafe7d · inbound

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants cites this paper.

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:30:34.179334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:30:34.179334Z digest=sha256:4b089a1fb07f2e9e0c1a1be0a4c2402e5f8f3e7f2f47f601be70955012a84cf7

Observation 256e3517-1a63-4ca0-b573-407c72110a94 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.474782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.474782Z digest=sha256:ad5e51ab6d3f3e3cdd696d475a1d21fb5b203159ae68345da9b069a41ba07ac4

Observation fec8ed68-93ac-427b-a97e-d3dea908ede9 · inbound

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces cites this paper.

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:33.274250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:33.274250Z digest=sha256:337cae9bde1cdb39ddfbe4a7fec534688e52b4b4af3621879d98344dfacdc689