Pith. sign in

Paper Citation Record · LEDGER

Marco-Voice Technical Report

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2508.02038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02038 v4

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:17:06.772619Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T12:54:20.815371Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:06:23.916814Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact4
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d657cd9c-ea7a-489f-b957-772d85182335 · outbound

This paper cites Barakat, O.

Marco-Voice Technical Report Barakat, O

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:17:08.544656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:05.606977Z digest=sha256:2b29237d574d41d01b206555c6a624ab97ff975ef9eb3ab53d19df55b4aee43d

Observation a6328654-ac4e-4765-a3e6-51717d90e619 · outbound

This paper cites EmoSpeech: Guiding FastSpeech2 Towards Emotional Text to Speech.

Marco-Voice Technical Report EmoSpeech: Guiding FastSpeech2 Towards Emotional Text to Speech

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:05.775601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:05.775601Z digest=sha256:2ec742af970917d53e9ac39dd2def6720f337ed696a4d98f2b5a55d09df7f2ce

Observation 30badf42-c085-493f-ba47-67cea15fd789 · outbound

This paper cites Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis.

Marco-Voice Technical Report Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:17:07.522436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:06.064532Z digest=sha256:8f00c9c3d40710a21c83ac77558feb796c7ddd145a5a0d631660b0b7821c8dc6

Observation 5128c4bc-6773-4381-bdbf-52782b84470e · outbound

This paper cites EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark.

Marco-Voice Technical Report EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.215755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.215755Z digest=sha256:39f98e823f37f43cbfebcc00fcfab1d580b5f22e88477e1e3b34c0c09a9fd062

Observation 39a25de5-0fe5-4e27-b05d-6ca480cb2e41 · outbound

This paper cites DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation.

Marco-Voice Technical Report DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:17:07.107479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:06.334514Z digest=sha256:a50fae98beb09a47c5956e2b078eb04b9db9f6608fccc4e2887aa771440512ec

Observation 3dd2c9a9-4ed4-4bde-be05-e8b4c8dbfb64 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Marco-Voice Technical Report NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.401559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.401559Z digest=sha256:6fd7052cdc93378ec0d79b5f8ec2bad4ac9cc4d152520621e76e5438dfe204de

Observation 6834a970-203c-428b-8fbb-b7a01beff663 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Marco-Voice Technical Report A Survey on Neural Speech Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.469172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.469172Z digest=sha256:9a1c527469103fadef934b11f13805e1183103b8c43aab1144a6966e6148ddf1

Observation 9505d7d4-4fef-46d4-9a41-e7b1e62bcb5b · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Marco-Voice Technical Report Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.530717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.530717Z digest=sha256:9db7c18626e3e6c2c270cf0f2cee572e8eca3eecc0e64e86abbce6d72b3b4254

Observation 605f7043-c24c-42d9-82e5-de471472d079 · outbound

This paper cites A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI.

Marco-Voice Technical Report A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.655722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.655722Z digest=sha256:e1aeb91f46c98db0537d8e29a53779a78af93b85c25769a9c78e4155c7379915

Observation e0776a0c-c7fe-4d93-9961-eff89e5c652a · outbound

This paper cites an unresolved cited work.

Marco-Voice Technical Report Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:17:08.073198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:06.772619Z digest=sha256:db37d63de2f63154333275c856e5a51929cd2bb1307007f2dddf47fcd42fe1e8

Observation 88e9e5bb-5305-4156-89a3-d955a6123bc4 · outbound

This paper cites an unresolved cited work.

Marco-Voice Technical Report Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:17:08.328850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:06.599240Z digest=sha256:451f60da06ba3ab5dc5f40c026c85036a9b3bd71ff309c6202bfd01de99b58b1

Observation 867aa1ca-1fb9-4924-9e32-ac2e1e0b9e10 · outbound

This paper cites Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech.

Marco-Voice Technical Report Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:05.972642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:05.972642Z digest=sha256:71e88c8e22be12a3566379fb22541ece2d1afdb861d34ff556e74f5d41be044b

Observation 38364a48-0d8c-4fa9-956d-10e6cce7a254 · outbound

This paper cites Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models.

Marco-Voice Technical Report Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:17:07.314148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:06.124035Z digest=sha256:5635960d77f59a5dc39e5e442d4040b672c72d0a6059f3bb9dce1b0cc93cf805

Observation 7272df6f-6938-436b-b387-22b64fdb9620 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Marco-Voice Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:05.882842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:05.882842Z digest=sha256:44c7b5302a0dd2da84d46d620600f73d87eb8ed82cf060dc7868b5390b04fabc

Observation 171d1df3-2360-4aa0-bc5b-a4d07157c2bd · outbound

This paper cites ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency.

Marco-Voice Technical Report ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:17:07.837325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:17:05.681978Z digest=sha256:bd9f0f0398107552e91c450bb0c207370750844542c99b7b76632b29eba2df1a

Pith citing papers

Observation c5e8dea3-6b19-48c0-a1b7-2b0eca1b5b19 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Marco-Voice Technical Report

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.323332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:6b5619224164aabc5284664afc267d4dd3f0ec60e7fc83ab08de616e481d4342

Observation ae97962b-e266-4109-8d59-ed4aa6f9e464 · inbound

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing cites this paper.

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing Marco-Voice Technical Report

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:06:23.918542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T12:54:20.815371Z digest=sha256:87b5156fc67b8b42fed2cba424e30a3ad385b3841b883f62471a2b53e8f722aa