Pith. sign in

Paper Citation Record · LEDGER

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.01391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01391 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:40:10.557940Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ff54f5e-682b-4d85-bd31-17cf031384a9 · outbound

This paper cites A Brief Overview of Unsupervised Neural Speech Representation Learning.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model A Brief Overview of Unsupervised Neural Speech Representation Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:40:10.681418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.427945Z digest=sha256:94fbd911520522af0930cd64024800991586a244626ec03533e125e1486506ac

Observation 90f5f69d-626c-4cd5-ae28-15d4a175a960 · outbound

This paper cites Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.066460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.433506Z digest=sha256:2eec218ee047c4e9dc8432a3087f7373c4f8d3c079be858ab6ef6f89450d8279

Observation 8e44e2db-0d4c-446a-8314-cd30ad8b20d0 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.049538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.439036Z digest=sha256:588195c3c0f72247a1d07db9e8efa2e1f52871f5e6f2a7958b354b56426ca362

Observation 9172d4dd-ad1a-421f-b1c9-3208c80ac016 · outbound

This paper cites On generative spoken language modeling from raw audio,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model On generative spoken language modeling from raw audio,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.031437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.444464Z digest=sha256:11fd11fce6dc544ed54ae9011ee19dac90c24629c8ef7d3c30ce2122fab53083

Observation e34ba3e0-1027-4ac8-8070-ab1221fa4917 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.013639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.450324Z digest=sha256:b92ce833a8ae032634dabc55ca81941176e45997a7bcce886bafc25aead26155

Observation c6abd5db-d1d3-4674-9079-3c8f01a02db8 · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.996936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.455036Z digest=sha256:a04e57ac8e6e8a876fceeab8a71a97a5723afb513b61ffccc6490379e523e180

Observation 5e76a905-5d1d-4d2e-a1bb-11c2f2f79f8e · outbound

This paper cites A Unified Accent Estimation Method Based on Multi-Task Learn- ing for Japanese Text-to-Speech,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model A Unified Accent Estimation Method Based on Multi-Task Learn- ing for Japanese Text-to-Speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.979543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.460355Z digest=sha256:6b7b8b2c4d8879c30855ad44decbf5d0d0739d06e75a886dd7018c09c8004b7a

Observation c4c91880-fe62-47fb-9d30-c05813665721 · outbound

This paper cites Reazonspeech: A free and massive corpus for Japanese ASR,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Reazonspeech: A free and massive corpus for Japanese ASR,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.961640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.465377Z digest=sha256:386d2ff277b2c5a56f6b9b3617685e987e96591a13f67c526551a9979a9491c5

Observation daf8e8ad-0319-407d-99de-c4f775461830 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to- text transformer,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Exploring the limits of transfer learning with a unified text-to- text transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.943901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.469772Z digest=sha256:447453fd2da31fa97272c56e6849bde36260cc53ddd7b87272e12fc58fbd0b42

Observation 4326f94e-a02e-415e-a5c6-8d8e5a9fc790 · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.474344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.474344Z digest=sha256:d3964f7e7e2784c637e05fea67337961ef638eb3e849b1a8e7f510b500c46ed4

Observation 7237d5cf-fe98-43df-a014-99330b8e6193 · outbound

This paper cites JVS corpus: free Japanese multi-speaker voice corpus.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model JVS corpus: free Japanese multi-speaker voice corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.479287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.479287Z digest=sha256:4bf010779486da479110d68144e2b0b4bc0bfd548be4bd7d99a38396521dbedd

Observation 02a76c6d-7964-40a0-8132-93a3e365dd55 · outbound

This paper cites Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.926464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.484916Z digest=sha256:dce514897fc33099a04d86ac95960e21fa3986879952c519c7087849785a4649

Observation 9d4ba3ca-6227-4aa6-8e29-535cc43dc83b · outbound

This paper cites J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.908829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.489843Z digest=sha256:68b8eee299753f40edc81a1cc3276cb64802c0a16d9cb232f34e4b9dbe0ffc16

Observation f41b6dad-0f1d-4279-8df4-09aa58a892d9 · outbound

This paper cites JSSS: free Japanese speech corpus for summarization and simplification.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model JSSS: free Japanese speech corpus for summarization and simplification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.494546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.494546Z digest=sha256:0bbe22c8d80e5eae5c105d8b3e5272b8a7eb064dc5e0148c1aafc831e8e618d9

Observation 0d808ecc-efea-487b-9f18-88eb319e06ef · outbound

This paper cites T5g2p: Us- ing text-to-text transfer transformer for grapheme-to- phoneme conversion,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model T5g2p: Us- ing text-to-text transfer transformer for grapheme-to- phoneme conversion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.890489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.500200Z digest=sha256:0a56f00a8e750991b8911159caf4fcc2d9d3ee7f99cce8492982c76a98707772

Observation 5af588d6-e2ec-44fe-819d-ace2c0e3a93e · outbound

This paper cites SpeechT5: Unified- modal encoder-decoder pre-training for spoken language processing,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model SpeechT5: Unified- modal encoder-decoder pre-training for spoken language processing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.872317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.504454Z digest=sha256:665df64d5aaebb676b3f0592863241552eb9a0e388d6c8bc8648954140196510

Observation 36323435-2a5d-4930-81db-cf68a5d70b52 · outbound

This paper cites Neural ma- chine translation for multilingual grapheme-to-phoneme conversion,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Neural ma- chine translation for multilingual grapheme-to-phoneme conversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.856534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.510410Z digest=sha256:522d9e30faa49c326c54ad63a282c74ca1d95ed6992025d7d9c9fbe34ba84fa2

Observation 13135d83-6c53-4f40-81dc-c563038ebeb5 · outbound

This paper cites One model to pronounce them all: Multilingual grapheme- to-phoneme conversion with a transformer ensemble,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model One model to pronounce them all: Multilingual grapheme- to-phoneme conversion with a transformer ensemble,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.838454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.515754Z digest=sha256:6df63ce45196461c12cd3c39dbcdc608d193f6dc6cc426b8ebf2f74b9095feaf

Observation d50b1644-9ac2-43ee-bb72-12798424b365 · outbound

This paper cites Byt5 model for mas- sively multilingual grapheme-to-phoneme conversion,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Byt5 model for mas- sively multilingual grapheme-to-phoneme conversion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.820105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.520427Z digest=sha256:43866976b1baeb3ba3de153687ea33aafc2848b807c9ba0c2a59d0191046599d

Observation 369ce2af-1ef8-40ed-84bf-32853dd86534 · outbound

This paper cites ContentVec: An improved self-supervised speech representation by dis- entangling speakers,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model ContentVec: An improved self-supervised speech representation by dis- entangling speakers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.803937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.524640Z digest=sha256:8cbc3a9f8ab01918679b171e42622a33f6ead7442bf4971aa5130e4c713c54e6

Observation 1257031f-77b7-42fc-a4b6-f15c4431988c · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.529671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.529671Z digest=sha256:c45a7b4d9d443e963eaa08c70fda87b87858699c7bbb57bbb0fc24d65d9f058a

Observation b7c8b8da-ac01-4958-b309-6a285538d60a · outbound

This paper cites Radford, K.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Radford, K

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.785436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.534763Z digest=sha256:5985c5f225ca8fe04de3e98bbbe63218c894819f13afa4f239ad4c1284f1e43e

Observation e515bfbf-108a-4b89-929a-49b24a59a620 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for VoiceMOS challenge 2022,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model UTMOS: UTokyo-SaruLab system for VoiceMOS challenge 2022,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.768827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.539047Z digest=sha256:f6a8c112b8d0b333e608066edc8ac4f06e8a54b5a40490bdc63199cbdecf7045

Observation 2ae9a126-9e21-46af-a0d5-77bf7ccb7e17 · outbound

This paper cites Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.750702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.543959Z digest=sha256:ca59463f94990b0d4eae99b81036b875bd506ea1063ad24304925ca00676eacb

Observation 90e04134-54e2-4c17-8272-80191bbcc018 · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.732694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.548721Z digest=sha256:fe04a67edb2c28a25779f7b6b2a63ae4133352152c9f2087f73d26ae71d20b93

Observation 07250b96-e2a1-4327-9ce1-90001fc03b31 · outbound

This paper cites mT5: A mas- sively multilingual pre-trained text-to-text transformer,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model mT5: A mas- sively multilingual pre-trained text-to-text transformer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.713327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.553662Z digest=sha256:b21f300c935cfaac03900ae5178f84d8e59347bd5d1d08663fe1512ac14301f0

Observation d9cf1ded-3c45-4a9f-957d-e7bd01a8cb69 · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte-to-byte models,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model ByT5: Towards a token-free future with pre-trained byte-to-byte models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.697841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T12:40:10.557940Z digest=sha256:3e63ae33a7e7fc6a8d5043485dfc1ebab9437f8b2c3ec10a966f70ad77fd67f0

Pith citing papers

No inbound Pith citation observations are available.