Pith. sign in

Paper Citation Record · LEDGER

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.01391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01391 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:40:10.557940Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ff54f5e-682b-4d85-bd31-17cf031384a9 · outbound

This paper cites A Brief Overview of Unsupervised Neural Speech Representation Learning.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model A Brief Overview of Unsupervised Neural Speech Representation Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:40:10.681418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.427945Z digest=sha256:bba9bc4f8681b97e23d363ede18e429e67a97aa4de5a55bec95eb5bb090deee9

Observation 90f5f69d-626c-4cd5-ae28-15d4a175a960 · outbound

This paper cites Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.066460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.433506Z digest=sha256:2825dce368b164c5f2dd4e8dc57596c2e12bd8b4e8353778d328ba98cef74a73

Observation 8e44e2db-0d4c-446a-8314-cd30ad8b20d0 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.049538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.439036Z digest=sha256:d987d9d3d5d743f1cdaae8fff23f5cfdf345e3d2949ddb3d11bb489a4ce92621

Observation 9172d4dd-ad1a-421f-b1c9-3208c80ac016 · outbound

This paper cites On generative spoken language modeling from raw audio,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model On generative spoken language modeling from raw audio,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.031437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.444464Z digest=sha256:d30b556feb3de1048dcfcdd3bc20c1179ed91aa620bb8ed24024a0a1f376dc7e

Observation e34ba3e0-1027-4ac8-8070-ab1221fa4917 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:11.013639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.450324Z digest=sha256:cf5c8dc5a35ea2d086df561ad3fab331aff146fc2f0a000f4f572d255975c016

Observation c6abd5db-d1d3-4674-9079-3c8f01a02db8 · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.996936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.455036Z digest=sha256:7806396ae2bae9892d7f297537d55bab140873a86a05db5d3347253c9b25b112

Observation 5e76a905-5d1d-4d2e-a1bb-11c2f2f79f8e · outbound

This paper cites A Unified Accent Estimation Method Based on Multi-Task Learn- ing for Japanese Text-to-Speech,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model A Unified Accent Estimation Method Based on Multi-Task Learn- ing for Japanese Text-to-Speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.979543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.460355Z digest=sha256:aa6129010244cbbd20556bc5427090fc0ea939882d4d8f950b4a21db650bacd9

Observation c4c91880-fe62-47fb-9d30-c05813665721 · outbound

This paper cites Reazonspeech: A free and massive corpus for Japanese ASR,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Reazonspeech: A free and massive corpus for Japanese ASR,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.961640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.465377Z digest=sha256:fd4be1e44eb96d05627bdb23e2f537c65b780699eca8e380a380fab84320ee61

Observation daf8e8ad-0319-407d-99de-c4f775461830 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to- text transformer,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Exploring the limits of transfer learning with a unified text-to- text transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.943901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.469772Z digest=sha256:74c20d45fb102b77bedb1b3004e40940d9a811bc8cbeada09ccc5f42bc67bcd7

Observation 4326f94e-a02e-415e-a5c6-8d8e5a9fc790 · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.474344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.474344Z digest=sha256:d3964f7e7e2784c637e05fea67337961ef638eb3e849b1a8e7f510b500c46ed4

Observation 7237d5cf-fe98-43df-a014-99330b8e6193 · outbound

This paper cites JVS corpus: free Japanese multi-speaker voice corpus.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model JVS corpus: free Japanese multi-speaker voice corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.479287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.479287Z digest=sha256:4bf010779486da479110d68144e2b0b4bc0bfd548be4bd7d99a38396521dbedd

Observation 02a76c6d-7964-40a0-8132-93a3e365dd55 · outbound

This paper cites Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.926464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.484916Z digest=sha256:d6931d320b34e5942fdb9d7cb296855c0e6ee03ae5dd325e83528e7044e6d612

Observation 9d4ba3ca-6227-4aa6-8e29-535cc43dc83b · outbound

This paper cites J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.908829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.489843Z digest=sha256:4643fe70b7fb088e291df1ebe11320aecfda433f73c07fbd3d90abe0a4d711e6

Observation f41b6dad-0f1d-4279-8df4-09aa58a892d9 · outbound

This paper cites JSSS: free Japanese speech corpus for summarization and simplification.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model JSSS: free Japanese speech corpus for summarization and simplification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.494546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.494546Z digest=sha256:c948a1250e5c2f9c0042e0f5afe7cf1d5fbae77256965d28c95a5fb7042036b0

Observation 0d808ecc-efea-487b-9f18-88eb319e06ef · outbound

This paper cites T5g2p: Us- ing text-to-text transfer transformer for grapheme-to- phoneme conversion,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model T5g2p: Us- ing text-to-text transfer transformer for grapheme-to- phoneme conversion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.890489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.500200Z digest=sha256:63d4d77988fa4535f032a82fdb6dd948df4f29c89feec9b0183e8fc41faa8fe5

Observation 5af588d6-e2ec-44fe-819d-ace2c0e3a93e · outbound

This paper cites SpeechT5: Unified- modal encoder-decoder pre-training for spoken language processing,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model SpeechT5: Unified- modal encoder-decoder pre-training for spoken language processing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.872317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.504454Z digest=sha256:8d20a76abc08538789aab822a3567b2749b86d88a81a8b3e52e51d50db79416a

Observation 36323435-2a5d-4930-81db-cf68a5d70b52 · outbound

This paper cites Neural ma- chine translation for multilingual grapheme-to-phoneme conversion,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Neural ma- chine translation for multilingual grapheme-to-phoneme conversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.856534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.510410Z digest=sha256:76e6852e7e0f16284d6f03e7dac086086fc9453b2cc3ae54bb65603317c33a8c

Observation 13135d83-6c53-4f40-81dc-c563038ebeb5 · outbound

This paper cites One model to pronounce them all: Multilingual grapheme- to-phoneme conversion with a transformer ensemble,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model One model to pronounce them all: Multilingual grapheme- to-phoneme conversion with a transformer ensemble,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.838454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.515754Z digest=sha256:5c6628d6fbb9bb3f7396be05669fbfa8e96e77cbb7b440855cadf1aac2456882

Observation d50b1644-9ac2-43ee-bb72-12798424b365 · outbound

This paper cites Byt5 model for mas- sively multilingual grapheme-to-phoneme conversion,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Byt5 model for mas- sively multilingual grapheme-to-phoneme conversion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.820105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.520427Z digest=sha256:6178e9de70b96e39d80266c622a5084b2f4de62788df0d74578c8abd35e5f23c

Observation 369ce2af-1ef8-40ed-84bf-32853dd86534 · outbound

This paper cites ContentVec: An improved self-supervised speech representation by dis- entangling speakers,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model ContentVec: An improved self-supervised speech representation by dis- entangling speakers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.803937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.524640Z digest=sha256:90f028521175759e4ae064db76cead7a158379dfa3cc2cf71bef8cccc26cc99d

Observation 1257031f-77b7-42fc-a4b6-f15c4431988c · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.529671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.529671Z digest=sha256:c45a7b4d9d443e963eaa08c70fda87b87858699c7bbb57bbb0fc24d65d9f058a

Observation b7c8b8da-ac01-4958-b309-6a285538d60a · outbound

This paper cites Radford, K.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Radford, K

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.785436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.534763Z digest=sha256:64120c5815f32a15bbe459cb3f069b0436d0e7f02d5692b0eb6db68bb7674d3e

Observation e515bfbf-108a-4b89-929a-49b24a59a620 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for VoiceMOS challenge 2022,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model UTMOS: UTokyo-SaruLab system for VoiceMOS challenge 2022,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.768827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.539047Z digest=sha256:ef62ca37836321bea29644783e382dfcded77fe79da9a2fe4d62f20e900c2fd1

Observation 2ae9a126-9e21-46af-a0d5-77bf7ccb7e17 · outbound

This paper cites Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.750702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.543959Z digest=sha256:ae58026335da83da937466b63606366ca2523f836d5631bc7dc24dfbdfe38413

Observation 90e04134-54e2-4c17-8272-80191bbcc018 · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.732694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.548721Z digest=sha256:74e33d3a58ea583b846080b55802cc4cded67d928f9df26fd74dd2d7de14685a

Observation 07250b96-e2a1-4327-9ce1-90001fc03b31 · outbound

This paper cites mT5: A mas- sively multilingual pre-trained text-to-text transformer,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model mT5: A mas- sively multilingual pre-trained text-to-text transformer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.713327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.553662Z digest=sha256:a2dcd2af9ffa94401a2e712e876ce98fc10986e9d6506a8b39ef364e46bc836d

Observation d9cf1ded-3c45-4a9f-957d-e7bd01a8cb69 · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte-to-byte models,.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model ByT5: Towards a token-free future with pre-trained byte-to-byte models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:40:10.697841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T12:40:10.557940Z digest=sha256:13bb7a2fd09dea662ddc8d42783f7d3f1bc93cb9840606810d8f02d62188f983

Pith citing papers

No inbound Pith citation observations are available.