Pith. sign in

Paper Citation Record · LEDGER

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2506.04527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04527 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:15.786786Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:15.390168Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:45:15.943522Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d183f01-308a-4a4d-b3a1-f50a73b0244b · outbound

This paper cites For training high-quality and diverse-styled TTS models, a large amount of text-speech paired data is re- quired [4, 5].

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning For training high-quality and diverse-styled TTS models, a large amount of text-speech paired data is re- quired [4, 5]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.968584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.376258Z digest=sha256:573b87e6c64ad93585dfe0f98e801e19e28589272450407faa738c861bbac35e

Observation 3828ef2d-8103-4827-9723-0e4ff294770b · outbound

This paper cites Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:45:15.953103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.390168Z digest=sha256:69f1315c28943173107ca069de8a728e0e50f4e08d23785538f410f55cefa20c

Observation d83c57ff-e9c2-4051-a2b8-5ffde5a2fcff · outbound

This paper cites <blank>.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning <blank>

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.947649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.397152Z digest=sha256:ee8525e2a83ec33f78054fb5ad0288bf2bc55b86fb5f0cfbc013f871066c33ea

Observation 283541ee-6f26-4a93-adb5-f166d39e8b93 · outbound

This paper cites Evaluation of proposed annotation model 4.1.1.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Evaluation of proposed annotation model 4.1.1

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.928154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.408097Z digest=sha256:a2c6951429119ef0cbe824bc56de967e7b2bed1282cb7df732b2596efb6fb7d1

Observation dd5fb3c8-6404-4130-b3ef-a50ee30c9db0 · outbound

This paper cites Specifically, we utilized the basic5000 subset along with its manually annotated TTS labels1.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Specifically, we utilized the basic5000 subset along with its manually annotated TTS labels1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.903592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.416347Z digest=sha256:889c64cbd1a3d16b0c76c4f702469a4f2711e303bea3078a7270a412ea5b14c1

Observation 88bb5455-04e4-464a-b6ce-7b7f26498546 · outbound

This paper cites ”, (2) Accent change from low to high “[.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning ”, (2) Accent change from low to high “[

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.884422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.422376Z digest=sha256:a159ddcfe5a023f444ec74a658a7e31da8ba4536dd858f2f04e39f0beb0c7ed3

Observation 8a7ac17f-3c7a-49fd-b1c8-b4a1f7ba2937 · outbound

This paper cites Applying this method to downstream tasks be- yond textual accent estimation is a challenge for future work.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Applying this method to downstream tasks be- yond textual accent estimation is a challenge for future work

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.865438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.428341Z digest=sha256:f5577873cd636cf803aebf4ffbdde987a887c1f9a9d7dc2416f7b7e35d9ac681

Observation 471f62fb-b9d2-4ecb-8445-85cff0353220 · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Tacotron: Towards end-to-end speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.843284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.436595Z digest=sha256:2807e636fc8033a00a1df584c06c6f69acd2d85ccb55d2d7a6c3b81a8bd1dd98

Observation 0d18b20c-8ab5-4230-87dd-921714906ef3 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.820773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.444523Z digest=sha256:dbbbee2b327a7622dabd8a25d6169b52e213e13dc778801352f82118f911ab7f

Observation 76c0dd0c-e7cd-4f32-8924-f033a51a921c · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.455133Z digest=sha256:0816efa7ef221cf2c8f488ad4850c6b1c80d4a659910f9c6bf21e607438989cf

Observation 30c4032d-9850-4d5d-b62e-ae3728234cc1 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesiz- ers,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesiz- ers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.780616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.461525Z digest=sha256:72ddbc688c09d7a8523c65de45ef2518795b865d648d8a492251101f6c721b0b

Observation b78bc22a-fcd7-4a2c-84c2-66b3954ae4d2 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.469169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.469169Z digest=sha256:bb04e457163f180b71463001767547f4fc83ccd21da70611857448143d2fac73

Observation f942b0ff-5a44-45d4-8b9d-f6c1fa025cb7 · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning V oicebox: Text-guided multilin- gual universal speech generation at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.760835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.479055Z digest=sha256:fa153b72afcfa82d4f95db3af4f124863eb5725b77ea5e5c4545d315199ed886

Observation bc11e927-d6f0-4d57-8504-d2201b043660 · outbound

This paper cites Listening while speaking: Speech chain by deep learning,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Listening while speaking: Speech chain by deep learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.741415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.515750Z digest=sha256:3bd07e0e55af5952f99144baf1f867223f4bc081000be5557536940a354c1f21

Observation 868b8ddd-ff18-4177-8dfd-146abb81dbe7 · outbound

This paper cites Almost unsupervised text to speech and automatic speech recognition,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Almost unsupervised text to speech and automatic speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.720787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.522695Z digest=sha256:6ce7049a088adc3b562da8601f421e65e3b89d8fd973c01ff8bc44974aa282b4

Observation f8a5af9e-90d8-4ad9-a796-5392a21170be · outbound

This paper cites Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.701489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.530970Z digest=sha256:bef1f4d9ba8bb5b73febf82635c6bbbc3a2632e291d039d1b3b9cff445ac7714

Observation ef53aded-23e9-4299-9d45-52066e46cc17 · outbound

This paper cites Investigation of enhanced tacotron text-to-speech synthesis systems with self- attention for pitch accent language,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Investigation of enhanced tacotron text-to-speech synthesis systems with self- attention for pitch accent language,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.678475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.536055Z digest=sha256:a64a8f1e1e21450a8acf20727a85a1bae0bcbe1547976029651303069e33e59c

Observation c08a13df-5cfb-450e-9d4b-c0447b39af09 · outbound

This paper cites A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.646962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.543348Z digest=sha256:757bfa69d06b73c6d99354413a48b7e47f05e4900211173e52177e8e6be0bb7f

Observation 9435527f-a9d1-4d16-9010-aa869e1c7848 · outbound

This paper cites A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.619564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.549588Z digest=sha256:efe26d4ad1a26c612ff7de410a9f21b496ee9817fdf49bcba54a84d3a680ad6a

Observation a50ed950-a929-4c7e-93a5-db966a4ba857 · outbound

This paper cites Polyphone disambigua- tion and accent prediction using pre-trained language models in Japanese TTS front-end,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Polyphone disambigua- tion and accent prediction using pre-trained language models in Japanese TTS front-end,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.599984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.555619Z digest=sha256:eceabe95a35753b602eedd1c26d5b07073f07ede9fdd62bef6d4d9ab1601af4a

Observation f3392d7b-db93-4477-98bf-158eef3a6723 · outbound

This paper cites Enhancing Japanese text-to-speech ac- curacy with a novel combination Transformer-BERT-based G2P: Integrating pronunciation dictionaries and accent sandhi,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Enhancing Japanese text-to-speech ac- curacy with a novel combination Transformer-BERT-based G2P: Integrating pronunciation dictionaries and accent sandhi,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.570769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.565465Z digest=sha256:1c04d364022f019cd54d6166411ef8a4c9b3200aba8f23fcdb3e3c8d399d4143

Observation f552c249-4b3b-47c8-8c9d-4d3c2c4c0196 · outbound

This paper cites Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.544361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.572435Z digest=sha256:59e3f06da9993ad82b5e8ce21236c3b914df21012a7395847bdc94e64091594a

Observation d1ed3e51-589b-422a-9cc4-5904314ad472 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Robust speech recognition via large-scale weak supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.517621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.577753Z digest=sha256:b2d33ac3f9c74ae9009f5944381cb13dd53e2a91bd8455503ab5ecc7ff535405

Observation c724ebc8-0b5b-4c22-b3c5-84e02602b416 · outbound

This paper cites Reazonspeech: A free and massive corpus for Japanese asr,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Reazonspeech: A free and massive corpus for Japanese asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.493668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.583381Z digest=sha256:e538227aab56126735466ae6f164a0a86bcf91c586dd532d43ca9fb4c1033163

Observation 6e54d56c-00de-446f-88a5-9df6dee64ba7 · outbound

This paper cites YODAS: Youtube-oriented dataset for audio and speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning YODAS: Youtube-oriented dataset for audio and speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.471599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.591574Z digest=sha256:52cd8653ce4c5df305a76443736149a21f9e4ae3b288681022b59f402517274d

Observation 3f41bf25-c650-4f98-9b8a-8920bdf1e30a · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recog- nition, translation, and language identification,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning OWSM-CTC: An open encoder-only speech foundation model for speech recog- nition, translation, and language identification,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.440054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.598323Z digest=sha256:516ec05d01e5ff8a0383caf93812a49ae37e4684e36c9ba3bd239a8b1a244657

Observation 19463acd-44b6-48c9-86c0-122a72c50bce · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.603784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.603784Z digest=sha256:e4ec22edaf32eabfbf8d450cd130b70534b6c4904b69fa040b1d1cb5d5cd84d1

Observation 5d8ac709-023a-4726-bbaa-09be72f8e728 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.616398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.616398Z digest=sha256:45d19a5950956cb91621bf7e7020990f4926df63e5bfcfd9c240e122e8055ab9

Observation 17825b3e-71ae-4e4a-a71a-3f083d27a2c4 · outbound

This paper cites PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.411517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.626660Z digest=sha256:ca508b5a9d441b1675fbe9259bfaaa0879222e8dd31940cd6705f014764215e2

Observation 4fb61dce-a47f-449a-bdf0-d4cc80b894ab · outbound

This paper cites Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme pre- dictions,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme pre- dictions,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.385283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.633625Z digest=sha256:afc224cf818821643624ca9874f0760bbc4945f59618836d046528fa371eb9cb

Observation 2ce861f4-1eea-4496-9484-b587b5dc1434 · outbound

This paper cites Miipher: A robust speech restoration model integrating self-supervised speech and text rep- resentations,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Miipher: A robust speech restoration model integrating self-supervised speech and text rep- resentations,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.366480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.648375Z digest=sha256:9c201992b777b1c2d42ea8b42e59ca95335879f3a3338914ce8714cd7a831ae4

Observation 31d5dd8b-a730-4d18-9463-76b826057f19 · outbound

This paper cites Japanese text-to-speech syn- thesis system: Open JTalk,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Japanese text-to-speech syn- thesis system: Open JTalk,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.346905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.656518Z digest=sha256:98ed8462f6fce72914c34a594931340ab0908adf152987fc3fe988b09819fe18

Observation 732fba98-843d-4fa3-a9c3-04c6d9d80334 · outbound

This paper cites Release of pre-trained mod- els for the Japanese language,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Release of pre-trained mod- els for the Japanese language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.327855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.665499Z digest=sha256:199b6b7cd62d84c3505470cbf0a1bf810c61b8cf113003a0339b4599eaffb8e9

Observation 583a47e6-ec4b-4d60-9037-fd4384f561c0 · outbound

This paper cites Joint ctc-attention based end- to-end speech recognition using multi-task learning,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Joint ctc-attention based end- to-end speech recognition using multi-task learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.304977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.671693Z digest=sha256:31bd198dab7dbe5bb34c43c9786c46960205a477f27072970fd12ce4dfbe548e

Observation 43022c1b-e346-4bc8-8cde-d01027fbcd08 · outbound

This paper cites Relaxing the conditional indepen- dence assumption of ctc-based asr by conditioning on interme- diate predictions,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Relaxing the conditional indepen- dence assumption of ctc-based asr by conditioning on interme- diate predictions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.285924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.678701Z digest=sha256:d5b99b613c33139571504f7a58040452d208fc183441b11ecf947beb0d5d7a71

Observation 47071f34-7d6e-493e-900b-270466e0f06d · outbound

This paper cites BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.256179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.686812Z digest=sha256:42c647d7fae326d3763ba0e5cb7ad1ca548ccfdf1d4df6f3851fc72e277cbf35

Observation fc735343-fa35-4e9a-a7d3-262905b81062 · outbound

This paper cites Improving speech recognition error prediction for modern and off-the-shelf speech recognizers,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Improving speech recognition error prediction for modern and off-the-shelf speech recognizers,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.236603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.692902Z digest=sha256:9f215b9026ab2290b4c04c9faa07e912efe345c8fbb2e415488e501fad2c2580

Observation b26f4c93-ca02-47c9-bb40-7bb3d732ce36 · outbound

This paper cites The theory of dynamic programming,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning The theory of dynamic programming,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.216220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.698387Z digest=sha256:6669997591492e444c44ce2c8f982a8eb48113f12ee672f62356177628b77e1a

Observation cbda27c5-89e2-45a0-8e39-e82525343f51 · outbound

This paper cites JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis re- search,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis re- search,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.195859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.704647Z digest=sha256:0dbf5ed7a098eb8601470e21c0e86330c06910e1b631108d6aa219b8e2fcf3a6

Observation 2059ae9e-9677-44c5-8b28-8d2c822ac4c3 · outbound

This paper cites Applying conditional random fields to Japanese morphological analysis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Applying conditional random fields to Japanese morphological analysis,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.176497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.711298Z digest=sha256:309500cd5a0791697d03207d0f3c731871d457e32bac340834bb8272cf8d56a1

Observation 5822a7ae-f875-4f4f-8bca-30cb0505e99b · outbound

This paper cites A proper approach to Japanese morphological analysis: Dictionary, model, and eval- uation.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A proper approach to Japanese morphological analysis: Dictionary, model, and eval- uation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.151737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.717972Z digest=sha256:6e2f532eb9baa92ccc8b435fcdff66de064f788b4a2d12c3367b019a11bf0edd

Observation eaa20038-d109-494e-84f1-59dd387d048a · outbound

This paper cites Period VITS: Varia- tional inference with explicit pitch modeling for end-to-end emo- tional speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Period VITS: Varia- tional inference with explicit pitch modeling for end-to-end emo- tional speech synthesis,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.129554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.723592Z digest=sha256:010e2e4e5f5323703e552ac909e6cf2357296a9c295cd18ff15b8bc5a93aa7e2

Observation dd86049e-a605-499f-986c-042b9a825cad · outbound

This paper cites DEMAND: A col- lection of multi-channel recordings of acoustic noise in diverse environments,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning DEMAND: A col- lection of multi-channel recordings of acoustic noise in diverse environments,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.111311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.731270Z digest=sha256:947448f88c68ec20496993d0ae322e8999a956e82bb9f47c562f693b01f19c44

Observation 1a35e2c5-68bb-42df-8bc1-73b3a4f86434 · outbound

This paper cites The ACE challenge—corpus description and performance evaluation,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning The ACE challenge—corpus description and performance evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.088377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.740563Z digest=sha256:bb165a7eac46ed225f46d382a7916ade0bd8f412da77dbd16d4627fc90ccc845

Observation 296c0a6e-4eca-4d0f-9891-b8969d28c888 · outbound

This paper cites End-to-end ASR to jointly predict transcriptions and linguistic annotations,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning End-to-end ASR to jointly predict transcriptions and linguistic annotations,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.069677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.750147Z digest=sha256:795be4011660bc8f6b1e404046924158a3576c7f2b3fbd51183fee5d69c8317a

Observation 2aaf2e9b-4a5f-415d-b22c-84b3264f98ac · outbound

This paper cites Building competitive direct acoustics-to-word models for english conversa- tional speech recognition,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Building competitive direct acoustics-to-word models for english conversa- tional speech recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.051902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.760108Z digest=sha256:2e53dff0c4378bc65aae7d8b68f540138fffcdb74ba8cd79240ea73199144eb2

Observation f0633d88-a512-411c-8e16-5b87fe41b9e0 · outbound

This paper cites Joint speech recognition and speaker diarization via sequence transduction,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Joint speech recognition and speaker diarization via sequence transduction,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.028606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.769922Z digest=sha256:99277987cd34a9112d2b142c22b23f25b23f6aab4226c78069f056eca7707b20

Observation 1f62b35d-1388-466c-87de-fbdaad28620a · outbound

This paper cites Uncon- strained many-to-many alignment for automatic pronunciation an- notation,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Uncon- strained many-to-many alignment for automatic pronunciation an- notation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.004563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.777135Z digest=sha256:58e1eb46130fbe716535c6f9b8cf85734a67feed75403ffb07526238bc3be88f

Observation 800a57a2-ed89-4474-9bb9-d7ba9a01ac0e · outbound

This paper cites Evaluation of many-to-many alignment algorithm by auto- matic pronunciation annotation using web text mining,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Evaluation of many-to-many alignment algorithm by auto- matic pronunciation annotation using web text mining,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:15.982204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.786786Z digest=sha256:4700c17be2d19b423d2e60494968ff20af4abd7a012706e98afce20453f2a407

Pith citing papers

Observation 3828ef2d-8103-4827-9723-0e4ff294770b · inbound

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning cites this paper.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:45:15.953103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:45:15.390168Z digest=sha256:69f1315c28943173107ca069de8a728e0e50f4e08d23785538f410f55cefa20c