Pith. sign in

Paper Citation Record · LEDGER

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2505.20341.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20341 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:48.254853Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:44.020833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:28:48.777510Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f217865-cea5-42ed-87f4-44b0d1568529 · outbound

This paper cites Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.852397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.020833Z digest=sha256:d22d553bf0469fab4996dacfebfd3a8138cb69c14b89583266f7516fc7377717

Observation 8eee5942-3a49-489d-a7e9-4e21042c05b8 · outbound

This paper cites sadness” as an ex- ample): “Now, I will give you a sentence. Please modify only one or two words to change the emotion to sadness. Please output only one modified sentence.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset sadness” as an ex- ample): “Now, I will give you a sentence. Please modify only one or two words to change the emotion to sadness. Please output only one modified sentence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.751761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.161073Z digest=sha256:f155742df5b915d5aff9c0397c16b732cde63d0aeef8c7532659a4d33a0e1fb6

Observation 43ed3c79-9812-44be-b80f-5831f9cd1cd9 · outbound

This paper cites The first two modules require pre- training, and then the third module is trained end-to-end.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset The first two modules require pre- training, and then the third module is trained end-to-end

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.248425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.341967Z digest=sha256:cd034070c90cc1fc0ddd6cdc24dc21d548552e1252a76986ed1f358768828b05

Observation e2ce15a0-d0e7-44e1-ac88-97f32f3890ae · outbound

This paper cites Experimental Setup We evaluate EmoCorrector on the ECD-TSE.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Experimental Setup We evaluate EmoCorrector on the ECD-TSE

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.011124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.448642Z digest=sha256:4ff2f2b830a52128bb547b797aaa0a1d254c66393b0c68317e5a18cb123f9b21

Observation 135f474c-6533-4f97-aec9-9636001b2759 · outbound

This paper cites Experimental results demonstrate that the proposed framework improves emotional consistency in the edited speech.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Experimental results demonstrate that the proposed framework improves emotional consistency in the edited speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.764375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.578272Z digest=sha256:71ccb961de62cf4c06c57bc5e2b83155bbad72f2e7b7dfa10cdb03af83c5d25b

Observation 7ada997d-43c7-48da-b9e4-be65e03beab1 · outbound

This paper cites 62206136), the General Program (No.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset 62206136), the General Program (No

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.558362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.687416Z digest=sha256:3c602b1443b438c8441be1d8cbc366511d2d79e46d8d8c0354e2acc1bc4d2646

Observation 2b4da371-5ea5-4768-b7db-d6dab5a649ce · outbound

This paper cites Emotional voice con- version: Theory, databases and esd,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Emotional voice con- version: Theory, databases and esd,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:45.477922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:45.477922Z digest=sha256:9743127ae044a4ea585e39483dbec8b3bdae3b5370c5150068f75d5abfacb939

Observation 5bf0d3c4-5ff9-412c-b02f-b716b75d951d · outbound

This paper cites Fluentspeech: Stutter-oriented automatic speech editing with context-aware diffusion models,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Fluentspeech: Stutter-oriented automatic speech editing with context-aware diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.302583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.859707Z digest=sha256:ca2a21102b74baf20602b29d55120107c4466fab3b95bb9d48dcfbec301c8a3d

Observation a957226c-a261-4f04-90b0-3968a953c34a · outbound

This paper cites A3t: Alignment-aware acoustic and text pretraining for speech synthe- sis and editing,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset A3t: Alignment-aware acoustic and text pretraining for speech synthe- sis and editing,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.069985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.980567Z digest=sha256:c463a0d83493745791ad1fddc8db00b6b0de6b78430c21a1f75cb169c87ca4cb

Observation 0b8cb2be-9168-4c7a-87d8-d7fd5af52ee8 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:45.077048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:45.077048Z digest=sha256:ea9f436afdd4e4c401e77c3e5064f4db4b83b74350e194308fdb86bd8827a9a0

Observation 0a4f1427-3dea-4a21-bdb3-d8564011123e · outbound

This paper cites Speechx: Neural codec language model as a versatile speech transformer,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Speechx: Neural codec language model as a versatile speech transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.829602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.176163Z digest=sha256:6a3ef79a57c74ca4edfbd583964a4cedd4ba015151b586f9444b52713755071d

Observation 96ad5d9e-5159-48cf-a12f-b918656bc5c2 · outbound

This paper cites V oice- craft: Zero-shot speech editing and text-to-speech in the wild,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset V oice- craft: Zero-shot speech editing and text-to-speech in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.574513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.291603Z digest=sha256:c2b9aacba46a34c5bf49990f9f3f0c1966b93bcb1dae2a0fbc6edd5266ebd5f5

Observation 673ba17a-ef8f-4be3-af29-a216d3511379 · outbound

This paper cites Tackling modality heterogeneity with multi-view calibration net- work for multimodal sentiment detection,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Tackling modality heterogeneity with multi-view calibration net- work for multimodal sentiment detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.338641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.392572Z digest=sha256:e565d8ca1e864b9490b74393f40485c4af8a1d05e6757ab102ba588e77b6d920

Observation 31dd990b-f13a-4ad6-8b46-7f45de78aaaa · outbound

This paper cites MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.273638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.273638Z digest=sha256:cd2be686a7e08254ffe43939855cd24ae082a62579c01001d4ce3484c8bb00bb

Observation 4c16ec55-e528-412b-b2bc-b0758292b0cf · outbound

This paper cites Decoupling speaker-independent emotions for voice conversion via source-filter networks,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Decoupling speaker-independent emotions for voice conversion via source-filter networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.135108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.573461Z digest=sha256:a6a72fbad6589606b252c7903e254b8c43c263609005b07f80e776a0c7d10845

Observation 9a8e4392-99f9-488b-8d6f-500b10c82f14 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Retrieval- augmented generation for knowledge-intensive nlp tasks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.903067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.691045Z digest=sha256:2f95d04e78b3022634ba71b7328507230d4dc2ab710dfe954a1ad9d34aef78ae

Observation e2d3ddee-1e00-47c9-a852-ce3304452638 · outbound

This paper cites CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.621297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.816756Z digest=sha256:1afe66919bb6f81b49080489f662741eb3e44f0e65a0fe1ed7f0bc4fd192ac72

Observation 0e07d195-a889-4926-be63-00af7372900b · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.659972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:45.921557Z digest=sha256:b3486126ef97b23f1af5c7392a282a4ac6de9c761966beb63c2074c1f745d33e

Observation 7ce033de-03a7-4c40-ad50-22df9f5efbfa · outbound

This paper cites Ultimately, the Azure system synthesizes speech for 5 speakers, CosyV oice2 synthe- sizes speech for another 5 speakers, and F5-TTS synthesizes speech for an additional 2 speakers.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Ultimately, the Azure system synthesizes speech for 5 speakers, CosyV oice2 synthe- sizes speech for another 5 speakers, and F5-TTS synthesizes speech for an additional 2 speakers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.460940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.265368Z digest=sha256:e766f5b65a4ac5b50fc92fe2464672183204bfc08f65defdfb7d7d4db0929b81

Observation 2642dd13-49f6-4254-a0bf-594957080dcc · outbound

This paper cites Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factoriza- tion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factoriza- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.328227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:46.042428Z digest=sha256:b243a70a289a96b564f49e9709dc950146b404262d5bb5305f23600ae67ea462

Observation 10093d43-8a68-498b-bdda-59b7196ec64b · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.073704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:46.173639Z digest=sha256:5929a64c8d0dcc2ce7472f8dc36ec2951676c8f5dc6ef5df73fa972e409b05c1

Observation d76e586a-0303-4c6d-9aeb-3a27757aee85 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Iemocap: Interactive emotional dyadic motion capture database,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.386970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.386970Z digest=sha256:2cae857df427cd34df08f25503fe283e3fc283138134a0187f889d83b1d28d4d

Observation 68eb2da2-3e00-45b8-b83c-4b10dff8cee3 · outbound

This paper cites Azure speech studio,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Azure speech studio,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.841506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:46.491925Z digest=sha256:4869d5114e0cead76754d41f3972dd2c514d32f7dacd8d0b671bf612a14b5a5a

Observation 25852870-80ad-436a-ad82-6a8eac9a38a9 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.620609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.620609Z digest=sha256:9b072883ee38d4f3b7e13f24cc31e9ce1438bc4895eafc6f10a062fed95e095d

Observation c41608d8-a453-4986-893b-393c8db9e335 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.725703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.725703Z digest=sha256:b5403f11af38999a011087a71ad006b28bf7a755f57aa257e948e52bf2cc7a0c

Observation 4013cdb8-27d9-47ae-ae3c-7eec84735a0c · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.578718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:46.821401Z digest=sha256:a4e37e1fdd10951b3a881f47eacece4da45eb1831ca7cf16e2ce5f74db64e27b

Observation 778845df-7422-4fb4-91fb-f946b3164c71 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Clap learning audio concepts from natural language supervision,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.937080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.937080Z digest=sha256:6f055f68e4ee42fd04df3f9c2d258834ea735ea129bd61fcdaba646485ccfab6

Observation 1d26394d-f402-4627-92fc-1e5aaf6f69b8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.071239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.071239Z digest=sha256:cb16529a7943b31345390d8188f0fa3c9768fcd6e93996f88c1bc453510c8d1c

Observation 6594f6f7-d31f-4e70-84cd-8276353c998b · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.256556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:47.187645Z digest=sha256:7797a75a0f346201e650aed8f177a4cb8902df2781c8c9ec31b1a83a30c5fd5e

Observation e872d0b9-17ac-46a6-bac7-d619c8bda871 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.029173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:47.306630Z digest=sha256:3ea06ae3ef5e8f3f1b6fde615a28c8dc5a9bb02674e3d81c5db96903028d8568

Observation b207e2eb-9bbc-49bf-b31f-a04100c052dc · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Montreal forced aligner: Trainable text-speech align- ment using kaldi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.452788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.452788Z digest=sha256:3c94583264e5c3216ea8ba2aaf8f94447df0a9cd3059d2b297e63a93977b1d45

Observation 9c9d3d00-fd43-4ea4-a043-9bed53fb2a6a · outbound

This paper cites Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.785963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:47.581107Z digest=sha256:f9d008b550e7b4b3dca6baefa70b01bcd08ca0ab22d0e31e33f6c237886be5c8

Observation 70bae382-b6d5-4cdf-819f-89643bb7203d · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.720959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.720959Z digest=sha256:783207e25bbd15c5f3e0f42967cf1191c0350277711f525163e4de05194250d2

Observation 0f2aecc1-bdad-4bd6-96b2-0e5b86de5ee5 · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Qwen2-Audio Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.826489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.826489Z digest=sha256:80a9b225682be34b99d7fb9f1813dfd3971cefa4cf55c8b4d8607251d62b1989

Observation 443f867f-75c4-4789-91af-1d32735db711 · outbound

This paper cites Editspeech: A text based speech editing system using partial in- ference and bidirectional fusion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Editspeech: A text based speech editing system using partial in- ference and bidirectional fusion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.387912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:47.969837Z digest=sha256:06767ea041a232156db96575b2e64d037905aeb2ed97542db10d4569ff8b92c2

Observation 5b8b15e4-92e7-44f9-be3b-f24a4c2ec7f6 · outbound

This paper cites Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:48.097024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:48.097024Z digest=sha256:a69f2d55476c5e470c6c620926e2884c1d5e592ebb136593d564c1f370d38d13

Observation 8900faa3-d864-46b1-b66e-89201e0ed880 · outbound

This paper cites Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.073961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:48.254853Z digest=sha256:d8fc9e666f6e9f269b3863873aeab5dc9252562bec6f57a911203dfaa1790a96

Pith citing papers

Observation 6f217865-cea5-42ed-87f4-44b0d1568529 · inbound

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset cites this paper.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.852397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:28:44.020833Z digest=sha256:d22d553bf0469fab4996dacfebfd3a8138cb69c14b89583266f7516fc7377717