Pith. sign in

Paper Citation Record · LEDGER

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2505.20341.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20341 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:48.254853Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:44.020833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:28:48.777510Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f217865-cea5-42ed-87f4-44b0d1568529 · outbound

This paper cites Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.852397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.020833Z digest=sha256:12d8916bf603c33a106c854d9e0475d4f5811ae3670deba55fc1301ce8afbc26

Observation 8eee5942-3a49-489d-a7e9-4e21042c05b8 · outbound

This paper cites sadness” as an ex- ample): “Now, I will give you a sentence. Please modify only one or two words to change the emotion to sadness. Please output only one modified sentence.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset sadness” as an ex- ample): “Now, I will give you a sentence. Please modify only one or two words to change the emotion to sadness. Please output only one modified sentence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.751761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.161073Z digest=sha256:be896394c51b54a6b248d8c20389f638c28d72d90cde07ef66b4e19fd185fc2d

Observation 43ed3c79-9812-44be-b80f-5831f9cd1cd9 · outbound

This paper cites The first two modules require pre- training, and then the third module is trained end-to-end.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset The first two modules require pre- training, and then the third module is trained end-to-end

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.248425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.341967Z digest=sha256:03eb8da451117f6ce2600ac25696acd8fe1d9f73da473cc277ebf919fee1172b

Observation e2ce15a0-d0e7-44e1-ac88-97f32f3890ae · outbound

This paper cites Experimental Setup We evaluate EmoCorrector on the ECD-TSE.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Experimental Setup We evaluate EmoCorrector on the ECD-TSE

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.011124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.448642Z digest=sha256:2d0647c7b64b3093d8419778d018866379722ea25549eae15411b5b18818e96f

Observation 135f474c-6533-4f97-aec9-9636001b2759 · outbound

This paper cites Experimental results demonstrate that the proposed framework improves emotional consistency in the edited speech.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Experimental results demonstrate that the proposed framework improves emotional consistency in the edited speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.764375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.578272Z digest=sha256:5300e01066128a7e16a8137f48d466c0b28897df85eb13f6bba4d042adb97b83

Observation 7ada997d-43c7-48da-b9e4-be65e03beab1 · outbound

This paper cites 62206136), the General Program (No.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset 62206136), the General Program (No

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.558362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.687416Z digest=sha256:94b05694b566fe866f18a0f64367b532e4afa711d341d6bc3b2ed22530f649a2

Observation 2b4da371-5ea5-4768-b7db-d6dab5a649ce · outbound

This paper cites Emotional voice con- version: Theory, databases and esd,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Emotional voice con- version: Theory, databases and esd,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:45.477922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:45.477922Z digest=sha256:9743127ae044a4ea585e39483dbec8b3bdae3b5370c5150068f75d5abfacb939

Observation 5bf0d3c4-5ff9-412c-b02f-b716b75d951d · outbound

This paper cites Fluentspeech: Stutter-oriented automatic speech editing with context-aware diffusion models,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Fluentspeech: Stutter-oriented automatic speech editing with context-aware diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.302583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.859707Z digest=sha256:ffcda6ecafc0198cb7474dcd97b2e9b67433e5237232133f847d46a47fb2c10c

Observation a957226c-a261-4f04-90b0-3968a953c34a · outbound

This paper cites A3t: Alignment-aware acoustic and text pretraining for speech synthe- sis and editing,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset A3t: Alignment-aware acoustic and text pretraining for speech synthe- sis and editing,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.069985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.980567Z digest=sha256:2e1a821e6ec4755a522f6c111e066e7bba0acd21ac0598a9b79ec188dc87d62d

Observation 0b8cb2be-9168-4c7a-87d8-d7fd5af52ee8 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:45.077048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:45.077048Z digest=sha256:7267c65a2f7991088afd040141277e3ff9e5941ccfaaf26edbc2cd514f2c659f

Observation 0a4f1427-3dea-4a21-bdb3-d8564011123e · outbound

This paper cites Speechx: Neural codec language model as a versatile speech transformer,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Speechx: Neural codec language model as a versatile speech transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.829602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.176163Z digest=sha256:4e9f2cf0a8a494d113c90bd82c8a30f39d4351553a5be1e6db1b2e81212b4446

Observation 96ad5d9e-5159-48cf-a12f-b918656bc5c2 · outbound

This paper cites V oice- craft: Zero-shot speech editing and text-to-speech in the wild,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset V oice- craft: Zero-shot speech editing and text-to-speech in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.574513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.291603Z digest=sha256:b2c477d0aa6dce8c43441abcbfc20d8cf7018d51c49bd3a23655fbf6644b926b

Observation 673ba17a-ef8f-4be3-af29-a216d3511379 · outbound

This paper cites Tackling modality heterogeneity with multi-view calibration net- work for multimodal sentiment detection,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Tackling modality heterogeneity with multi-view calibration net- work for multimodal sentiment detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.338641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.392572Z digest=sha256:fbfa007d683673dcdb41c198747c9e230fd5e5d1409184b80dcfc2b413c24b2a

Observation 31dd990b-f13a-4ad6-8b46-7f45de78aaaa · outbound

This paper cites MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.273638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.273638Z digest=sha256:d52c78ad2c9d54df60e917222a90459b95c574d8e3c5c0cdb888e366034b2cec

Observation 4c16ec55-e528-412b-b2bc-b0758292b0cf · outbound

This paper cites Decoupling speaker-independent emotions for voice conversion via source-filter networks,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Decoupling speaker-independent emotions for voice conversion via source-filter networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.135108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.573461Z digest=sha256:3b93f393b2ebc8f4c2f4af258c14be475a4e12728a6bfc5a914171dfe9085798

Observation 9a8e4392-99f9-488b-8d6f-500b10c82f14 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Retrieval- augmented generation for knowledge-intensive nlp tasks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.903067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.691045Z digest=sha256:f22a52efcfeab6f9c94f287b68cbe332a4cf417e13d29cb2aa1c3b07119a982a

Observation e2d3ddee-1e00-47c9-a852-ce3304452638 · outbound

This paper cites CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.621297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.816756Z digest=sha256:230c11d2788a52e024d644511d4eb56debccd3ebfbbc14e40c10dbed6d1a17ea

Observation 0e07d195-a889-4926-be63-00af7372900b · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.659972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:45.921557Z digest=sha256:532768085222d92ac308801cce1a5ed6eb4bccb34149e447f7065b26cd2dde8f

Observation 7ce033de-03a7-4c40-ad50-22df9f5efbfa · outbound

This paper cites Ultimately, the Azure system synthesizes speech for 5 speakers, CosyV oice2 synthe- sizes speech for another 5 speakers, and F5-TTS synthesizes speech for an additional 2 speakers.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Ultimately, the Azure system synthesizes speech for 5 speakers, CosyV oice2 synthe- sizes speech for another 5 speakers, and F5-TTS synthesizes speech for an additional 2 speakers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.460940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.265368Z digest=sha256:340faa81f2f7b5690300e6416732378acceeebdcd28c05c72a6be37fb821faeb

Observation 2642dd13-49f6-4254-a0bf-594957080dcc · outbound

This paper cites Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factoriza- tion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factoriza- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.328227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:46.042428Z digest=sha256:32d8b053b079bfd6aa0fbe7d67f2dfc3fa3699e6bb840e6f8508a9a8072b0d72

Observation 10093d43-8a68-498b-bdda-59b7196ec64b · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.073704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:46.173639Z digest=sha256:dd22f668bff080570ccb3957829a96377c46e942c34d26586b21cdd9ae7af2d9

Observation d76e586a-0303-4c6d-9aeb-3a27757aee85 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Iemocap: Interactive emotional dyadic motion capture database,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.386970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.386970Z digest=sha256:2cae857df427cd34df08f25503fe283e3fc283138134a0187f889d83b1d28d4d

Observation 68eb2da2-3e00-45b8-b83c-4b10dff8cee3 · outbound

This paper cites Azure speech studio,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Azure speech studio,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.841506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:46.491925Z digest=sha256:72d7e86ce26ea9389c0fbea6252bd441456186b541064cfe6c6811662c53410a

Observation 25852870-80ad-436a-ad82-6a8eac9a38a9 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.620609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.620609Z digest=sha256:9b072883ee38d4f3b7e13f24cc31e9ce1438bc4895eafc6f10a062fed95e095d

Observation c41608d8-a453-4986-893b-393c8db9e335 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.725703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.725703Z digest=sha256:b5403f11af38999a011087a71ad006b28bf7a755f57aa257e948e52bf2cc7a0c

Observation 4013cdb8-27d9-47ae-ae3c-7eec84735a0c · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.578718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:46.821401Z digest=sha256:6ef50b57b0bd8ff961769963e8059f2d1808ddbd74e62a36a5aa560dd5c14a3a

Observation 778845df-7422-4fb4-91fb-f946b3164c71 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Clap learning audio concepts from natural language supervision,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.937080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.937080Z digest=sha256:6f055f68e4ee42fd04df3f9c2d258834ea735ea129bd61fcdaba646485ccfab6

Observation 1d26394d-f402-4627-92fc-1e5aaf6f69b8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.071239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.071239Z digest=sha256:cb16529a7943b31345390d8188f0fa3c9768fcd6e93996f88c1bc453510c8d1c

Observation 6594f6f7-d31f-4e70-84cd-8276353c998b · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.256556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:47.187645Z digest=sha256:767ef1a5b1d5837551e29f7182b7a756bf1f1c9a3aa6dc7e509b748760f35d23

Observation e872d0b9-17ac-46a6-bac7-d619c8bda871 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.029173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:47.306630Z digest=sha256:34988edadeabb8572a8cea9c92434e100dc527f145aa9439a010492a26926d94

Observation b207e2eb-9bbc-49bf-b31f-a04100c052dc · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Montreal forced aligner: Trainable text-speech align- ment using kaldi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.452788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.452788Z digest=sha256:3c94583264e5c3216ea8ba2aaf8f94447df0a9cd3059d2b297e63a93977b1d45

Observation 9c9d3d00-fd43-4ea4-a043-9bed53fb2a6a · outbound

This paper cites Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.785963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:47.581107Z digest=sha256:09b016674da43ddf25a747be87a81af79fa39477b856d1e5e7760fab756c8179

Observation 70bae382-b6d5-4cdf-819f-89643bb7203d · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.720959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.720959Z digest=sha256:783207e25bbd15c5f3e0f42967cf1191c0350277711f525163e4de05194250d2

Observation 0f2aecc1-bdad-4bd6-96b2-0e5b86de5ee5 · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Qwen2-Audio Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.826489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.826489Z digest=sha256:6c49ccc3208ee81da9f33f7c70e2738d1b825046162eff848b4f7bee0e12aac2

Observation 443f867f-75c4-4789-91af-1d32735db711 · outbound

This paper cites Editspeech: A text based speech editing system using partial in- ference and bidirectional fusion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Editspeech: A text based speech editing system using partial in- ference and bidirectional fusion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.387912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:47.969837Z digest=sha256:c894beb1ecc869771cdcd62137c493ee8046a6988b9869a66b3732597be894e6

Observation 5b8b15e4-92e7-44f9-be3b-f24a4c2ec7f6 · outbound

This paper cites Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:48.097024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:48.097024Z digest=sha256:a69f2d55476c5e470c6c620926e2884c1d5e592ebb136593d564c1f370d38d13

Observation 8900faa3-d864-46b1-b66e-89201e0ed880 · outbound

This paper cites Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.073961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:48.254853Z digest=sha256:c4d7f73f8dc06c781d34bf55bc8b15098682bac2bce6fb2be63092498746b179

Pith citing papers

Observation 6f217865-cea5-42ed-87f4-44b0d1568529 · inbound

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset cites this paper.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.852397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:28:44.020833Z digest=sha256:12d8916bf603c33a106c854d9e0475d4f5811ae3670deba55fc1301ce8afbc26