Pith. sign in

Paper Citation Record · LEDGER

Voice Adaptation for Swiss German

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.22054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22054 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:12.222004Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:04.116674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:23:12.457477Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9360fa95-2080-4783-9ccf-5cf0dfdfd366 · outbound

This paper cites It is now possible to clone a voice across languages with less than a minute of audio required [3, 4].

Voice Adaptation for Swiss German It is now possible to clone a voice across languages with less than a minute of audio required [3, 4]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.935149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.038461Z digest=sha256:743104af256b2c40b0b36358c58114b21e965b1b40cd2ed7c1b1707c9133307f

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · outbound

This paper cites Voice Adaptation for Swiss German.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:48addb8dbd22ce0608793e4f862b8e65389032d1ea44191627a94ac6ef8b2771

Observation c3ddeacb-2b21-408f-8d4f-af3d3e24ce66 · outbound

This paper cites For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix.

Voice Adaptation for Swiss German For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.674564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.252180Z digest=sha256:a6a04a6a14afb1088dd4a1d6fa112c2cdf16f8ae112ef6fbacbd2b82ad6aa850

Observation 8cef58c1-db32-4a48-b1e9-610963262ba3 · outbound

This paper cites The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files.

Voice Adaptation for Swiss German The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.458662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.357760Z digest=sha256:eb2c8a03a8dc951c39790214f6327125f7ebce5da703fa17c9f987a751bbe413

Observation 6f335fa5-5f71-4d6a-899f-8b6aa9af0001 · outbound

This paper cites The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits.

Voice Adaptation for Swiss German The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.199244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.486066Z digest=sha256:cb705324b7f40670bec63cfd42c9142c2af95f47a3cf66301b4a11de19c8978a

Observation 0d4607ae-b7db-4300-820e-fe1f53b6300b · outbound

This paper cites We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results.

Voice Adaptation for Swiss German We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.666971Z digest=sha256:1df15be521bda4447bc9eec1acbebef27c83e7af5bd7612231987430edd8be49

Observation 07c2fa73-8827-4748-8982-68848deb21f1 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:04.803829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:04.803829Z digest=sha256:91823fb47165ff6763e3406b976d30b65d3bc236bbf07ddad233163b57ee28f3

Observation 78f0abd3-8987-4043-b332-5830629f06d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:05.343214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:05.343214Z digest=sha256:6dcbfe997160d48f56569bbc69d84162526678680c71adf61e8c5068135fd6b1

Observation f148466c-20d9-4340-a9ed-b9aedfd7a3f6 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Voice Adaptation for Swiss German Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:07.320866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:07.320866Z digest=sha256:6ef503e9c52bb93c19ba77b4884d058faeb9f9ac6ec14413a5630f1db41d1a47

Observation c72a5b82-fef3-45bd-bdd5-26bfe6611af9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,.

Voice Adaptation for Swiss German XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.733994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:07.455851Z digest=sha256:aa1cd28da5acd5f717a1348786090f23241990730f71f47579a82fd0a38943df

Observation b9c8fbaa-7db7-4946-8806-96a193e5e12e · outbound

This paper cites Atten- tion is all you need,.

Voice Adaptation for Swiss German Atten- tion is all you need,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.526781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:07.623190Z digest=sha256:1498428e551dae5bdc7eee89fa9bb33df2a52ad6a69bbda368624fb4e659eca2

Observation eee9640c-9dc1-4413-b0db-56dd9709b897 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,.

Voice Adaptation for Swiss German Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.233252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:08.110373Z digest=sha256:0f312d827d5898cace63d5aeb9a61d7770f4b67c520b430d66352299cbcad8b6

Observation 2ff7b0fe-c7bf-4b3f-a4ad-2c442e8cb43e · outbound

This paper cites Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,.

Voice Adaptation for Swiss German Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.048074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.048074Z digest=sha256:c4fd3dfb6cdc02f914eed59472985f59a64286cfaf7e69029a7c60902f81f772

Observation 4f9c09f7-e28a-4e74-951e-cec81e9bae97 · outbound

This paper cites Autoprep: An automatic preprocessing framework for in-the-wild speech data,.

Voice Adaptation for Swiss German Autoprep: An automatic preprocessing framework for in-the-wild speech data,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.060135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:10.189848Z digest=sha256:f6b26c04bba1417ed3ed47b9b5a2850d3d3b4406b6e054dfb0cb90ee6eed7a12

Observation 9610d4f2-2459-4f65-bc20-0db5b7d752f4 · outbound

This paper cites SDS-200: A Swiss German speech to Standard German text corpus,.

Voice Adaptation for Swiss German SDS-200: A Swiss German speech to Standard German text corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.853461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:10.299302Z digest=sha256:64b952bdf12f273eda488c97819dd13791446a5254971fe9a1d1b9abf9a9fd23

Observation ccab4ab0-f4ef-4840-b927-e429cad7b4ef · outbound

This paper cites STT4SG-350: A speech corpus for all Swiss German dialect regions,.

Voice Adaptation for Swiss German STT4SG-350: A speech corpus for all Swiss German dialect regions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.674152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:10.426587Z digest=sha256:dd71526c00e6a7024c6da7508e274ddefa6629c2e0c572e3d468bcf141aabc46

Observation c96d4746-dcf6-4885-9db6-e185de5cc42b · outbound

This paper cites Fine-tuning Whisper on Low-Resource Languages for Real-World Applications.

Voice Adaptation for Swiss German Fine-tuning Whisper on Low-Resource Languages for Real-World Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.577473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.577473Z digest=sha256:e4afc2c3946731db0679420c8e0c00e86053a082dc502e44accaa5aeda143cf0

Observation d36cf610-c113-49cd-ae69-a0f96157642f · outbound

This paper cites SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German.

Voice Adaptation for Swiss German SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.730467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.730467Z digest=sha256:e5cdb2593637921b6de61ca8cc43927e04c40f8f61e9dd177714f3d527860cfd

Observation 07d8835e-427e-4c65-a01e-8b43d7627762 · outbound

This paper cites Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,.

Voice Adaptation for Swiss German Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.426205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:10.900109Z digest=sha256:25f8aa3ac3dbad0f9eb793028e8cbf8b75378a94275d3b04292ca7229eb36c9e

Observation 6c2527e8-c2fa-417a-b316-56e2321ca07d · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,.

Voice Adaptation for Swiss German pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.127369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:11.089755Z digest=sha256:1336a4581922a14300f88eb5c6ca87f1278b51df536572bd40ed60e820074532

Observation 45c10899-69c6-4eb3-9ffc-da8cd68cb9ed · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

Voice Adaptation for Swiss German Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.763019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:11.230304Z digest=sha256:4f5678843342043dcb00c5bf1cbafe7a7d7fdc1081447a5a622f0862ce373275

Observation 383a0702-687b-4fad-a8b0-68d07dd8345f · outbound

This paper cites ELAN (Version 6.8) [Computer software],.

Voice Adaptation for Swiss German ELAN (Version 6.8) [Computer software],

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.405616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:11.428780Z digest=sha256:99ed969eed1c462db3d9f57bd8fe317f04fab2c7d3128f5918bff69543e4e8e1

Observation 3cacc949-1138-4e5d-bfb0-bf129a7e2be8 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Voice Adaptation for Swiss German Robust speech recognition via large-scale weak su- pervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.567094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.567094Z digest=sha256:c14f4e3d8c8d84d2fee56588d4fc1773c50bdfaa1f3f7f7f17d7be1e60b14231

Observation fa6ec5c3-4192-4638-a18f-9b3c9a5926cf · outbound

This paper cites Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,.

Voice Adaptation for Swiss German Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.150766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:11.681298Z digest=sha256:882dda729fd87e2cd5fc0fb94b02f3dffd16d16398e192decb8d84b9efb16397

Observation bd712c5f-effc-41a8-9975-fac9d42d2d4b · outbound

This paper cites Simple and Effective Zero-shot Cross-lingual Phoneme Recognition.

Voice Adaptation for Swiss German Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.843471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.843471Z digest=sha256:ada51fb38bf20dfce23c48ee7613cbd88c174aff332b3251a7d55c8eb7e4f62c

Observation 2125b9b7-a794-4ec2-9470-b4b64e7caa41 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Voice Adaptation for Swiss German Common voice: A massively-multilingual speech corpus,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.955633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.955633Z digest=sha256:f9ee7e5ce70e5a5d9e3d4680edb2178a8b5a1782bbd7b37007175448439bec05

Observation 7b48d7f8-d108-4065-92a3-d65c95915eeb · outbound

This paper cites Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,.

Voice Adaptation for Swiss German Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:12.102534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:12.102534Z digest=sha256:ab8647d9d1522d9b5f95e9b3caec19e700b7d471b6078c0f33e79890987895d1

Observation e8a41c36-d42a-4dff-8cf8-28f5895874fa · outbound

This paper cites Dialect transfer for Swiss German speech translation,.

Voice Adaptation for Swiss German Dialect transfer for Swiss German speech translation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:12.860467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:12.222004Z digest=sha256:4f10dec0edcbb9b6486f9d9c3907d941e6be690828d42bb3f98807a977bbf41b

Pith citing papers

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:48addb8dbd22ce0608793e4f862b8e65389032d1ea44191627a94ac6ef8b2771