Pith. sign in

REVIEW 3 major objections 5 minor 84 references

A Survey on Spoken Italian Datasets and Corpora

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey catalogs 66 spoken Italian datasets, organizes them by speech type, source, and demographic and linguistic features, and publishes the complete inventory openly, while identifying persistent gaps in dialect, child, and…

desk verdict Useful inventory for Italian speech researchers, but the 'comprehensive' claim can't be audited without the full list and search protocol in the paper itself. read the letter →

arxiv 2501.06557 v2 pith:2OHSQ2YD submitted 2025-01-11 cs.CL cs.AIcs.DL

classification cs.CLcs.AIcs.DL
keywords spokenItaliandatasetsspeechcorporalanguageresourcesautomaticrecognitiondatasetsurveydialectalvariationtechnologyannotationstandards
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that the landscape of spoken Italian data is broader than commonly assumed, yet still falls short of what Italian's dialectal and demographic diversity requires. It identifies, examines, and catalogs 66 spoken Italian datasets, organizes them by speech type, source, and linguistic and sociodemographic features, and publishes the complete inventory openly on GitHub and Zenodo. It also documents recurring gaps, including underrepresentation of dialects and minority languages, demographic imbalance, restricted access, and inconsistent annotation standards, and offers recommendations for future dataset creation and sharing. A sympathetic reader would care because anyone building Italian speech systems or doing Italian linguistics needs a reliable map of what data exists and where it is thin.

What carries the argument

The central artifact is the curated inventory of 66 datasets, organized along three axes: speech type, source and context, and demographic and linguistic features. The inventory, publicly available on GitHub and archived on Zenodo, is what carries the survey's claims, because the gap analysis, the application tables, and the recommendations all derive from the way these datasets are sorted and compared. The paper also uses a set of named tool references, such as Praat, ELAN, and WebMAUS, to illustrate annotation practice, but those serve as examples of methodology rather than as carriers of the central claim.

What would settle it

Check the public Zenodo archive of the inventory and look for an actively maintained spoken Italian dataset, documented as of November 2024, that is absent from the 66; finding a significant missing dataset would falsify the comprehensiveness claim. A second check is to verify a random sample of the 66 entries against their cited source pages, and if many sizes, speech types, or availability statuses are wrong, the categorization would collapse.

Watch

Extended reading notes

Core claim

The paper's central claim is that it provides a comprehensive examination of 66 spoken Italian datasets, characterized by speech type, source and context, and demographic and linguistic features, with the full inventory publicly available via GitHub and archived on Zenodo. It finds that the existing resources concentrate on read and conversational standard Italian, leaving dialects, minority languages, children's speech, field recordings, and specialized domains comparatively underrepresented. It also identifies technical and ethical challenges, such as inconsistent annotation, large-file overhead, privacy concerns, and access restrictions, and proposes standardized annotation, open-access models, and collaborative collection as the main remedies.

Load-bearing premise

The survey assumes that the 66 datasets it assembled, with sizes and features taken from creators' own documentation, accurately represent the full landscape of usable spoken Italian resources as of November 2024, and that this documentation is reliable.

Editorial extensions

If this is right

  • Researchers can use the public inventory to locate an Italian speech dataset by speech type, source, or demographic need, instead of rediscovering resources through scattered catalogs.
  • The gap analysis gives concrete targets for new data collection: dialects and minority languages, child speech, outdoor field recordings, and specialized domains such as medical or non-native speech.
  • Adopting the paper's recommendations for standardized annotation formats, open-access licensing, and richer metadata would make Italian speech datasets easier to combine and compare.
  • The survey's reliance on creator documentation means its characterizations are a starting point for selection, not an independent quality certification of the datasets themselves.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The three-axis categorization is simple enough to serve as a lightweight template for comparable surveys of other under-resourced languages, not just Italian.
  • Because the inventory is frozen as of November 2024 and hosted on a public repository, it is positioned to become a living resource; community updates would be a natural next step that the paper only gestures at.
  • The paper highlights a few dozen datasets in its tables and discussion while the full inventory contains 66 entries, so a reader who wants the complete map must consult the GitHub or Zenodo archive rather than rely only on the body text.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a survey of 66 spoken Italian datasets and corpora, organizing them by speech type, source/context, and demographic/linguistic features, and discussing collection and annotation methodologies, applications, challenges, and future directions. The authors state that the complete inventory is available on GitHub and archived on Zenodo with DOI 10.5281/zenodo.14246196, while the article itself presents selected examples in three tables. The survey's stated aim is to provide a comprehensive overview of usable spoken Italian resources as of November 2024, along with a gap analysis and recommendations.

Significance. If the inventory is accurate, this survey fills a genuine gap: there is no comparable recent reference cataloging spoken Italian resources across academic, commercial, and crowdsourced sources. The authors' decision to publish the full inventory as an open, versioned artifact on GitHub/Zenodo is a concrete contribution that can be updated and reused, and the categorization by speech type, source, and demographic/linguistic features is useful for navigating the resource landscape. The paper is also transparent in stating that dataset quality and availability were not independently verified. However, the central 'comprehensive' claim and the resulting gap analysis in Sections V and VI depend on the completeness and correct attribution of the 66-item inventory, and this dependence is currently the paper's weakest point because the selection protocol is unspecified and the full list is not in the manuscript.

major comments (3)
  1. [I.B, V, VII] The claim of comprehensiveness over 66 datasets is not auditable from the manuscript. Section I.B does not state the search queries, databases queried, inclusion/exclusion criteria, or a working definition of 'actively maintained,' and Section VII directs readers to an external GitHub/Zenodo archive for the full list. Within the paper, Table 1 shows 23 rows, Table 2 shows 23 datasets, and Table 3 shows 9 datasets, so the reader cannot check membership, duplicates, or overlap handling from the article itself. Because Section V's scarcity/accessibility analysis and Section VI's recommendations are derived from what is absent from this list, an incomplete or misattributed inventory would change the conclusions, not just the numbers. The authors should provide the full 66-item inventory as a supplementary appendix or, at minimum, an appendix table with the same fields as Table 1, and they should document the search and inclusion protocol.
  2. [Table 1, II.A.2] Several table entries appear to report multilingual totals rather than Italian-specific figures, which is inconsistent with the survey's stated scope. The VoxPopuli row lists '400,000 hours' and labels the source as 'Standard European languages,' but the paper's subject is spoken Italian datasets; no Italian-specific hour count is given in the table or in Section II.A.2. Similarly, the C-ORAL-ROM row reports '1,200,000 words' for 'Spontaneous speech in Romance languages,' which is a multilingual total. If Italian-specific quantities are unavailable or not reported, the table should state 'not reported' rather than a multilingual total that can be misread as an Italian dataset size.
  3. [V.A, I.B] The gap analysis in Section V.A draws conclusions about restricted access and commercial availability (e.g., 'commercial datasets such as Aurora Project [28] and Defined.ai [3], [7], [25] require significant financial investment') from data that the authors explicitly did not independently verify in Section I.B. While the authors' caveat is honest, it does not resolve the inconsistency between the caveat and the strength of the claims in Section V.A. The manuscript should either verify accessibility at the time of writing or mark each entry's availability as reported rather than confirmed, and it should define 'actively maintained' as used in Section I.B.
minor comments (5)
  1. [I.C] Section I.C refers to 'Section 2,' 'Section 3,' etc., while the actual section headings use Roman numerals (II, III, etc.); the numbering should be made consistent.
  2. [Table 3] The MuST-C row in Table 3 lists 'Creative Commons (temporarily suspended)' as the availability, which is an odd status for a table entitled 'Summary of Publicly Available Italian Speech Datasets'; the table should clarify what 'temporarily suspended' means and whether the dataset is currently downloadable.
  3. [II.C.1] In Section II.C.1, C-ORAL-ROM is described as 'though multilingual, includes a substantial component of Italian speech,' but neither the text nor Table 1 quantifies the Italian component; adding the Italian-specific size would strengthen the categorization.
  4. [IV.A.1] Reference [56], cited for EMOVO, is incomplete and lacks a publisher, venue, or URL; it should be completed for readers wishing to locate the resource.
  5. [III.C] The sentence 'according to our current knowledge, none of the documented datasets made use of this tool in validation processes' reads as an aside and would be clearer as a footnote or a statement of the limitation of the survey's information, since the authors already disclaim independent verification.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey derives no fitted quantities, makes no predictions from its own inputs, and rests on external dataset documentation rather than on a self-citation chain.

full rationale

This paper is a cataloging survey, not a derivation. It contains no equations, no fitted parameters, and no prediction that is later compared with the data that defined it. The central claim, that the paper provides a comprehensive overview of 66 spoken Italian datasets, is supported by citations to external dataset pages, repositories, and published corpus descriptions, not by the authors' own prior results. The only self-referential element is the statement in Section I.B and the Conclusion that the complete inventory is hosted on GitHub and archived on Zenodo (DOI: 10.5281/zenodo.14246196). That is an availability claim, not a circular step: the paper does not derive its categorization or gap analysis from the contents of that archive, nor does it treat the archive as evidence for its completeness. The manuscript explicitly discloses in Section I.B that 'dataset quality and availability were not independently verified' and that the survey 'synthesizes information provided by creators or documentation,' which correctly identifies that the inventory is compiled from external sources. Concerns about whether the 66-item inventory is auditable, whether inclusion criteria are stated, and whether some table entries attribute multilingual totals to Italian resources are legitimate reproducibility and verifiability risks, but they are not circularity: the conclusions do not reduce by construction to a fitted input or to an unverified self-citation. No load-bearing self-citations, uniqueness theorems, or ansatz-smuggling citations appear in the text. Accordingly, no circular step can be quoted or exhibited, and the honest finding is that the paper exhibits no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities are present because this is a survey. Two domain assumptions are load-bearing: trust in dataset creators' documentation, and completeness of the 66-dataset set.

assumptions (2)
  • domain assumption Dataset documentation supplied by creators is accurate
    The survey states in Section I.B that it 'synthesizes information provided by creators or documentation' and that quality and availability were not independently verified. All dataset sizes and features in Tables 1-3 inherit this assumption.
  • domain assumption The set of 66 datasets identified by the authors is comprehensive
    The paper does not describe a systematic search protocol or inclusion criteria beyond 'actively maintained as of November 2024 with sufficient available information' (Section I.B). The claim of comprehensiveness is therefore an unargued premise.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Spoken Italian Datasets and Corpora." pith.science (2026). https://pith.science/paper/2OHSQ2YD

@misc{pith2026250106557,
  author       = {Pith},
  title        = {Pith review of: A Survey on Spoken Italian Datasets and Corpora},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2OHSQ2YD}},
  note         = {Machine review of arXiv:2501.06557}
}
read the original abstract

Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored compared to major languages like English or Mandarin. This survey provides a comprehensive analysis of 66 spoken Italian datasets, highlighting their characteristics, methodologies, and applications. The datasets are categorized by speech type, source and context, and demographic and linguistic features, with a focus on their utility in fields such as Automatic Speech Recognition, emotion detection, and education. Challenges related to dataset scarcity, representativeness, and accessibility are discussed alongside recommendations for enhancing dataset creation and utilization. The full dataset inventory is publicly accessible via GitHub and archived on Zenodo, serving as a valuable resource for researchers and developers. By addressing current gaps and proposing future directions, this work aims to support the advancement of Italian speech technologies and linguistic research.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

84 extracted references · 75 canonical work pages

  1. [28]

    AURORA Project database - Subset of SpeechDat-Car - Italian database - Evaluation Package – ELRA Catalogue

    “AURORA Project database - Subset of SpeechDat-Car - Italian database - Evaluation Package – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-AURORA-CD0003_05/

  2. [3]

    Italian Spontaneous Dialogue Dataset | Defined.ai

    “Italian Spontaneous Dialogue Dataset | Defined.ai.” [Online]. Available: https://www.defined.ai/datasets/italian-spontaneous-dialogue

  3. [7]

    Italian Scripted Monologue Dataset | Defined.ai

    “Italian Scripted Monologue Dataset | Defined.ai.” [Online]. Available: https://www.defined.ai/datasets/italian-scripted-monologue

  4. [25]

    Italian Lifestyle Podcast Dataset | Defined.ai

    “Italian Lifestyle Podcast Dataset | Defined.ai.” [Online]. Available: https://www.defined.ai/datasets/italian-lifestyle-podcast

  5. [1]

    Corpus KIParla – L’italiano parlato e chi parla italiano

    “Corpus KIParla – L’italiano parlato e chi parla italiano.” [Online]. Available: https://kiparla.it/

  6. [2]

    KIParla Corpus: A New Resource for Spoken Italian,

    C. Mauri, S. Ballarè, E. Goria, M. Cerruti, and F. Suriano, “KIParla Corpus: A New Resource for Spoken Italian,” Proceedings of the Sixth Italian Conference on Computational Linguistics, vol. 2481

  7. [4]

    DIA-Dialogic ItAlian corpus,

    A. Vietti and D. Mereu, “DIA-Dialogic ItAlian corpus,” 2021, publisher: The Language Archive - Max Planck Institute for Psycholinguistics. [Online]. Available: https://bia.unibz.it/esploro/outputs/ dataset/DIA-Dialogic-ItAlian-corpus/991006608798401241

  8. [5]

    Robust Speech Recognition via Large-Scale Weak Supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” in Proceedings of the 40th International Conference on Machine Learning. 14 VOLUME XX, 202X PMLR, Jul. 2023, pp. 28 492–28 518, iSSN: 2640-3498. [Online]. Available: https://proceedings.mlr.press/v202/radford23a.html

Show all 84 references
  1. [6]

    WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,” Jun. 2022, arXiv:2110....

  2. [8]

    voxpopuli,

    “voxpopuli,” Nov. 2024, original-date: 2021-01-08T18:47:36Z. [Online]. Available: https://github.com/facebookresearch/voxpopuli

  3. [9]

    V oxPopuli: A Large- Scale Multilingual Speech Corpus for Representation Learning, Semi- Supervised Learning and Interpretation,

    C. Wang, M. Rivière, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “V oxPopuli: A Large- Scale Multilingual Speech Corpus for Representation Learning, Semi- Supervised Learning and Interpretation,” Jul. 2021, arXiv:2101.00390. [Online]. Availabl...

  4. [10]

    Spontaneous Speech Characterization and Detection in Large Audio Database,

    R. Dufour, V . Jousse, Y . Estève, F. Béchet, and G. Linarès, “Spontaneous Speech Characterization and Detection in Large Audio Database,” in 13-th International Conference on Speech and Computer (SPECOM 2009), St Petersburg, Russia, 2009. [Online]. Available: https://hal.scie...

  5. [11]

    A comparison of ASR and human errors for transcription of non- native spontaneous speech,

    M. Mulholland, M. Lopez, K. Evanini, A. Loukina, and Y . Qian, “A comparison of ASR and human errors for transcription of non- native spontaneous speech,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 5855–5859, iSSN:...

  6. [12]

    The LABLITA Speech Resources,

    E. CRESTI, L. GREGORI, M. MONEGLIA, C. NICOLÁS, and A. PANUNZI, “The LABLITA Speech Resources,” Corpora e Studi Linguistici, no. 6, pp. 85–108, 2022. [Online]. Available: https: //doi.org/10.17469/O2106SLI000005

  7. [13]

    Lablita

    “Lablita.” [Online]. Available: http://corpus.lablita.it/?locale=en

  8. [14]

    Multi-media edition; tools of analysis; standard linguistic mea- surements for validation in HLT – ELRA Catalogue.” [Online]

    “C-ORAL-ROM - Integrated reference corpora for spoken romance languages. Multi-media edition; tools of analysis; standard linguistic mea- surements for validation in HLT – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ELRA-S0172/

  9. [15]

    Hey ASR System! Why Aren’t You More Inclusive? Automatic Speech Recognition Systems’ Bias and Proposed Bias Mitigation Techniques. A Literature Review,

    M. K. Ngueajio and G. Washington, “Hey ASR System! Why Aren’t You More Inclusive? Automatic Speech Recognition Systems’ Bias and Proposed Bias Mitigation Techniques. A Literature Review,” Nov. 2022, arXiv:2211.09511. [Online]. Available: http://arxiv.org/abs/2211.09511

  10. [16]

    Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,

    E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,” Feb. 2023, arXiv:2302.03540. [Online]. Available: http://arxiv.org/abs/2302.03540

  11. [17]

    MLS: A Large-Scale Multilingual Dataset for Speech Research,

    V . Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” in Interspeech 2020, Oct. 2020, pp. 2757–2761, arXiv:2012.03411 [cs, eess]. [Online]. Available: http://arxiv.org/abs/2012.03411

  12. [18]

    MLS - Multilingual Librispeech

    “MLS - Multilingual Librispeech.” [Online]. Available: https://www. openslr.org/94/

  13. [19]

    Common V oice Mozilla

    “Common V oice Mozilla.” [Online]. Available: https://commonvoice. mozilla.org/

  14. [20]

    Combining Deep Learning with Domain Adaptation and Filtering Techniques for Speech Recognition in Noisy Environments,

    E. De J. Velásquez-Martínez, A. Becerra-Sánchez, J. I. De La Rosa- Vargas, E. González-Ramírez, A. Rodarte-Rodríguez, G. Zepeda-Valles, N. I. Escalante-García, and J. E. Olvera-González, “Combining Deep Learning with Domain Adaptation and Filtering Techniques for Speech Recogn...

  15. [21]

    Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech,

    T. Menne, I. Sklyar, R. Schlüter, and H. Ney, “Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech,” Sep. 2019, arXiv:1905.03500. [Online]. Available: http://arxiv.org/abs/1905.03500

  16. [22]

    Identifying Fake News on Social Networks Based on Natural Language Processing: Trends and Challenges,

    N. R. De Oliveira, P. S. Pisa, M. A. Lopez, D. S. V . De Medeiros, and D. M. F. Mattos, “Identifying Fake News on Social Networks Based on Natural Language Processing: Trends and Challenges,” Information, vol. 12, no. 1, p. 38, Jan. 2021. [Online]. Available: https://www.mdpi....

  17. [23]

    The MGB challenge: Evaluating multi-genre broadcast media recognition,

    P. Bell, M. J. F. Gales, T. Hain, J. Kilgour, P. Lanchantin, X. Liu, A. McParland, S. Renals, O. Saz, M. Wester, and P. C. Woodland, “The MGB challenge: Evaluating multi-genre broadcast media recognition,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding...

  18. [24]

    IBNC - An Italian Broadcast News Corpus – ELRA Catalogue

    “IBNC - An Italian Broadcast News Corpus – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-S0093/

  19. [26]

    Exploring AI-Driven Customer Service: Evolution, Archi- tectures, Opportunities, Challenges and Future Directions,

    S. M. I. , “Exploring AI-Driven Customer Service: Evolution, Archi- tectures, Opportunities, Challenges and Future Directions,” International Journal For Multidisciplinary Research, vol. 6, no. 3, p. 22283, Jun. 2024. [Online]. Available: https://www.ijfmr.com/research-paper.p...

  20. [27]

    Italian(Italy) Spontaneous Dialogue Telephony speech dataset - Nexdata

    “Italian(Italy) Spontaneous Dialogue Telephony speech dataset - Nexdata.” [Online]. Available: https://m.nexdata.ai/datasets/speechrecog/ 1232?source=Huggingface

  21. [29]

    MuST-C: a Multilingual Speech Translation Corpus,

    M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, V olume 1 ...

  22. [30]

    EmoFilm - A multilingual emotional speech corpus,

    E. Parada-Cabaleiro, G. Costantini, A. Batliner, A. Baird, and B. Schuller, “EmoFilm - A multilingual emotional speech corpus,” Sep. 2018. [Online]. Available: https://zenodo.org/records/7665999

  23. [31]

    Vivaldi

    “Vivaldi.” [Online]. Available: https://www2.hu-berlin.de/vivaldi/index. php

  24. [32]

    DEMoS: an Italian emotional speech corpus,

    E. Parada-Cabaleiro, G. Costantini, A. Batliner, M. Schmitt, and B. W. Schuller, “DEMoS: an Italian emotional speech corpus,” Language Resources and Evaluation, vol. 54, no. 2, pp. 341–383, Jun. 2020. [Online]. Available: https://doi.org/10.1007/s10579-019-09450-y

  25. [33]

    DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception,

    E. Parada-Cabaleiro, G. Costantini, A. Batliner, M. Schmitt, and B. Schuller, “DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception,” Feb. 2019. [Online]. Available: https://zenodo.org/records/2544829

  26. [34]

    ITALIC: An Italian Intent Classification Dataset,

    A. Koudounas, M. L. Quatra, L. Vaiani, L. Colomba, G. Attanasio, E. Pastor, L. Cagliero, and E. Baralis, “ITALIC: An Italian Intent Classification Dataset,” Jun. 2023, arXiv:2306.08502. [Online]. Available: http://arxiv.org/abs/2306.08502

  27. [35]

    ITALIC: An Italian Intent Classification Dataset,

    A. Koudounas, M. La Quatra, L. Vaiani, L. Colomba, G. Attanasio, E. Pastor, L. Cagliero, and E. Baralis, “ITALIC: An Italian Intent Classification Dataset,” Jun. 2023. [Online]. Available: https://zenodo.org/ records/8040649

  28. [36]

    Europarl ST corpus

    “Europarl ST corpus.” [Online]. Available: https://mllp.upv.es/europarl-st/

  29. [37]

    Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates,

    J. Iranzo-Sánchez, J. A. Silvestre-Cerdà, J. Jorge, N. Roselló, A. Giménez, A. Sanchis, J. Civera, and A. Juan, “Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Si...

  30. [38]

    PortMedia French and Italian corpus – ELRA Catalogue

    “PortMedia French and Italian corpus – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-S0371/

  31. [39]

    m-ailabs-dataset,

    i. celeste [witchzard, “m-ailabs-dataset,” Nov. 2024, original-date: 2019- 03-21T07:46:35Z. [Online]. Available: https://github.com/imdatceleste/ m-ailabs-dataset

  32. [40]

    MULTEXT Prosodic database – ELRA Catalogue

    “MULTEXT Prosodic database – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ELRA-S0060/

  33. [41]

    APASCI – ELRA Catalogue

    “APASCI – ELRA Catalogue.” [Online]. Available: https://catalogue.elra. info/en-us/repository/browse/ELRA-S0039/

  34. [42]

    Italian Kids Speech Recognition Corpus (Desktop) – ELRA Catalogue

    “Italian Kids Speech Recognition Corpus (Desktop) – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-S0228_98/

  35. [43]

    VOLIP: a corpus of spoken Italian and a virtuous example of reuse of linguistic resources

    I. Alfano, F. Cutugno, A. D. Rosa, C. Iacobini, and R. Savy, “VOLIP: a corpus of spoken Italian and a virtuous example of reuse of linguistic resources.”

  36. [44]

    Methods and Tools for Prosodic Analysis of a Spoken Italian Corpus

    M. Savino, M. Refice, and D. Daleno, “Methods and Tools for Prosodic Analysis of a Spoken Italian Corpus.” VOLUME XX, 202X 15

  37. [45]

    More Harm than Good: Why Dictionaries Using Orthographic Transcription Instead of the IPA Should Be Handled with Care,

    A. Bryła-Cruz, “More Harm than Good: Why Dictionaries Using Orthographic Transcription Instead of the IPA Should Be Handled with Care,” Research in Language, vol. 20, no. 2, pp. 133–152, Dec. 2022, number: 2. [Online]. Available: https://czasopisma.uni.lodz.pl/research/ articl...

  38. [46]

    RECOApy: Data Recording, Pre-Processing and Phonetic Transcription for End-to-End Speech-Based Applications,

    A. Stan, “RECOApy: Data Recording, Pre-Processing and Phonetic Transcription for End-to-End Speech-Based Applications,” Interspeech 2020, pp. 586–590, Oct. 2020, conference Name: Interspeech 2020 Publisher: ISCA. [Online]. Available: https://www.isca-archive.org/ interspeech_2...

  39. [47]

    Praat: doing Phonetics by Computer

    “Praat: doing Phonetics by Computer.” [Online]. Available: https: //www.fon.hum.uva.nl/praat/

  40. [48]

    ELAN Annotator | The Language Archive

    “ELAN Annotator | The Language Archive.” [Online]. Available: https://archive.mpi.nl/tla/elan

  41. [49]

    WebMAUS | Bavarian Archive for Speech Signals

    “WebMAUS | Bavarian Archive for Speech Signals.” [Online]. Available: https://clarin.phonetik.uni-muenchen.de/BASWebServices/ interface/WebMAUSGeneral

  42. [50]

    EMU-SDMS: Advanced speech database management and analysis in R,

    R. Winkelmann, J. Harrington, and K. Jänsch, “EMU-SDMS: Advanced speech database management and analysis in R,” Computer Speech & Language, vol. 45, pp. 392–410, Sep. 2017. [Online]. Available: https://linkinghub.elsevier.com/retrieve/pii/S0885230816302601

  43. [51]

    SPPAS - SPeech Phonetization Alignment and Syllabification

    “SPPAS - SPeech Phonetization Alignment and Syllabification.” [Online]. Available: https://sppas.org/

  44. [52]

    SPPAS: a tool for the phonetic segmentations of Speech,

    B. Bigi, “SPPAS: a tool for the phonetic segmentations of Speech,” in The eighth international conference on Language Resources and Evaluation, Istanbul, Turkey, May 2012, pp. 1748–1755. [Online]. Available: https://hal.science/hal-00983701

  45. [53]

    Annotald

    “Annotald.” [Online]. Available: https://annotald.github.io/

  46. [54]

    Inter-annotator Agreement,

    R. Artstein, “Inter-annotator Agreement,” in Handbook of Linguistic Annotation, N. Ide and J. Pustejovsky, Eds. Dordrecht: Springer Netherlands, 2017, pp. 297–313. [Online]. Available: https://doi.org/10. 1007/978-94-024-0881-2_11

  47. [55]

    Audacity ® | Free Audio editor, recorder, music making and more!

    “Audacity ® | Free Audio editor, recorder, music making and more!” [Online]. Available: https://www.audacityteam.org/

  48. [56]

    EMOVO Corpus: an Italian Emotional Speech Database

    G. Costantini, I. Iadarola, A. Paoloni, and M. Todisco, “EMOVO Corpus: an Italian Emotional Speech Database.”

  49. [57]

    Speech emotion recognition with artificial intelligence for contact tracing in the COVID-19 pandemic,

    F. Pucci, P. Fedele, and G. M. Dimitri, “Speech emotion recognition with artificial intelligence for contact tracing in the COVID-19 pandemic,” Cognitive Computation and Systems, vol. 5, no. 1, pp. 71–85, 2023, _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1049/ccs2.1207...

  50. [58]

    Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing,

    M. Federico, Y . Virkar, R. Enyedi, and R. Barra-Chicote, “Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing,” in Interspeech

  51. [59]

    For a mapping of the languages/dialects of Italy and regional varieties of Italian,

    P. Boula de Mareüil, E. Bilinski, F. Vernier, V . de Iacovo, and A. Romano, “For a mapping of the languages/dialects of Italy and regional varieties of Italian,” in New Ways of Analyzing Dialectal Variation, A. Thibault, M. Avanzi, N. L. Vecchio, and AliceMillour, Eds. Édition...

  52. [60]

    The "untamed

    B. Huszthy, “The "untamed" /s/ of Italian dialects: An overview of the singular behaviour of Italo-Romance sibilants,” Verbum – Analecta Neolatina, vol. 18, no. 1-2, pp. 189–214, Dec. 2017, number: 1-2. [Online]. Available: https://ojs.ppke.hu/verbum/article/view/345

  53. [61]

    Speech Analysis of Language Varieties in Italy,

    M. La Quatra, A. Koudounas, E. Baralis, and S. M. Siniscalchi, “Speech Analysis of Language Varieties in Italy,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M.-Y . K...

  54. [62]

    Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based V oice Conversion,

    A. Kamble, A. Tathe, S. Kumbharkar, A. Bhandare, and A. C. Mitra, “Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based V oice Conversion,” 2023, publisher: arXiv Version Number: 3. [Online]. Available: https://arxiv.org/abs/2311.14836

  55. [63]

    ITAcotron 2: Trans- fering English Speech Synthesis Architectures and Speech Features to Italian

    A. Favaro, L. Sbattella, R. Tedesco, and V . Scotti, “ITAcotron 2: Trans- fering English Speech Synthesis Architectures and Speech Features to Italian.”

  56. [64]

    Italian SpeechDat-Car database – ELRA Catalogue

    “Italian SpeechDat-Car database – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ELRA-S0144/

  57. [65]

    Wake Words & V oice Commands Speech Data: Italian (Italy)

    “Wake Words & V oice Commands Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/ monologue-speech-dataset/wake-words-and-commands-italian-italy

  58. [66]

    Travel Scripted Monologue Speech Data: Italian (Italy)

    “Travel Scripted Monologue Speech Data: Italian (Italy).” [Online]. Avail- able: https://www.futurebeeai.com/dataset/monologue-speech-dataset/ travel-scripted-speech-monologues-italian-italy

  59. [67]

    Telecom Scripted Monologue Speech Data: Italian (Italy)

    “Telecom Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https: //www.futurebeeai.com/dataset/monologue-speech-dataset/ telecom-scripted-speech-monologues-italian-italy

  60. [68]

    Retail & E-commerce Scripted Monologue Speech Data: Italian (Italy)

    “Retail & E-commerce Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/monologue-speech-dataset/ retail-scripted-speech-monologues-italian-italy

  61. [69]

    Healthcare Scripted Monologue Speech Data: Italian (Italy)

    “Healthcare Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https: //www.futurebeeai.com/dataset/monologue-speech-dataset/ healthcare-scripted-speech-monologues-italian-italy

  62. [70]

    Delivery & Logistics Scripted Monologue Speech Data: Italian (Italy)

    “Delivery & Logistics Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https: //www.futurebeeai.com/dataset/monologue-speech-dataset/ delivery-scripted-speech-monologues-italian-italy

  63. [71]

    BFSI Scripted Monologue Speech Data: Italian (Italy)

    “BFSI Scripted Monologue Speech Data: Italian (Italy).” [Online]. Avail- able: https://www.futurebeeai.com/dataset/monologue-speech-dataset/ bfsi-scripted-speech-monologues-italian-italy

  64. [72]

    Travel Call Center Speech Data: Italian (Italy)

    “Travel Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ travel-call-center-conversation-italian-italy

  65. [73]

    Telecom Call Center Speech Data: Italian (Italy)

    “Telecom Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ telecom-call-center-conversation-italian-italy

  66. [74]

    Real Estate Call Center Speech Data: Italian (Italy)

    “Real Estate Call Center Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ realestate-call-center-conversation-italian-italy

  67. [75]

    Healthcare Call Center Speech Data: Italian (Italy)

    “Healthcare Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ healthcare-call-center-conversation-italian-italy

  68. [76]

    BFSI Call Center Speech Data: Italian (Italy)

    “BFSI Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ bfsi-call-center-conversation-italian-italy

  69. [77]

    Retail & E-commerce Call Center Speech Data: Italian (Italy)

    “Retail & E-commerce Call Center Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ retail-call-center-conversation-italian-italy

  70. [78]

    Delivery & Logistics Call Center Speech Data: Italian (Italy)

    “Delivery & Logistics Call Center Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ delivery-call-center-conversation-italian-italy

  71. [79]

    Italian (Italy) General Conversation Speech Dataset

    “Italian (Italy) General Conversation Speech Dataset.” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ general-conversation-italian-italy

  72. [80]

    voxforge.org

    “voxforge.org.” [Online]. Available: https://www.voxforge.org/home

  73. [81]

    ASR-ItaCSC: An Italian Conversational Speech Corpus - MagicHub

    “ASR-ItaCSC: An Italian Conversational Speech Corpus - MagicHub.” [Online]. Available: https://magichub.com/datasets/ italian-conversational-speech-corpus/

  74. [82]

    Consonant gemination in Italian: the nasal and liquid case,

    M.-G. D. Benedetto and L. D. Nardis, “Consonant gemination in Italian: the nasal and liquid case,” Sep. 2020, arXiv:2005.06960. [Online]. Available: http://arxiv.org/abs/2005.06960

  75. [83]

    GDPR - Regulation - 2016/679 - EN - EUR-Lex,

    “GDPR - Regulation - 2016/679 - EN - EUR-Lex,” doc ID: 32016R0679 Doc Sector: 3 Doc Title: Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the fre...

  76. [2020]

    2020, pp

    ISCA, Oct. 2020, pp. 1481–1485. [Online]. Available: https: //www.isca-archive.org/interspeech_2020/federico20_interspeech.html

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.