Pith. sign in

REVIEW 2 cited by

ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.08720 v1 pith:KYN6QN45 submitted 2023-07-17 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords hebrewivritspeechdatasetresearchadvancingcomprehensivedata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce "ivrit.ai", a comprehensive Hebrew speech dataset, addressing the distinct lack of extensive, high-quality resources for advancing Automated Speech Recognition (ASR) technology in Hebrew. With over 3,300 speech hours and a over a thousand diverse speakers, ivrit.ai offers a substantial compilation of Hebrew speech across various contexts. It is delivered in three forms to cater to varying research needs: raw unprocessed audio; data post-Voice Activity Detection, and partially transcribed data. The dataset stands out for its legal accessibility, permitting use at no cost, thereby serving as a crucial resource for researchers, developers, and commercial entities. ivrit.ai opens up numerous applications, offering vast potential to enhance AI capabilities in Hebrew. Future efforts aim to expand ivrit.ai further, thereby advancing Hebrew's standing in AI research and technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A Hebrew G2P system that outputs fully specified IPA with stress, along with a new IPA-annotated speech corpus, improves phonetic accuracy of small real-time TTS models.

  2. HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A 30-second counting task, embedded with speaker-identification models, predicts male sleep apnea (AUC 0.64) and shows gender- and condition-specific model rankings across a new 7,188-recording clinical speech benchmark.

Pith tools