Pith. sign in

REVIEW 4 major objections 6 minor 35 references

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

T0 review · 4 major / 6 minor · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Indic DiarBench is the first open joint diarization-and-ASR benchmark spanning all 22 scheduled Indian languages, with roughly 108 hours of human-corrected multi-speaker speech.

desk verdict Real infrastructure gap filled: all-22 Indic joint diarization+ASR labels with a public release and sane baselines—thin hours on 12 languages just mean you should not over-read the per-language rankings. read the letter →

arxiv 2607.23808 v1 pith:62AXTH25 submitted 2026-07-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords speakerdiarizationIndianlanguagesmultilingualbenchmarkspeaker-attributedASRcode-mixingconversationalspeechDERcpWER
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Progress in Indian-language speech recognition has mostly targeted clean single-speaker audio, yet real meetings and conversations involve overlapping talkers, English code-mixing, and wide dialectal range. This paper releases Indic DiarBench: about 108 hours of natural multi-speaker recordings that cover every one of India's 22 scheduled languages, drawn from near-field virtual meetings, far-field rooms, and in-the-wild YouTube conversations. Every file carries human-corrected, time-aligned transcripts that attribute each word to a speaker. The authors evaluate commercial speech APIs and multimodal models on the same mixed single-channel audio and show that even the strongest systems still produce large errors under high overlap and on lower-resource languages. The corpus, labels, and protocols are released openly so the field can measure joint speaker diarization and recognition under conditions that match Indian conversational speech.

What carries the argument

Indic DiarBench—the multilingual multi-condition corpus together with the joint metrics DER (acoustic segmentation), cpWER, and WDER (speaker-attributed transcription)—is the central object. It forces every system to be scored on identical mixed single-channel audio so diarization mistakes and recognition mistakes are measured together rather than in isolation.

What would settle it

Independent re-annotation of a large high-overlap subset that systematically changes speaker boundaries or word sequences and thereby reverses the reported ranking of systems on DER or cpWER would show the released labels are not yet a stable joint benchmark.

Watch

Extended reading notes

Core claim

No prior open resource jointly evaluates speaker diarization and speaker-attributed ASR across all 22 scheduled Indian languages under realistic multi-speaker conditions. Indic DiarBench supplies roughly 108 hours of such audio with human-corrected, time-aligned speaker transcripts, and the accompanying baselines show that current commercial APIs and multimodal models remain far from reliable—especially when speakers overlap or the language is lower-resource.

Load-bearing premise

The multi-stage human correction pipeline is assumed to yield speaker labels and transcripts accurate enough that differences in system error rates reflect true capability rather than leftover annotation noise, especially on heavily overlapping speech.

Editorial extensions

If this is right

  • Comparable joint diarization-plus-ASR numbers can now be reported on all 22 scheduled Indian languages instead of English-only or single-speaker sets.
  • Public baselines identify high-overlap segments and lower-resource languages as the dominant remaining failure modes.
  • Dual native-script and Romanized English reference transcripts allow fairer scoring of code-mixed system output.
  • Open RTTM and segment-level labels enable development of tightly coupled diarization-ASR pipelines rather than cascaded ones.
  • The same collection and annotation protocol can be extended to finish in-the-wild coverage for the remaining twelve languages.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because multimodal models show high missed-detection rates yet competitive transcription once segments are found, hybrid stacks that keep a strong diarizer in front of an LLM decoder are a natural architecture to test next.
  • The reported correlation between overlap ratio and error for the best Indic system implies that overlap-aware separation or training will move aggregate scores more than language-specific fine-tuning alone.
  • Weighting far-field and YouTube subsets more heavily in future leaderboards would better reflect true single-microphone difficulty than near-field virtual meetings alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Indic DiarBench, an open-access benchmark for joint speaker diarization and speaker-attributed ASR covering all 22 scheduled Indian languages. The corpus comprises ~108 hours of multi-speaker audio in three conditions: near-field meetings (~53 h, one microphone per non-co-located speaker, all 22 languages), far-field meetings (~27 h, 8 languages), and in-the-wild YouTube audio (~28 h, 10 languages). Annotations are produced by a five-stage human-in-the-loop pipeline (multi-system ASR bootstrap, professional correction, dual-format code-mixed transcription, QC, expert supercheck) yielding RTTM speaker timing and speaker-attributed transcripts. The authors evaluate seven systems (commercial APIs, multimodal LLMs, and the Sarvam pipeline) using DER (no collar, overlap included), cpWER, and WDER, with duration-weighted aggregates, DER decomposition, per-language heatmaps, and an overlap-correlation analysis. Headline findings: the Indic-specialized Sarvam pipeline leads (16.0% DER, 38.8% cpWER), multimodal LLMs trade strong transcription for very high missed detection, and overlap ratio correlates strongly with both DER and cpWER.

Significance. If the data and annotations are as described, this is a useful and genuinely novel resource: the first joint diarization + speaker-attributed ASR benchmark spanning all 22 scheduled Indian languages, a clear gap left by DISPLACE (which decoupled the ASR track and covered fewer languages). Strengths worth naming: the corpus and evaluation protocols are publicly released (reproducible resource); the near-field subset's one-mic-per-speaker, non-co-located design gives near-ground-truth speaker timing for ~53 hours and structurally prevents speaker-label invention there; the dual-format (native-script / Romanized) WER convention is a principled, stated choice for code-mixed evaluation; and the baseline suite produces concrete, falsifiable reference numbers across seven named systems with a DER error decomposition. The metrics (DER, cpWER, WDER) are standard community definitions, so there is no circularity concern. The benchmark is likely to be used and cited by the Indic speech community regardless of the reservations below.

major comments (4)
  1. [§3.2 (Annotation Pipeline)] No quantitative evidence of gold-label quality is provided. The five-stage pipeline is described, but there is no inter-annotator agreement measurement (e.g., DER/cpWER between double-annotated files, or kappa on speaker labels) on any subset. This matters most exactly where the benchmark is most valuable: the in-the-wild subset, where annotators may add/merge/remove speakers, and the high-overlap segments the paper itself says 'often require multiple rounds of review.' Without a label-noise estimate, the reader cannot tell whether, e.g., the 4–6 pp WDER gaps between systems in Table 3 exceed annotation noise. Please double-annotate a stratified sample (including high-overlap and in-the-wild files) and report agreement in the same units as the benchmark metrics.
  2. [§5 (Performance across languages) and Table 2] Per-language and language-family conclusions are drawn on very thin data without uncertainty quantification. Twelve languages (Assamese, Bodo, Dogri, Kashmiri, Konkani, Maithili, Manipuri, Nepali, Sanskrit, Santali, Sindhi, Urdu) exist only as ~1.1–1.6 h of near-field audio — a few hundred speaker turns each — yet §5 ranks them ('Santali ... 9.7% DER and Urdu ... 12.8% DER are among the easiest') and asserts a cross-family effect ('Dravidian languages show near-field cpWER roughly 5 percentage points above Indo-Aryan languages at comparable DER'). The Malayalam near-field row (1.3 h) also feeds the Dravidian average. Point estimates at this volume plausibly carry ±5–15 pp uncertainty. Please add bootstrap confidence intervals (or at minimum per-language session counts and a significance caveat) and soften the family-level claim accordingly. The duration-weighted aggregates in Table 3 are
  3. [§5 (Performance across languages) vs. Table 2] The text states 'Telugu emerges as the most challenging language, exhibiting the highest overlap (24.7%)', but Table 2 lists Telugu's overlap as 20.4%; 24.7% is Maithili's value (and Maithili has only near-field audio, so its table value equals its near-field value). Either the §5 numbers come from a near-field-only overlap computation not shown, or the value is misattributed. Since the 'Telugu most challenging' claim partly rests on this figure, please reconcile the text with Table 2 and state explicitly which condition each quoted overlap number refers to.
  4. [§4 (Models) and Table 3] The top-ranked system (Sarvam) is a product of the first authors' employer, and its reference [19] is a blog post. This is disclosed via affiliations and is not disqualifying, but the comparison's credibility requires stronger reproducibility commitments than are currently stated: exact model/API versions and access dates for all systems, the prompting/endpoint configuration used for GPT-4o and Gemini 3 Pro, and release of the scoring scripts and system outputs alongside the dataset. Additionally, excluding diarization-only models (e.g., Pyannote) is defensible for the joint task, but a DER-only baseline would contextualize the DER column and cost little; at minimum, the paper should note that Table 3 DERs are not comparable to diarization-only literature for that reason.
minor comments (6)
  1. [Figure 2] The heatmaps are difficult to parse at print size: language codes are non-standard abbreviations (Brx, Doi, Kok, Sat, etc.) without a legend, and several cells exceed 100% (e.g., Gemini cpWER 103, 116) which deserves a one-line explanation (insertions on low-word-count languages).
  2. [§4 (Metrics)] DER is computed 'without a forgiveness collar and including overlapping speech' — good — but please state the scoring tool/version (e.g., dscore / pyannote.metrics) and confirm the same convention for the Miss/FA/Conf decomposition, since collar conventions differ across prior benchmarks.
  3. [§3.3 / Table 2] The 'Overlap %' column is described as 'averaged over all conditions,' but for 12 languages only one condition exists; clarify whether overlap is time-weighted or session-averaged, and how the total row (12.8%) is computed.
  4. [§3.3 (Limitations)] Stating that no speaker IDs are released for the in-the-wild subset is appreciated; please also state whether the in-the-wild audio itself is redistributed or released as YouTube URLs + timestamps, since link rot will affect reproducibility.
  5. [Acknowledgments] Typo: 'Sshubam' likely should be 'Shubam'.
  6. [Table 1] The DISPLACE '24 row lists 38 hours; the text (§2) says 158 hours with 38 labelled. Align the table cell with the labelled-hours figure and note the distinction in the caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: benchmark paper with external metrics and held-out evaluation; results do not reduce to fitted inputs or self-definition.

full rationale

Indic DiarBench is a dataset-and-baselines paper, not a first-principles derivation. Its load-bearing claims are (i) construction of ~108h multi-speaker audio with human-corrected RTTM and speaker-attributed transcripts across 22 languages, and (ii) evaluation of third-party and in-house systems under community metrics (DER without collar, cpWER, WDER). Those metrics are imported from prior external literature (CHiME-6, El Shafey et al.), not redefined in terms of the paper’s outputs. Baselines are run on the collected evaluation audio; no parameter is fitted to a subset and then re-reported as a “prediction.” Self-citations (IndicVoices for recruitment principles; Sarvam ASR as one evaluated system) are ordinary context or COI texture—they do not force the ranking or the coverage claim by construction. Author-affiliated Sarvam achieving the best aggregate numbers is a conflict-of-interest concern, not definitional circularity. There is no uniqueness theorem, ansatz smuggled via self-citation, or renaming of a known empirical law. Per the analyzer rules, honest non-finding applies: score 0, empty steps.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

As a dataset-and-baseline paper, load-bearing commitments are domain conventions (standard diarization/ASR metrics, human gold labels, single-channel mixed evaluation) rather than fitted physical constants or invented particles. Free parameters are essentially none in a theory sense; the main soft choices are curation filters, dual code-mix transcript acceptance, and which commercial systems to call.

assumptions (5)
  • domain assumption DER without forgiveness collar including overlap, cpWER, and WDER are appropriate joint measures of diarization and speaker-attributed ASR.
    Invoked in Evaluation Setup §4; standard in CHiME/diarization literature the paper cites, not re-derived.
  • domain assumption Human-corrected transcripts and RTTM after the described multi-stage pipeline constitute gold reference for ranking systems.
    Annotation Pipeline §3.2; no IAA statistics provided, so this is an unquantified quality assumption.
  • domain assumption Evaluating all systems on the same single-channel mixed audio is a fair comparison of joint ASR–diarization capability.
    Stated in Results §5; excludes multi-channel oracle advantages available in the near-field collection setup.
  • ad hoc to paper Accepting either native-script or Romanized-English normalized references in WER avoids unfairly penalizing orthographic convention differences.
    Code-Mixed Transcription step in §3.2; reasonable but paper-specific scoring choice.
  • ad hoc to paper Diarization-only models (e.g., Pyannote) can be excluded because modern applications require joint outputs.
    Models paragraph in §4; narrows the baseline set by design.
invented entities (1)
  • Indic DiarBench corpus independent evidence
    purpose: Provide the evaluation resource and labels on which all baseline claims rest.
    New dataset constructed by the authors; existence is evidenced by the release link and statistics, but quality depends on the annotation axiom above.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages." pith.science (2026). https://pith.science/paper/62AXTH25

@misc{pith2026260723808,
  author       = {Pith},
  title        = {Pith review of: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/62AXTH25}},
  note         = {Machine review of arXiv:2607.23808}
}
read the original abstract

In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.

Figures

Figures reproduced from arXiv: 2607.23808 by the authors.

Figure 1
Figure 1. Speaker Data Collection by District and Language To address this gap, we introduce Indic DiarBench, an open-access benchmark for evaluating speaker-attributed ASR in multilingual Indian conversational speech. Our contributions can be summarised as follows: • A corpus of 108 hours of conversational speech covering all 22 scheduled Indian languages, collected across near-field meetings, far-field recordings, and in-th… view at source ↗
Figure 2
Figure 2. Per-language (a) cpWER and (b) WDER (%) heatmap [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Metrics variation vs overlap ratio 5. Results and Analysis For evaluation, all systems were provided the same single￾channel mixed audio to ensure fairness [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 1 linked inside Pith

  1. [19]

    SDBench: A comprehensive benchmark suite for speaker diarization,

    B. Durmus, B. Munyampirwa, E. Pacheco, A. Orhon, and A. Leonov, “SDBench: A comprehensive benchmark suite for speaker diarization,” inProc. Interspeech, 2025, pp. 1598–1602

  2. [1]

    Large scale data collection efforts such as IndicV oices [1] have enabled multilingual ASR systems that begin to cover India’s linguis- tic diversity

    Introduction Recent years have seen significant progress in automatic speech recognition (ASR) for Indian languages. Large scale data collection efforts such as IndicV oices [1] have enabled multilingual ASR systems that begin to cover India’s linguis- tic diversity. However, most of this progress has focused on single speaker speech, while many real worl...

  3. [2]

    In- dic DiarBench is the first to cover all 22 Indian scheduled lan- guages with joint ASR + diarization labels

    Related Work Early diarization benchmarks such as the AMI [3] and ICSI [4] meeting corpora provided multi-channel English arXiv:2607.23808v1 [cs.CL] 26 Jul 2026 Table 1:Comparison of diarization evaluation datasets. In- dic DiarBench is the first to cover all 22 Indian scheduled lan- guages with joint ASR + diarization labels. Dataset Lang. Hrs Domain Spk...

  4. [3]

    The follow- ing subsections describe the data collection process, annotation pipeline, and dataset statistics

    The Indic DiarBench Corpus Indic DiarBench is a multilingual conversational speech benchmark designed to evaluate speaker attributed ASR in re- alistic multi speaker settings for Indian languages. The follow- ing subsections describe the data collection process, annotation pipeline, and dataset statistics. 3.1. Data Collection We now describe the recordin...

  5. [4]

    For acoustic segmentation, we report Diarization Error Rate (DER) computed without a forgiveness collar and includ- ing overlapping speech

    Evaluation Setup Metrics.We report evaluation metrics along two complemen- tary axes: acoustic diarization and word-level speaker attribu- tion. For acoustic segmentation, we report Diarization Error Rate (DER) computed without a forgiveness collar and includ- ing overlapping speech. To jointly evaluate ASR and diarization performance, we Table 2:Per-lang...

  6. [5]

    Figure 2 presents per-language cpWER and WDER for all evaluated systems across 22 Indic languages

    Results and Analysis For evaluation, all systems were provided the same single- channel mixed audio to ensure fairness. Figure 2 presents per-language cpWER and WDER for all evaluated systems across 22 Indic languages. Grey cells indicate unsupported lan- guages; consequently, global averages over languages are inher- ently skewed, and we focus instead on...

  7. [6]

    Conclusion We presentIndic DiarBench, the first open benchmark for joint diarization and speaker-attributed ASR spanning all 22 scheduled Indian languages. By unifying near-field meetings, far-field recordings, and in-the-wild conversations, the bench- mark captures realistic variation in speaker counts, overlap ra- tios, and acoustic conditions, establis...

  8. [7]

    We also thank the language experts at Sarvam AI and AI4Bharat for their excellent work; this effort would not have been possible without their contributions

    Acknowledgments We thank Sshubam, Sadakopa, and Vamsi from Sarvam AI for generously giving their time and helping with the YouTube data collection effort. We also thank the language experts at Sarvam AI and AI4Bharat for their excellent work; this effort would not have been possible without their contributions

Show all 35 references
  1. [8]

    All technical content, analyses, results, and conclusions were produced and verified by the authors

    Generative AI use disclosure Generative AI tools were used only for limited language editing and polishing of parts of the manuscript. All technical content, analyses, results, and conclusions were produced and verified by the authors

  2. [9]

    IndicV oices: Towards building an inclusive multilingual speech dataset for Indian languages,

    T. Javed, J. Nawale, E. I. George, S. Joshi, K. S. Bhogaleet al., “IndicV oices: Towards building an inclusive multilingual speech dataset for Indian languages,” inFindings of ACL, 2024, pp. 10 740–10 782

  3. [10]

    A review of speaker diarization: Recent advances with deep learning,

    T. J. Park, N. Kanda, D. Dimitriadis, K. J. Han, S. Watanabe, and S. Narayanan, “A review of speaker diarization: Recent advances with deep learning,”Computer Speech & Language, vol. 72, p. 101317, 2022

  4. [11]

    The AMI meeting corpus: A pre-announcement,

    J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hainet al., “The AMI meeting corpus: A pre-announcement,” inProc. Machine Learning for Multimodal Interaction (MLMI), 2005, pp. 28–39

  5. [12]

    The ICSI meeting corpus,

    A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, and C. Wooters, “The ICSI meeting corpus,” inProc. ICASSP, 2003, pp. 364–367

  6. [13]

    CALLHOME American English speech,

    A. Canavan, D. Graff, and G. Zipperlen, “CALLHOME American English speech,” 1997, lDC97S42

  7. [14]

    The third DI- HARD diarization challenge,

    N. Ryant, P. Singh, V . Krishnamohan, R. Varma, K. Church, C. Cieri, J. Du, S. Ganapathy, and M. Liberman, “The third DI- HARD diarization challenge,” inProc. Interspeech, 2021, pp. 3570–3574

  8. [15]

    Spot the conversation: Speaker diarisation in the wild,

    J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, “Spot the conversation: Speaker diarisation in the wild,” inProc. Interspeech, 2020, pp. 299–303

  9. [16]

    M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,

    F. Yu, S. Zhang, Y . Fu, L. Xieet al., “M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,” inProc. ICASSP, 2022, pp. 6167–6171

  10. [17]

    AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,

    Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” inProc. Inter- speech, 2021, pp. 3665–3669

  11. [18]

    Continuous speech separation: Dataset and anal- ysis,

    Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y . Luo, X. Xiao, J. Li, and J. Wu, “Continuous speech separation: Dataset and anal- ysis,” inProc. ICASSP, 2020, pp. 7284–7288

  12. [20]

    NOTSOFAR-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,

    A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gur- vich, S. Peer, X. Xiao, B. M. Elizalde, N. Kanda, X. Wang, S. Shaer, S. Yagev, Y . Asher, S. Sivasankaran, Y . Gong, M. Tang, H. Wang, and E. Krupka, “NOTSOFAR-1 challenge: New datasets, baseline, and tasks for...

  13. [21]

    Common voice: A massively-multilingual speech corpus,

    R. Ardilaet al., “Common voice: A massively-multilingual speech corpus,” inProc. LREC, 2020, pp. 4218–4222

  14. [22]

    FLEURS: Few-shot learning evaluation of universal representations of speech,

    A. Conneau, M. Ma, S. Khanuja, Y . Zhang, V . Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “FLEURS: Few-shot learning evaluation of universal representations of speech,” in2022 IEEE Spoken Language Technology Workshop (SLT), 2023, pp. 798– 805

  15. [23]

    The DISPLACE challenge 2023 – DIarization of SPeaker and LAnguage in Conversational Environments,

    S. Baghel, S. Ramoji, Sidharth, R. H, P. Singh, S. Jain, P. R. Chowdhuri, K. Kulkarni, S. Padhi, D. Vijayasenan, and S. Gana- pathy, “The DISPLACE challenge 2023 – DIarization of SPeaker and LAnguage in Conversational Environments,” inProc. Inter- speech, 2023, pp. 3562–3566

  16. [24]

    The second DISPLACE challenge: DIariza- tion of SPeaker and LAnguage in Conversational Environments,

    S. B. Kalluri, P. Singh, P. R. Chowdhuri, A. Kulkarni, S. Baghel, P. Hegde, S. Sontakke, D. K T, S. R. M. Prasanna, D. Vijayasenan, and S. Ganapathy, “The second DISPLACE challenge: DIariza- tion of SPeaker and LAnguage in Conversational Environments,” inProc. Interspeech, 202...

  17. [25]

    CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,

    S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Araki, X. Changet al., “CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,” inProc. CHiME Workshop, 2020

  18. [26]

    Joint speech recogni- tion and speaker diarization via sequence transduction,

    L. El Shafey, H. Soltau, and I. Shafran, “Joint speech recogni- tion and speaker diarization via sequence transduction,” inProc. Interspeech, 2019, pp. 396–400

  19. [27]

    Sarvam ASR,

    Sarvam AI, “Sarvam ASR,” https://www.sarvam.ai/blogs/asr/, ac- cessed: 2026-03-05

  20. [28]

    Introducing Nova-3 Speech-to-Text API,

    Deepgram, “Introducing Nova-3 Speech-to-Text API,” https:// deepgram.com/learn/introducing-nova-3-speech-to-text-api, ac- cessed: 2026-03-05

  21. [29]

    Speech to Text Capabilities,

    ElevenLabs, “Speech to Text Capabilities,” https://elevenlabs.io/ docs/overview/capabilities/speech-to-text, accessed: 2026-03-05

  22. [30]

    Universal-2 Speech Recognition Model,

    AssemblyAI, “Universal-2 Speech Recognition Model,” https:// www.assemblyai.com/universal-2, accessed: 2026-03-05

  23. [31]

    Azure Speech-to-Text,

    Microsoft Azure, “Azure Speech-to-Text,” https://learn.microsoft. com/en-us/azure/ai-services/speech-service/speech-to-text, ac- cessed: 2026-03-05

  24. [32]

    Amazon Transcribe,

    Amazon Web Services, “Amazon Transcribe,” https://aws. amazon.com/transcribe/, accessed: 2026-03-05

  25. [33]

    Gemini 3,

    Google, “Gemini 3,” https://blog.google/products-and-platforms/ products/gemini/gemini-3/, accessed: 2026-03-05

  26. [34]

    GPT-4o Transcribe Model Documentation,

    OpenAI, “GPT-4o Transcribe Model Documentation,” https: //developers.openai.com/api/docs/models/gpt-4o-transcribe, ac- cessed: 2026-03-05

  27. [35]

    pyannote.audio 2.1 speaker diarization pipeline: Prin- ciple, benchmark, and recipe,

    H. Bredin, “pyannote.audio 2.1 speaker diarization pipeline: Prin- ciple, benchmark, and recipe,” inProc. Interspeech, 2023, pp. 1983–1987

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.