{"id":"9a8aa62f-a0cc-4967-8831-8ef5183d8bee","arxiv_id":"2605.28253","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"First dedicated ASR corpus of 66 hours and systematic benchmarks for Puno Quechua using participatory collection and open release of data and fine-tuned models.","lead":"The authors created the largest speech corpus for Puno Quechua with 66 hours of recordings and 36 hours of transcriptions via community participation, plus the first ASR benchmarks on models like Whisper and wav2vec2. A smart generalist might read it to see how NLP tools can be built for endangered languages in an inclusive way.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption flags missing quality metrics, but those metrics are not load-bearing for the paper's actual claim of having collected and released the first dedicated resources and benchmark. The claim is primarily descriptive and release-oriented; quality concerns would matter for downstream use but do not undermine the reported existence or novelty of the contribution.","tokens_in":1714,"tokens_out":239,"duration_ms":25136,"concrete_test":"Confirm that the released datasets on the linked repository match the reported 66 h / 36 h totals and that the fine-tuned model checkpoints load and run inference on the described test audio.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the presentation of new resources (corpus + benchmark) rather than a performance breakthrough. The participatory collection and manual transcription are described at a high level; absent any internal contradiction or missing step that would invalidate the existence of the released data, the claim that the resources exist and are the first of their kind for this variety holds on the paper's own terms. No quantitative quality metrics are required for the descriptive claim itself.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript claims to introduce the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): a 66-hour speech corpus (36 hours manually transcribed and validated) collected via participatory design for scripted and spontaneous speech; the first systematic ASR benchmark evaluating state-of-the-art models and fine-tuning Whisper-base, wav2vec2-base, and XLS-R-300M with and without continued pre-training; and the open release of all datasets and fine-tuned models.","tokens_in":1789,"tokens_out":256,"duration_ms":23764,"significance":"If the resources exist as described, the work is significant for under-resourced language preservation through community-centered data collection and open science practices. The participatory design campaign and explicit open release of datasets and models are explicit strengths that enable reproducibility and downstream research on Quechua varieties.","major_comments":[],"minor_comments":[{"comment":"Abstract: the benchmark description states that models were evaluated and fine-tuned but provides no quantitative results, error analysis, or data-split details; adding a brief summary of key metrics would improve reader assessment of the benchmark's scope.","section":"Abstract"}],"recommendation":"accept","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive review, recognition of the significance of the participatory data collection and open release practices, and recommendation to accept the manuscript.","responses":[],"tokens_in":1112,"tokens_out":49,"duration_ms":16532,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that the authors collected 66 hours of Puno Quechua speech (36 hours transcribed and validated) through a participatory campaign, ran a basic benchmark on Whisper, wav2vec2, and XLS-R with and without continued pre-training, and released the data and models. That fills a clear gap for this specific variety.\n\nThe participatory approach and open release are the parts that work well. They match the stated goal of community-centered resources and give others something concrete to build on. The claim of being the first systematic benchmark for this variety holds on the paper's own terms since no prior work is cited for it.\n\nThe soft spot is the absence of any quantitative details on transcription accuracy, speaker demographics, data splits, or actual model error rates. Without those, it's hard to judge how usable the corpus really is for downstream work or whether the fine-tuning experiments show meaningful gains. The abstract alone leaves the soundness of the benchmark claim thin, though the stress-test note is right that the core existence claim does not require those numbers.\n\nThis is for people working on low-resource ASR or language documentation in Andean languages. A reader focused on data collection methods or specific language varieties would find the release useful. It is not a methodological advance, but resource papers like this still matter for preservation.\n\nI would send it to peer review. The data release itself justifies the effort even if the technical results need more detail in revision.","headline":"This paper releases the first dedicated speech corpus and ASR benchmark for Puno Quechua, which is a straightforward resource contribution worth documenting.","tokens_in":2246,"tokens_out":366,"would_cite":false,"duration_ms":21451,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Puno Quechua gains its first dedicated speech corpus and ASR benchmarks from community-collected data.","keywords":["Puno Quechua","automatic speech recognition","speech corpus","participatory design","low-resource languages","model fine-tuning","Whisper","wav2vec2"],"falsifier":"If fine-tuned models show no measurable word-error-rate improvement over their zero-shot baselines when tested on held-out Puno Quechua audio, the claimed utility of the new corpus and benchmarks would be refuted.","tokens_in":2627,"feed_emoji":"🗣️","tokens_out":663,"duration_ms":35040,"temperature":0.7,"pith_summary":"The paper sets out to build the first automatic speech recognition resources tailored to Puno Quechua by running a participatory design campaign that gathers 66 hours of scripted and spontaneous recordings. Thirty-six hours receive manual transcription and validation, after which several current models receive evaluation and fine-tuning with and without continued pre-training before all data and models are released openly. A sympathetic reader would care because this supplies concrete digital tools for a language variety that previously had none, created with direct speaker involvement rather than external scraping alone. The work therefore links resource creation to community agency in language technology.","feed_headline":"66-hour corpus gives Puno Quechua first ASR benchmarks","feed_subtitle":"Participatory recordings yield largest Quechua-variety dataset plus open fine-tuned models for this under-resourced language.","key_machinery":"Participatory design campaign that collects, transcribes, and validates speech data, followed by fine-tuning of pre-trained ASR models on the resulting Puno Quechua corpus.","core_discovery":"The first dedicated ASR resources for Puno Quechua consist of the largest speech corpus for any single Quechua variety at 66 hours of recordings with 36 hours of manually transcribed and validated data collected via participatory design, the first systematic benchmark that evaluates state-of-the-art models and fine-tunes Whisper-base, wav2vec2-base, and XLS-R-300M with and without continued pre-training, plus open release of all datasets and fine-tuned models.","pith_inferences":["The same participatory collection method could be replicated for other Quechua varieties or unrelated low-resource languages.","Continued pre-training on small target-language corpora may prove a reusable tactic for improving ASR accuracy in similar settings.","The released models could serve as starting points for downstream tasks such as spoken language documentation or translation aids."],"forward_implications":["Puno Quechua speakers obtain openly available datasets and models that can support building local speech applications.","Subsequent ASR research on Quechua varieties obtains a concrete baseline for comparison.","Community-shaped data collection produces resources aligned with actual scripted and spontaneous usage patterns.","Open release enables other researchers to extend or adapt the resources without starting from scratch."],"fun_headline_variants":["Puno Quechua first ASR from 66-hour corpus","Participatory 66-hour corpus for Puno Quechua ASR","Benchmarks for Puno Quechua using 66-hour speech data","Open release of ASR models for Puno Quechua"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The participatory design campaign and manual transcription produce data of sufficient quality, diversity, and representativeness to support effective ASR model development and benchmarking.","fun_headline_variants_meta":{"raw":{"variants":["Puno Quechua first ASR from 66-hour corpus","Participatory 66-hour corpus for Puno Quechua ASR","Benchmarks for Puno Quechua using 66-hour speech data","Open release of ASR models for Puno Quechua"]},"model":"grok-4.3","cost_usd":0.012961,"raw_usage":{"total_tokens":5589,"prompt_tokens":594,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":129612000,"prompt_tokens_details":{"text_tokens":594,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4931,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":594,"tokens_out":64,"duration_ms":39771,"temperature":1.0,"reasoning_tokens":4931,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T13:22:51.571003+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If fine-tuned models show no measurable word-error-rate improvement over their zero-shot baselines when tested on held-out Puno Quechua audio, the claimed utility of the new corpus and benchmarks would be refuted.","supporting_citations":[],"review_version":1}