{"id":"16dfa2b6-9a06-4029-ac5d-e12b59e91bd6","arxiv_id":"2606.23144","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Koshur Pixel is the first large-scale synthetic OCR dataset for Kashmiri with 613,078 image-text pairs generated via SynthOCR-Gen from the KS-PRET-5M corpus across multiple fonts and granularities with 25+ augmentations.","lead":"This paper creates Koshur Pixel, a synthetic dataset of 613,078 image-text pairs for Kashmiri OCR in Nastaliq script using text from the KS-PRET-5M corpus and over 25 augmentations. A smart generalist might read it to learn how synthetic data can bootstrap language technologies for under-resourced scripts without manual labeling.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No empirical validation that synthetic pairs transfer to real Kashmiri documents","rationale":"The reader's weakest_assumption directly identifies the missing transfer evidence; the abstract-only view makes this the single load-bearing gap. No other internal inconsistency is visible from the given text.","tokens_in":1673,"tokens_out":291,"duration_ms":13289,"concrete_test":"Obtain or annotate a small held-out set of real scanned Kashmiri pages; train a baseline OCR model (e.g., CRNN or TrOCR) on the full Koshur Pixel training split and report character error rate on the real test set. If CER exceeds 25-30% or shows >2x degradation relative to performance on a synthetic validation split, the representativeness assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Koshur Pixel supplies a foundational, cost-effective resource for training effective OCR models. This requires the SynthOCR-Gen pipeline plus >25 augmentations to produce image-text pairs whose distribution of Nastaliq ligatures, contextual shaping, and degradations is close enough to real documents that models trained on them generalize. The abstract describes the corpus source, font variety, and augmentation count but supplies zero quantitative checks: no real-document test set, no CER/WER numbers, no ablation on augmentation realism, and no comparison against existing low-resource OCR baselines.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces Koshur Pixel as the first large-scale synthetic OCR dataset for Kashmiri, comprising 613,078 image-text pairs generated from the KS-PRET-5M corpus using the SynthOCR-Gen framework. It spans multiple fonts and textual granularities (words to full pages) and applies more than 25 augmentation strategies to emulate document degradations, positioning the dataset as a scalable, cost-effective alternative to manual annotation and a foundational resource for OCR training in this low-resource language with Nastaliq script challenges.","tokens_in":1775,"tokens_out":348,"duration_ms":16903,"significance":"If the synthetic generation process produces data that transfers effectively to real Kashmiri documents, the dataset could provide a valuable starting point for training OCR models where annotated real data is scarce, supporting digitization efforts for Kashmiri textual heritage. The scale and augmentation count are strengths, but the lack of any reported validation metrics means the practical significance remains unassessed.","major_comments":[{"comment":"Abstract: The central claim that Koshur Pixel 'establishes a foundational resource for training OCR systems' is load-bearing on the assumption that the SynthOCR-Gen pipeline and >25 augmentations produce image-text pairs representative of real Nastaliq ligatures, contextual shaping, and degradations. No quantitative support is supplied—no real-document test set, no CER/WER numbers on held-out data, no ablation on augmentation realism, and no comparison to existing low-resource OCR baselines—preventing assessment of whether models trained on the dataset generalize.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We agree that the abstract's claim requires qualification in the absence of empirical validation metrics, and we will revise the manuscript to address this directly while preserving the core contribution as a synthetic dataset release.","responses":[{"response":"We acknowledge the validity of this observation. The manuscript introduces a synthetic dataset to mitigate the absence of annotated Kashmiri OCR data, but the strong phrasing in the abstract does overstate the immediate utility without supporting experiments. Because no large-scale, publicly available annotated real-world Kashmiri document corpus exists for benchmarking, we cannot supply CER/WER results, ablation studies on real data, or baseline comparisons at present. We will revise the abstract to state that Koshur Pixel 'provides a large-scale synthetic resource intended to support the development of OCR systems' rather than claiming it 'establishes a foundational resource.' We will also add a dedicated limitations subsection discussing the synthetic nature of the data, the lack of real-document validation, and the need for future collection of annotated real samples. These changes will make the scope of the contribution explicit.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The central claim that Koshur Pixel 'establishes a foundational resource for training OCR systems' is load-bearing on the assumption that the SynthOCR-Gen pipeline and >25 augmentations produce image-text pairs representative of real Nastaliq ligatures, contextual shaping, and degradations. No quantitative support is supplied—no real-document test set, no CER/WER numbers on held-out data, no ablation on augmentation realism, and no comparison to existing low-resource OCR baselines—preventing assessment of whether models trained on the dataset generalize."}],"tokens_in":1316,"tokens_out":419,"duration_ms":11622,"standing_objections":["Provision of quantitative metrics (CER/WER, ablations, or real-document test sets) on actual Kashmiri documents, as no such annotated real-world resources are currently available for this low-resource language."]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main point is the release of Koshur Pixel, described as the first large synthetic OCR dataset for Kashmiri Nastaliq. It pulls text from the KS-PRET-5M corpus, renders it across fonts and word-to-page scales, and applies over 25 augmentations meant to copy real degradations.\n\nThis is new because Kashmiri has almost no public annotated OCR data, so a documented generation pipeline at this scale fills a clear gap for heritage work and low-resource script research.\n\nThe generation approach itself looks reasonable on paper: multiple fonts, contextual shaping, and standard augmentations. That part is straightforward and could save others from starting from scratch.\n\nThe soft spot is the complete absence of validation. The text gives scale and method details but no CER or WER on held-out real Kashmiri pages, no ablation on which augmentations matter, and no baseline comparison against existing low-resource OCR setups. Without those checks, readers cannot tell whether models trained on Koshur Pixel actually generalize.\n\nThe paper is aimed at people building OCR for under-resourced Perso-Arabic scripts or digitizing Kashmiri material. Anyone already working in that niche could use the data as a starting point, provided they run their own tests.\n\nIt deserves peer review as a data resource paper. Reviewers will almost certainly ask for transfer results, but the core claim of filling a documented gap is worth checking in detail.","headline":"Koshur Pixel gives a new 613k-pair synthetic Kashmiri OCR dataset but supplies no numbers showing it works on real documents.","tokens_in":2242,"tokens_out":371,"would_cite":false,"duration_ms":13031,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Koshur Pixel supplies the first large-scale synthetic OCR dataset for Kashmiri with 613,078 image-text pairs.","keywords":["synthetic OCR dataset","Kashmiri language","Nastaliq script","data augmentation","low-resource languages","optical character recognition","Perso-Arabic script"],"falsifier":"Train an OCR model on Koshur Pixel and measure its character or word error rate on a separate collection of scanned real-world Kashmiri documents.","tokens_in":2574,"feed_emoji":"","tokens_out":646,"duration_ms":18559,"temperature":0.7,"pith_summary":"The paper introduces Koshur Pixel to overcome the absence of annotated data that blocks optical character recognition for Kashmiri. Kashmiri text in Nastaliq script features contextual shaping, dense ligatures, and variability that demand large training sets. The authors generate the pairs from the KS-PRET-5M text corpus by rendering across fonts and layouts then applying more than 25 augmentations that mimic document wear and imaging artifacts. This synthetic route supplies a scalable substitute for hand-labeled data and creates a base resource for model training. If the generated examples transfer to real documents, the dataset supports digitization of Kashmiri texts and further language-technology work for this under-resourced language.","feed_headline":"First synthetic OCR dataset for Kashmiri offers 613k image-text pairs","feed_subtitle":"Generated from text corpus with over 25 augmentations to train recognition models for Nastaliq script","key_machinery":"SynthOCR-Gen framework that renders text from a large Kashmiri corpus into images across fonts and granularities while applying degradation augmentations.","core_discovery":"Koshur Pixel is a collection of 613,078 synthetic image-text pairs for Kashmiri, produced from the KS-PRET-5M corpus through the SynthOCR-Gen framework; the pairs cover multiple fonts, range from single words to full-page layouts, and include more than 25 augmentation strategies that emulate real-world degradations.","pith_inferences":["The same generation pipeline could be reused for other low-resource languages that employ Nastaliq or related Perso-Arabic scripts.","Models trained only on the synthetic set would probably benefit from later fine-tuning on even modest amounts of real scanned data.","Direct comparison of error rates between synthetic-only training and mixed synthetic-plus-real training would quantify the dataset's practical value."],"forward_implications":["Supplies a cost-effective substitute for manual annotation when building Kashmiri OCR systems.","Enables training of models that can digitize existing Kashmiri textual heritage.","Provides a foundation for language technologies serving a severely under-resourced language.","Covers multiple fonts and document scales from isolated words to complete pages."],"fun_headline_variants":["Koshur Pixel: 613k synthetic OCR pairs for Kashmiri","613k synthetic image-text pairs for Kashmiri OCR in Koshur Pixel","Koshur Pixel dataset with 613k pairs for Kashmiri text recognition","Synthetic Kashmiri OCR data: 613k pairs across fonts in Koshur Pixel","613k pairs from Koshur Pixel for Nastaliq script OCR training"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Synthetic images produced with the chosen fonts and augmentations are representative enough of real Kashmiri documents to train effective OCR models.","fun_headline_variants_meta":{"raw":{"variants":["Koshur Pixel: 613k synthetic OCR pairs for Kashmiri","613k synthetic image-text pairs for Kashmiri OCR in Koshur Pixel","Koshur Pixel dataset with 613k pairs for Kashmiri text recognition","Synthetic Kashmiri OCR data: 613k pairs across fonts in Koshur Pixel","613k pairs from Koshur Pixel for Nastaliq script OCR training"]},"model":"grok-4.3","cost_usd":0.008331,"raw_usage":{"total_tokens":3745,"prompt_tokens":610,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":83312000,"prompt_tokens_details":{"text_tokens":610,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3039,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":610,"tokens_out":96,"duration_ms":17391,"temperature":1.0,"reasoning_tokens":3039,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T09:17:36.511901+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Train an OCR model on Koshur Pixel and measure its character or word error rate on a separate collection of scanned real-world Kashmiri documents.","supporting_citations":[],"review_version":1}