{"id":"b42b2e13-2ddb-4383-909f-dd053d49a67d","arxiv_id":"2606.25246","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces WBCMor VQA benchmark with 110K bilingual QA pairs for hematology VQA on 20K cell images using existing annotations and a new Urdu dictionary.","lead":"The paper introduces WBCMor VQA, a new bilingual English-Urdu visual question answering dataset with 110K pairs for 20K white blood cell images focused on leukemia and normal cells. A smart generalist might read it to see how AI resources can be adapted for language barriers in healthcare settings like Pakistan where Urdu is primary but systems remain English-based.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No details provided on QA generation or clinical validation process","rationale":"The reader's weakest_assumption directly matches the load-bearing gap; the absence of construction and validation details prevents any stronger verdict even if full text were supplied.","tokens_in":1791,"tokens_out":223,"duration_ms":12573,"concrete_test":"Locate the methods section describing QA-pair construction and validation; if it contains the generation algorithm, inter-rater agreement on a sample, or expert review results, re-assess the claim; if these elements remain absent, the clinical-validation assertion stays unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that repurposed annotations from LeukemiaAttri/WBCAtt plus a new Urdu hematology dictionary yield 110K clinically accurate bilingual QA pairs. The abstract states this is 'clinically validated' and ensures 'clinical correctness' but supplies no generation procedure, translation rules, expert review protocol, or quantitative validation metrics (e.g., agreement scores). This assumption is therefore unevaluable from the given text.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces WBCMor VQA, a bilingual (English-Urdu) morphology-aware VQA benchmark for leukemia and normal white blood cell analysis. It is constructed from morphology-aware annotations in the LeukemiaAttri and WBCAtt datasets, augmented by a domain-specific Urdu hematology dictionary, yielding 110K bilingual QA pairs over 20K single-cell images. The work also reports results from a survey of healthcare professionals on language mismatches and provides baseline evaluations of several open-source VLMs on the new benchmark.","tokens_in":1858,"tokens_out":431,"duration_ms":18822,"significance":"If the construction process and clinical validation can be shown to be sound and reproducible, the resource would address a documented gap in multilingual medical VQA resources for Urdu-speaking regions. The bilingual design and morphology focus could support development of more accessible clinical AI tools; the reported baselines would serve as a useful reference point for subsequent work.","major_comments":[{"comment":"Abstract: the central claim that the benchmark is 'clinically validated' and ensures 'clinical correctness' rests on the repurposing of annotations from LeukemiaAttri/WBCAtt plus a new Urdu dictionary, yet no generation procedure, translation rules, expert review protocol, or quantitative validation metrics (e.g., agreement scores) are described. This information is required to assess the validity of the 110K QA pairs.","section":"Abstract"},{"comment":"Dataset construction description (wherever presented): the mapping from the 20K images and existing morphology annotations to the final 110K bilingual QA pairs is not specified. Details on question templates, how morphology awareness is preserved in Urdu, dictionary application rules, and any filtering or quality-control steps are absent, preventing evaluation of reproducibility and clinical fidelity.","section":"Dataset construction"}],"minor_comments":[{"comment":"Abstract: 'releveant' is a typographical error and should read 'relevant'.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive report. The comments correctly identify that the current manuscript lacks explicit descriptions of the QA-pair generation process, translation methodology, expert review steps, and quantitative validation metrics. We will revise the manuscript to supply these details in a dedicated methods subsection, thereby strengthening the claims of clinical validation and reproducibility.","responses":[{"response":"We agree that the abstract and main text currently assert clinical validation without providing the supporting procedural details. The benchmark re-uses morphology annotations already present in LeukemiaAttri and WBCAtt (which were produced by hematologists) and augments them with a newly compiled Urdu hematology dictionary; however, the exact template-based question generation rules, how morphology terms were mapped while preserving clinical meaning in Urdu, the expert review protocol used to verify the dictionary, and any inter-annotator agreement statistics are not reported. We will add a new subsection (e.g., Section 3.2) that documents these steps, including the dictionary construction process and any quantitative checks performed.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that the benchmark is 'clinically validated' and ensures 'clinical correctness' rests on the repurposing of annotations from LeukemiaAttri/WBCAtt plus a new Urdu dictionary, yet no generation procedure, translation rules, expert review protocol, or quantitative validation metrics (e.g., agreement scores) are described. This information is required to assess the validity of the 110K QA pairs."},{"response":"The referee is correct that the mapping procedure is underspecified. The 110K bilingual pairs are generated by applying a fixed set of morphology-aware question templates to the existing attribute annotations, followed by automatic translation via the domain dictionary and a small number of manual corrections. Because these templates, dictionary application rules, and quality-control filters are not described, reproducibility cannot be assessed from the current text. We will expand the dataset-construction section with pseudocode or explicit rules for template instantiation, the dictionary lookup procedure, and the filtering criteria applied to remove low-quality or duplicate pairs.","revision_made":"yes","referee_comment":"[Dataset construction] Dataset construction description (wherever presented): the mapping from the 20K images and existing morphology annotations to the final 110K bilingual QA pairs is not specified. Details on question templates, how morphology awareness is preserved in Urdu, dictionary application rules, and any filtering or quality-control steps are absent, preventing evaluation of reproducibility and clinical fidelity."}],"tokens_in":1390,"tokens_out":536,"duration_ms":13729,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that the authors release WBCMor VQA: 110K English-Urdu question-answer pairs over 20K single-cell white blood cell images, built from existing morphology annotations in LeukemiaAttri and WBCAtt plus a new Urdu hematology dictionary. They also run a small survey of healthcare workers to show language mismatches in Pakistani clinical settings and report baseline numbers for several open VLMs.\n\nThe dataset itself is new and targets a concrete regional gap. Extending English-only hematology resources to Urdu with morphology-aware labels is a reasonable step, and the motivation from the survey is straightforward. Providing baselines gives others a quick way to compare models on this data.\n\nThe weak point is the missing methodology. The abstract and text claim the pairs are \"clinically validated\" and ensure \"clinical correctness,\" yet supply no description of how the questions were generated from the annotations, what translation or adaptation rules were applied, who performed the review, or any quantitative checks such as agreement scores or error rates. Without those steps, the central claim cannot be assessed. The survey is mentioned but not detailed either.\n\nThis paper is for researchers building or testing medical VLMs who need non-English data in hematology. A reader focused on multilingual medical AI could extract value from the resource once the data and full construction details are public.\n\nI would send it for peer review. The targeted gap is real and a documented bilingual dataset would be worth referee attention, provided the authors add the missing generation and validation sections.","headline":"New bilingual hematology VQA dataset released but the construction and validation steps remain undescribed.","tokens_in":2325,"tokens_out":371,"would_cite":false,"duration_ms":16986,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A bilingual English-Urdu dataset supplies 110,000 question-answer pairs for 20,000 white blood cell images.","keywords":["multilingual VQA","hematology","leukemia","white blood cells","Urdu","visual question answering","bilingual dataset","medical imaging"],"falsifier":"An independent review by hematologists and Urdu-speaking clinicians that finds a high rate of clinical inaccuracies or inconsistencies in a random sample of the generated question-answer pairs would show the benchmark lacks clinical validation.","tokens_in":2709,"feed_emoji":"🩸","tokens_out":670,"duration_ms":23298,"temperature":0.7,"pith_summary":"The paper aims to address the English-only focus of existing medical vision-language resources by building a benchmark that supports visual question answering in both English and Urdu for leukemia and normal white blood cells. A survey of healthcare professionals showed frequent mismatches between English-based clinical systems and Urdu used in patient communication, particularly in South Asia. The authors construct the benchmark by combining morphology annotations from two existing cell datasets with a specialized Urdu hematology dictionary to generate consistent bilingual pairs. This creates a resource of 110K question-answer pairs across 20K single-cell images, along with baseline evaluations of open-source vision-language models.","feed_headline":"Bilingual dataset adds Urdu to leukemia cell VQA","feed_subtitle":"110K question-answer pairs now cover 20K cell images in English and Urdu for regions where clinical systems and patient talk differ in langu","key_machinery":"The WBCMor VQA benchmark, built by repurposing morphology annotations from existing datasets and extending them with an Urdu hematology dictionary to produce bilingual question-answer pairs.","core_discovery":"The authors introduce WBCMor VQA as a clinically validated bilingual benchmark containing 110K English-Urdu question-answer pairs that annotate 20K single-cell images of leukemic and normal white blood cells. The benchmark is assembled from morphology-aware annotations in prior datasets and supported by a domain-specific Urdu hematology dictionary to preserve clinical accuracy and linguistic consistency. Baseline performance results from multiple open-source vision-language models are reported on the new resource.","pith_inferences":["Performance differences between English and Urdu questions on the benchmark could highlight where current models struggle with technical medical terms across languages.","If the dictionary approach scales, similar resources could be built for other languages facing English-dominant medical AI systems.","Integration into clinical tools might allow direct Urdu responses to image-based queries without separate translation steps."],"forward_implications":["Open-source vision-language models can be tested and trained on medical image questions that include Urdu terminology.","The resource supports AI systems capable of handling both English documentation and Urdu patient communication in hematology.","The construction method offers a pattern for generating similar bilingual benchmarks in other medical imaging domains.","The 20K annotated images provide a base for additional tasks such as cell classification or report generation in two languages."],"fun_headline_variants":["WBCMor VQA adds Urdu to leukemia cell VQA","English Urdu benchmark for 20K white blood cell images","110K bilingual pairs for leukemia and normal cell VQA","Morphology aware Urdu English hematology VQA dataset"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Repurposing annotations from prior datasets and adding a new Urdu dictionary yields clinically accurate and consistent bilingual question-answer pairs.","fun_headline_variants_meta":{"raw":{"variants":["WBCMor VQA adds Urdu to leukemia cell VQA","English Urdu benchmark for 20K white blood cell images","110K bilingual pairs for leukemia and normal cell VQA","Morphology aware Urdu English hematology VQA dataset"]},"model":"grok-4.3","cost_usd":0.003964,"raw_usage":{"total_tokens":1957,"prompt_tokens":689,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":39640500,"prompt_tokens_details":{"text_tokens":689,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1203,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":689,"tokens_out":65,"duration_ms":8271,"temperature":1.0,"reasoning_tokens":1203,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-25T21:06:35.150778+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An independent review by hematologists and Urdu-speaking clinicians that finds a high rate of clinical inaccuracies or inconsistencies in a random sample of the generated question-answer pairs would show the benchmark lacks clinical validation.","supporting_citations":[],"review_version":1}