{"id":"d47ba08a-51dc-4604-bd79-9198dbdcd857","arxiv_id":"2508.09087","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A clustering and reweighting pipeline with a CLIP-plus-SimCLR joint objective reduces subgroup disparities in VLM glaucoma screening without protected attribute labels, per FairVLMed glaucoma-subset metrics.","lead":"The paper proposes a training method that reduces fairness gaps across automatically learned image groups for glaucoma screening with vision-language models, without using demographic labels. If it works, it offers a way to mitigate demographic bias in medical AI when protected attributes are unavailable, a common practical constraint.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cluster-to-demographic validity is the load-bearing premise; the abstract never validates that inferred clusters correspond to protected attributes, so parity across clusters may not constitute demographic debiasing.","rationale":"The reader's weakest assumption identifies exactly the cluster-to-demographic mapping as the load-bearing premise, and I agree. The abstract explicitly says fairness is measured 'across inferred demographic subgroups,' but no validation of that inference is visible. Since the provided full text is unreadable corruption, I cannot check whether the actual experiments validate clusters against ground-truth protected attributes. The concern is thus not resolved but also not confirmed fatal. The appropriate verdict remains UNVERDICTED, consistent with the reader's assessment. A definitive evaluation requires a readable manuscript and a targeted check of cluster-demographic alignment. The paper's method is plausible and the FairVLMed dataset provides the necessary attributes for validation, so the concern is empirically testable rather than a matter of internal inconsistency. I find no independent evidence (code, proofs, or readable results) that would alter the verdict.","tokens_in":12553,"tokens_out":3296,"duration_ms":37619,"concrete_test":"Obtain a readable version of the paper, specifically the experimental section (likely §4) and any appendix. Check two things: (1) Do the authors report quantitative alignment between inferred clusters and ground-truth protected attributes from FairVLMed (e.g., adjusted Rand index, normalized mutual information, or cluster composition tables by race/ethnicity/sex)? (2) Are fairness metrics (EOD, ES AUC, Groupwise AUC) also computed for the true demographic subgroups rather than only for the inferred clusters? If either is absent, the central fairness claim is unsupported: parity over clusters cannot be interpreted as demographic debiasing. If both are present and show non-trivial alignment, the concern is resolved. Additionally, if feasible, re-run the method with clusters initialized from known non-demographic factors (e.g., acquisition site) to see whether the reported disparity reducti","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the attribute-agnostic pipeline reduces subgroup disparities in VLM glaucoma screening. The load-bearing premise is that unsupervised clusters of image-image embeddings (step i) correspond to the demographic subgroups that matter for fairness. Harvard FairVLMed contains ground-truth protected attributes, yet the abstract reports EOD/ES AUC/Groupwise AUC only 'across inferred demographic subgroups.' If clusters instead capture imaging device, acquisition site, disease severity, or retinal morphology, then equalizing performance across clusters does not constitute demographic debiasing; it may even optimize a proxy that diverges from demographic parity. No evidence of cluster-demographic alignment (e.g., adjusted Rand index, cluster-by-attribute contingency, or evaluation on true demographic subgroups) is visible in the abstract, and the full text provided is corrupted to the point that no experimental validation can be inspected. The fairness metrics are measured on the same clusters the method optimizes, which risks circularity: the method reduces disparities over the very partition it constructs. This does not establish reduced demographic bias. A secondary structural concern is the top-k upweighting's accuracy trade-off, but the primary issue is proxy validity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an attribute-agnostic debiasing method for vision-language models (VLMs) in glaucoma screening from retinal fundus images. The pipeline is described in three steps: (i) infer 'proxy subgroups' by unsupervised clustering of image-image embeddings; (ii) compute gradient-similarity weights between the CLIP-style multimodal loss and a SimCLR-style image-pair contrastive loss; and (iii) apply these weights in a joint, top-k weighted objective that upweights underperforming clusters. The method is evaluated on the Harvard FairVLMed glaucoma subset, with the abstract reporting reductions in Equalized Odds Distance (EOD) and gains in Equalized Subgroup AUC (ES AUC) and Groupwise AUC across inferred demographic subgroups.","tokens_in":12737,"tokens_out":4701,"duration_ms":50451,"significance":"If the central claim is correct, the method would address a real and important problem: mitigating demographic bias in medical imaging when protected attribute labels are unavailable. The idea of using gradient-similarity between multimodal and contrastive losses to guide reweighting, combined with cluster-based proxy subgroups, is technically interesting. However, the manuscript as submitted provides no quantitative results in the abstract, no code release, and the full text is corrupted to the point of unreadability. The empirical claims are therefore not currently falsifiable or reproducible. In addition, the evaluation is partially circular (fairness metrics are computed over the same clusters the method optimizes), and the central premise that clusters correspond to demographic subgroups is never validated. The significance of the work is conditional on resolving these issues with concrete evidence.","major_comments":[{"comment":"The abstract states that the method 'reduces subgroup disparities' and lists EOD, ES AUC, and Groupwise AUC as evaluation metrics, but reports no numerical values. The provided full text is corrupted mojibake and includes an unrelated identifier 'arXiv:2508.09083v2 [physics.optics]', making it impossible to inspect any equations, tables, or experimental settings. Consequently, the central empirical claim cannot be verified. A clean, readable manuscript with quantitative results is essential.","section":"Abstract and Full Text"},{"comment":"The fairness metrics are computed over the same inferred clusters that the top-k weighted objective (step iii) is designed to optimize. Because the reweighting explicitly upweights underperforming clusters, improving disparity metrics on that same partition is expected by construction. This circularity means the reported fairness gains do not establish demographic debiasing. The authors must also evaluate on ground-truth protected attributes (available in Harvard FairVLMed) and report cluster-to-attribute alignment (e.g., adjusted Rand index, normalized mutual information, or a cluster-by-attribute contingency table).","section":"Evaluation metrics (EOD, ES AUC, Groupwise AUC)"},{"comment":"The load-bearing assumption is that unsupervised clusters of image embeddings are faithful proxies for demographic subgroups. The abstract never validates this. If the clusters capture imaging device, acquisition site, disease severity, or retinal morphology, then parity across clusters does not constitute reduced demographic bias; it may even optimize a proxy that diverges from demographic parity. Concrete tests are needed: report cluster-by-attribute contingency, ARI/NMI against known attributes, and evaluate EOD and Groupwise AUC on the true demographic labels. Without this, the central fairness claim collapses.","section":"Step (i): proxy subgroups via clustering"},{"comment":"The top-k reweighting may reduce group disparities at the expense of overall accuracy. The manuscript should report overall AUC or accuracy for the joint objective and compare it with the unweighted baseline, to show that any accuracy loss is acceptable. The abstract mentions only fairness metrics, and the corrupted full text does not allow checking this trade-off. This is a secondary but load-bearing issue for practical deployment.","section":"Overall accuracy trade-off"}],"minor_comments":[{"comment":"The full text is severely corrupted and unreadable; please resubmit a clean, properly encoded PDF or source file.","section":"Manuscript formatting"},{"comment":"The text contains an unrelated arXiv identifier 'arXiv:2508.09083v2 [physics.optics]' near the beginning; this appears to be a submission or rendering error and should be removed.","section":"Inserted passage"},{"comment":"The abstract uses 'inferred demographic subgroups' without explaining how clusters are interpreted as demographic. The authors should explicitly define the relationship between clusters and demographics, or use a more neutral term such as 'proxy subgroups' consistently.","section":"Terminology"},{"comment":"The phrase 'without protected attribute supervision' is potentially misleading: the method does use unsupervised clustering of image embeddings, which may indirectly encode demographic information. Please clarify what information the clustering is allowed to use and what exactly is 'attribute-agnostic'.","section":"Claim wording"}],"recommendation":"major_revision","confidential_remarks":"The submitted manuscript is physically corrupted, and the unrelated arXiv ID suggests a possible submission mix-up or rendering error. I recommend requiring a clean copy before further review. Beyond that, the circularity of evaluating fairness on the same clusters the method optimizes, and the unvalidated cluster-to-demographic mapping, are substantive issues that require new analyses. These are fixable within the manuscript's scope, hence major revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper proposes an attribute-agnostic debiasing method for VLM glaucoma screening: unsupervised clustering of image embeddings to create proxy subgroups, gradient-similarity weighting between CLIP and SimCLR losses, and a top-k reweighted joint objective that upweights underperforming clusters. If it works, it is a useful practical recipe for fairness when protected attributes are unavailable. The authors are honest that it builds on a reweighting-based contrastive framework, and the abstract is clear. The specific combination—unsupervised proxy subgroups plus gradient-similarity weighting between multimodal and image-pair losses—is a plausible extension, and the application to glaucoma screening is relevant. The evaluation metrics (EOD, ES AUC, groupwise AUC) are standard and appropriate.\n\nThe soft spot is structural and you should take it seriously. The whole fairness claim rests on the premise that unsupervised clusters are faithful proxies for demographic subgroups. Harvard FairVLMed contains ground-truth protected attributes, yet the abstract reports fairness \"across inferred demographic subgroups\" and never shows the clusters correspond to those attributes. If the clusters capture acquisition site, device, or disease severity, parity across clusters does not mean reduced demographic bias. The fairness metrics are also computed over the same clusters the optimizer upweights, so there is circularity: the method reduces disparities over the partition it constructs, and the demographic conclusion does not follow. This is fixable: cluster-by-attribute contingency, ARI, or evaluation on true demographic subgroups would settle it.\n\nI could not verify anything else. The full text I have is a corrupted extraction—it even includes a header from a different arXiv paper (physics.optics)—so there are no equations, tables, error bars, or ablations to check. That may be an extraction artifact rather than the authors' fault, but it means the abstract's numbers are unverifiable at this stage. On the evidence available, the paper is plausible but unverified.\n\nWho is this for? Researchers working on fairness in medical imaging, especially those dealing with missing protected attributes. The paper deserves a serious referee if the actual manuscript is readable, because the question is important and the method is coherent. But I would not cite it in its current form until the proxy validity and the numbers are confirmed.\n\nRecommendation: send it to peer review with the expectation of major revision; the authors must add cluster-demographic validation and report results on true subgroups.","headline":"Clever label-free debiasing recipe, but the fairness claim rests on an unvalidated cluster-demographic mapping.","tokens_in":13329,"tokens_out":2538,"would_cite":false,"duration_ms":29716,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an attribute-agnostic debiasing pipeline—unsupervised clustering plus gradient-similarity reweighting of a joint contrastive objective—reduces subgroup disparities in vision-language-model glaucoma screening without p","keywords":["vision-language models","glaucoma screening","fairness","contrastive learning","attribute-agnostic debiasing","unsupervised clustering","retinal fundus images","equalized odds"],"falsifier":"Take a retinal-image dataset with ground-truth demographic labels, run the clustering and reweighting pipeline, and compare the inferred clusters to the true groups. If the cluster assignments are nearly random with respect to demographics, or if Equalized Odds Distance computed with true labels does not improve even though cluster-based Equalized Odds Distance does, the central claim is falsified.","tokens_in":12355,"feed_emoji":"👁️","tokens_out":6593,"duration_ms":63704,"temperature":0.7,"pith_summary":"The paper tries to show that vision-language models for glaucoma screening from retinal fundus images can be made fairer without ever observing protected attributes. The proposed pipeline clusters image embeddings into proxy subgroups, measures how much each cluster's loss gradient agrees with a contrastive image-pair loss, and reweights the joint training objective toward the top-k hardest clusters. On a public medical fairness benchmark, this is reported to lower Equalized Odds Distance and raise Equalized Subgroup AUC and Groupwise AUC across the inferred subgroups. If right, it would let practitioners debias vision-language models in settings where collecting race, sex, or ethnicity labels is legally or practically impossible. The cost is that fairness is defined with respect to clusters, not true demographic groups.","feed_headline":"Clusters, not demographics, equalize glaucoma AI screening","feed_subtitle":"Unsupervised cluster reweighting cuts subgroup gaps in vision-language glaucoma screening without protected labels.","key_machinery":"The load-bearing mechanism is a joint contrastive objective with per-cluster reweighting. The two losses—an image-text alignment loss and an image-pair contrastive loss—define learning signals for the same encoder; the cosine similarity between their gradients tells how much a cluster would be served by extra contrastive pressure. Clusters are found by unsupervised clustering of image embeddings, so no protected attributes are needed. The top-k weighting concentrates training on the clusters whose gradients are least aligned, which the paper identifies as the underperforming subgroups.","core_discovery":"The central claim is that demographic disparities in a vision-language model trained for glaucoma detection from retinal fundus images can be reduced without any protected attribute labels. The method (i) clusters image embeddings into proxy subgroups; (ii) for each cluster, computes gradient-similarity weights between the image-text alignment loss and the image-pair contrastive loss; and (iii) optimizes a joint, top-k weighted objective that upweights clusters whose gradients disagree most, targeting the hardest examples. On a public fairness benchmark for medical vision-language models, the paper reports lower Equalized Odds Distance, higher Equalized Subgroup AUC, and higher Groupwise AUC","pith_inferences":["The central unvalidated premise is that clusters correspond to demographic groups; if they capture acquisition site, camera artifact, or disease severity, the fairness claim overstates real-world demographic parity. Testing on data with known demographic labels would resolve this.","Because top-k selection is unsupervised, it may also emphasize clusters that are hard because of label noise or rare pathology; ablating top-k against uniform cluster weighting would show whether hard-cluster selection or simple reweighting drives the gain.","The same weighting logic could be extended to other pairs of losses (e.g., supervised and self-supervised) as a general fairness-through-gradient-agreement tool."],"forward_implications":["Glaucoma screening models built on vision-language models can be debiased without collecting sensitive demographic data, easing privacy and regulatory constraints.","Fairness metrics can be computed and tracked during training even when protected attributes are absent, using the inferred clusters as stand-ins.","The gradient-similarity, top-k reweighting recipe is not specific to glaucoma: it can be applied to any vision-language contrastive training pipeline.","The reported gains (lower Equalized Odds Distance, higher subgroup AUC) would mean the worst-off inferred groups improve without requiring a separate debiasing stage."],"supporting_citations":[],"fun_headline_variants":["Cluster-based weights fix glaucoma AI bias sans protected data","Unsupervised reweighting narrows subgroup gaps in glaucoma AI","No demographics needed: clusters debias glaucoma models","Cluster upweighting targets hardest glaucoma cases for fairness"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The method only removes bias if the unsupervised clusters of image embeddings are faithful stand-ins for the demographic subgroups that matter; if the clusters instead reflect imaging artifacts or disease severity, parity across clusters is not demographic fairness.","fun_headline_variants_meta":{"raw":{"variants":["Cluster-based weights fix glaucoma AI bias sans protected data","Unsupervised reweighting narrows subgroup gaps in glaucoma AI","No demographics needed: clusters debias glaucoma models","Cluster upweighting targets hardest glaucoma cases for fairness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000669,"raw_usage":{"total_tokens":2872,"prompt_tokens":715,"completion_tokens":2157,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":2092}},"tokens_in":459,"tokens_out":2157,"duration_ms":15819,"temperature":1.0,"reasoning_tokens":2092,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:13:25.862245+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a retinal-image dataset with ground-truth demographic labels, run the clustering and reweighting pipeline, and compare the inferred clusters to the true groups. If the cluster assignments are nearly random with respect to demographics, or if Equalized Odds Distance computed with true labels does not improve even though cluster-based Equalized Odds Distance does, the central claim is falsified.","supporting_citations":[],"review_version":1}