{"id":"3714e0e0-3b39-4db7-a0c9-1b4cbd09e9f4","arxiv_id":"2508.16148","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"The claimed Japanese PDF QA framework is absent from the supplied body, which is an unrelated ACM paper on social media popularity prediction.","lead":"The submission's abstract claims a framework for Japanese PDF question-answering with vision-language models and a sub-question verification strategy. But the supplied full text is a different paper, about social media popularity prediction, so the claimed results cannot be checked.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplied full text is an unrelated ACM MM paper (arXiv:2508.16147); the abstract's claims about Japanese PDF QA have no derivable support in the submitted body.","rationale":"The reader correctly identified the submission mismatch as the weakest premise. I independently reviewed the supplied full text: the title, author list, task, and arXiv footer identifier are all for a different paper. There is no mention of ten-choice QA, Japanese language scenarios, PDF document layout, Colqwen, semantic verification via sub-questions, or any of the baselines/results claimed in the abstract. Because the abstract's claims are the only formulation of the paper's contribution, and they are completely disconnected from the provided body, no scientific assessment of the framework's correctness or robustness is possible. This is a claim_without_derivation red flag of the strongest kind—the derivation is missing not just a step but the entire paper. The appropriate verdict is UNVERDICTED, pending submission of a matching full text. I have no additional methodological concerns because there is no methodology to scrutinize. I agree with the reader's weakest_assumption exactly; no adjustment to the verdict is needed.","tokens_in":1877,"tokens_out":1969,"duration_ms":16966,"concrete_test":"Query the arXiv API for 2508.16148 and fetch its PDF. If the PDF matches the supplied full text, the abstract's claims are unsupported and the submission is a mismatched artifact. If the PDF instead matches the abstract (hierarchical reasoning, Colqwen, Japanese PDF MCQA), then the supplied full text is corrupted and the actual scientific content should be extracted, uploaded, and re-reviewed. Either outcome settles whether the central claim has any evidential basis in the provided material.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the abstract is that a framework combining hierarchical reasoning, Colqwen retrieval, and sub-question semantic verification improves ten-choice multimodal QA for Japanese PDF documents. For this claim to hold, the submitted full text must contain that framework, its datasets, baselines, and results. It does not. The full text is \"Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction,\" an ACM MM 2025 paper with a different task, different benchmarks, and even a different arXiv identifier in its footer (2508.16147). None of the methods, evaluation protocol, or experimental results described in the abstract appear anywhere in the body. Consequently, the paper's strongest empirical claim is entirely unsupported by any inspectable evidence. This is not a disagreement with consensus or an internal inconsistency; it is an absence of the required artifact. The only way the central claim could be true is if a different, matching manuscript exists, but that manuscript was not provided for review.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract submitted under arXiv:2508.16148 proposes a multimodal hierarchical reasoning framework for ten-choice Japanese PDF document question answering, combining Colqwen-based retrieval, sub-question decomposition, and semantic verification, and claims significant robustness gains over existing MLLMs. The supplied full text, however, is an entirely different paper titled 'Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction,' an ACM MM 2025 submission about social media popularity prediction. The body contains no mention of Japanese PDF QA, Colqwen, ten-choice evaluation, sub-question decomposition, or any datasets, baselines, or experimental results related to the abstract. The first-page footer even carries the identifier arXiv:2508.16147, not 2508.16148. Thus the paper's central empirical claim is unsupported by the submitted manuscript.","tokens_in":1992,"tokens_out":1531,"duration_ms":17681,"significance":"If the claimed framework existed and performed as stated, it could be a meaningful contribution to multilingual document understanding and multimodal QA for Japanese PDFs, particularly in addressing layout complexity and English-centric training bias. However, the submitted manuscript provides no evidence for these claims: there is no method description, no experimental setup, no results, and no ablation study. The paper also ships no code or machine-checked proofs. In its current form, the contribution cannot be evaluated, and the significance is therefore unsubstantiated.","major_comments":[{"comment":"The submitted full text is not the paper described in the abstract. The title, task, method, and benchmarks all differ: the body is a social media popularity prediction paper for ACM MM 2025, and its footer prints arXiv:2508.16147 rather than 2508.16148. The abstract's framework—hierarchical reasoning, Colqwen retrieval, sub-question semantic verification, ten-choice Japanese PDF QA—appears nowhere in the body. Consequently, the central claim of the paper, namely that 'our framework' improves deep semantic parsing and robustness for Japanese PDF documents, has no supporting material in the submitted manuscript. This is not a minor presentation gap; it removes the entire evidential basis for the claimed result.","section":"Full text (title, abstract, §1, first-page footer)"},{"comment":"The abstract asserts that experimental results demonstrate significant enhancement and superior robustness, but no experiments, datasets, baselines, metrics, or results are present in the submitted text. The body's experiments (if any) pertain to social media popularity prediction, a different task, and cannot support the abstract's claims about multimodal multiple-choice QA on Japanese PDFs. This unsupported empirical claim is the paper's central contribution and must be either supplied in full or withdrawn.","section":"Abstract, 'Experimental results demonstrate...'"},{"comment":"The motivation that current MLLMs 'suffer from a strong bias toward English training data, resulting in suboptimal performance for Japanese and other language scenarios' is presented as fact without citation, analysis, or control experiment. Even if the rest of the manuscript matched, this causal claim would need evidence (e.g., cross-lingual benchmark comparisons) to be load-bearing for the proposed method. As submitted, it is an unsupported assertion.","section":"Abstract, 'strong bias toward English training data'"}],"minor_comments":[{"comment":"The footer reads arXiv:2508.16147v1 [cs.IR], but the submission is labeled arXiv:2508.16148. The identifier mismatch is a clear sign of submission error and should be corrected or the correct manuscript provided.","section":"First-page footer"},{"comment":"The title, author affiliation header, CCS concepts, keywords, and ACM reference format all describe the social media popularity prediction paper, not the Japanese PDF QA paper. This inconsistency makes the manuscript impossible to review as submitted.","section":"Title and metadata"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission error: the supplied full text is an unrelated paper with a different arXiv identifier. The abstract's claims have no corresponding body, so there is nothing to review. A desk rejection is appropriate; if the authors intended to submit a different manuscript, they should resubmit the correct full text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me skip the courtesies: the PDF you sent is not the paper whose title and abstract appear on the first page. The abstract describes a framework for Japanese PDF document QA using hierarchical reasoning, Colqwen retrieval, and sub-question verification, and claims strong experimental results. The body is an ACM MM 2025 paper on social media popularity prediction, with a different task, different benchmarks, a partially different author list, and a footer that prints arXiv:2508.16147 rather than 2508.16148. None of the claimed framework, datasets, baselines, or metrics appear anywhere in the body. So the abstract's empirical claims have zero inspectable support.\n\nI want to give credit where it's due. The abstract itself is plausible: the problem of non-English multimodal document QA is real, and combining hierarchical decomposition with retrieval and verification is a reasonable direction. That's about all I can say, because there is no substance to evaluate.\n\nThe soft spot is decisive, not minor. This is not a case where the confident abstract slightly oversells a modest experiment. The submitted artifact is internally inconsistent: the abstract and body are about different problems. The only way the abstract could be true is if a different, matching manuscript exists somewhere, but it was not provided. As a reviewer, there is nothing to adjudicate. You cannot check the math, the data, or the comparisons because none of them are here.\n\nFor peer review, my recommendation is a desk reject, not a send-out. A serious referee cannot work with this. The right next step is for the authors to upload a corrected version with a matching abstract, title, and body, and then the community can assess the actual claim. If that happens, this could be a reasonable empirical paper within the established MLLM-plus-retrieval-plus-prompting paradigm, but that is speculation.","headline":"The abstract and full text are two different papers; the Japanese QA claims have no supporting body, so this cannot be reviewed as submitted.","tokens_in":2594,"tokens_out":2425,"would_cite":false,"duration_ms":23483,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The abstract claims that a multimodal hierarchical reasoning framework with Colqwen-optimized retrieval and sub-question verification improves ten-choice question answering on Japanese PDF documents.","keywords":["multimodal large language models","Japanese PDF documents","ten-choice question answering","hierarchical reasoning","Colqwen retrieval","sub-question decomposition","semantic verification","English training bias"],"falsifier":"Search the body for any section that reports a ten-choice Japanese PDF QA benchmark, a Colqwen-retrieval ablation, or a sub-question verification comparison; the attached body contains none of these, so the abstract's central claim is unsupported by this submission.","tokens_in":1665,"feed_emoji":"📄","tokens_out":6090,"duration_ms":62215,"temperature":0.7,"pith_summary":"The abstract claims that current multimodal large language models handle complex-layout, lengthy PDF documents poorly in ten-choice question answering, and that their English training bias makes Japanese performance worse than English. To fix this, the paper proposes a Japanese PDF understanding framework that combines multimodal hierarchical reasoning with Colqwen-optimized retrieval and a semantic verification strategy that decomposes questions into sub-questions. If the claim is right, robust document understanding should improve for Japanese and other non-English document scenarios without changing the underlying model's language training mix. The supplied full text, however, is a different paper about social media popularity prediction; it contains none of the proposed framework, data, baselines, or experiments described in the abstract. So the abstract stands alone as the source of the central claim, with no supporting material in this submission.","feed_headline":"Abstract promises Japanese PDF QA; full text is social media research","feed_subtitle":"The proposed ten-choice document-QA framework never appears in the attached body; only the abstract supports the claim.","key_machinery":"The central object named in the abstract is a multimodal hierarchical reasoning mechanism paired with Colqwen-optimized retrieval — retrieval tuned with the Colqwen document model — and a semantic verification strategy through sub-question decomposition, where a complex question is split into smaller sub-questions whose answers are checked before final selection. This machinery is supposed to compensate for MLLMs' English-data bias by grounding reasoning in retrieved document evidence and verifying each step. In the attached full text this machinery does not appear; instead the described system uses hierarchical prototypes, contrastive vision-text alignment, and dual-grained prompt learning","core_discovery":"On its own terms, the central discovery is the claim that existing MLLMs fail at ten-choice questions over complex PDF documents, especially Japanese, because they are biased toward English training data; the proposed remedy is a three-part framework — multimodal hierarchical reasoning over document structure, retrieval optimized via a model called Colqwen, and semantic verification through sub-question decomposition. The abstract reports that this framework significantly enhances deep semantic parsing and shows superior robustness in practice. The body of the submitted manuscript, which is titled for a different task, does not contain this framework or any experiments on Japanese PDF QA; th","pith_inferences":["If the abstract's mechanism is separable, the sub-question verification component could be tested alone against plain MLLM prompting on a standard multilingual document QA dataset; that would isolate whether verification or retrieval drives the gain.","The claim that English training bias causes Japanese degradation is testable by comparing model performance on matched Japanese/English PDF sets while controlling for layout; if the gap persists under identical question content, the bias explanation gains support.","A practical consequence the abstract implies: the same framework might be adapted for Chinese, Korean, or Arabic documents without retraining the base model, since the reasoning and verification layers are language-agnostic."],"forward_implications":["If the abstract's claim is correct, ten-choice QA on Japanese PDFs with complex layouts should improve over non-retrieval MLLM baselines.","Sub-question decomposition plus semantic verification should reduce hallucinated choices in long-document QA, since each answer is checked against retrieved evidence.","The approach should transfer to other non-English languages, as the underlying issue is not language fluency but English-centric training distributions.","Retrieval optimized for PDF layout should make deployed systems robust on real user documents, which are messier than standard benchmarks."],"supporting_citations":[],"fun_headline_variants":["Abstract claims PDF QA; body is social media study","Promised framework missing from manuscript body","Ten-choice QA in abstract only, not in text","Multimodal QA paper lacks its own framework","Japanese PDF QA claim unsupported in full text"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the submitted text describes the Japanese PDF QA framework the abstract advertises; in fact the attached full text is a different paper on social media popularity prediction, so the central claim currently has no experimental backing.","fun_headline_variants_meta":{"raw":{"variants":["Abstract claims PDF QA; body is social media study","Promised framework missing from manuscript body","Ten-choice QA in abstract only, not in text","Multimodal QA paper lacks its own framework","Japanese PDF QA claim unsupported in full text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000449,"raw_usage":{"total_tokens":2051,"prompt_tokens":646,"completion_tokens":1405,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":390,"completion_tokens_details":{"reasoning_tokens":1344}},"tokens_in":390,"tokens_out":1405,"duration_ms":11303,"temperature":1.0,"reasoning_tokens":1344,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:28:54.670598+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search the body for any section that reports a ten-choice Japanese PDF QA benchmark, a Colqwen-retrieval ablation, or a sub-question verification comparison; the attached body contains none of these, so the abstract's central claim is unsupported by this submission.","supporting_citations":[],"review_version":1}