{"id":"be746017-6aec-487d-b694-bb2c60815e15","arxiv_id":"2508.11170","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Integer-only labels with masked loss are claimed to improve VLM video quality assessment, ranking 3rd in VQualA 2025.","lead":"This paper proposes a fine-tuning method for vision language models that forces video quality scores to be whole numbers between 10 and 50, and computes loss only on the first two digits of the integer label. The claim is that this integer-only approach improves accuracy and consistency in video quality assessment, supported by a 3rd place finish in the VQualA 2025 challenge.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is unverified: no control ablation, corrupted full text, and target-mask specification ambiguous.","rationale":"The reader's UNVERDICTED verdict is appropriate: the evidence in the abstract and corrupted full text is insufficient. I agree with the reader's weakest assumption. I considered whether there is any independent support (e.g., reproducible code, machine-checked proofs) that could offset the missing evidence; the manuscript includes none in the accessible portion. I also considered whether the target-mask might be internally inconsistent; it appears ambiguous, but the corruption prevents a definitive internal-contradiction finding. Thus the load-bearing concern is evidential, not logical: the causal claim rests on a single leaderboard position and an underspecified mechanism. No verdict change is warranted beyond UNVERDICTED; if the clean paper confirms the ablation is present and significant, the claim may become credible.","tokens_in":10697,"tokens_out":4113,"duration_ms":48010,"concrete_test":"Obtain a clean version of the paper and run an ablation on the VQualA Track I validation set with Qwen2.5-VL: (a) standard fine-tuning with original decimal MOS labels and full loss; (b) integer labels with the proposed first-two-digit mask; (c) integer labels without mask; (d) decimal labels with mask. Fix dataset order, base model, optimizer, and number of steps across conditions; report Pearson/Spearman correlation and prediction consistency over at least three seeds. The central claim is supported only if (b) significantly outperforms both (a) and (c) by more than the seed variance; otherwise the integer label or the mask alone, or neither, is responsible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that the proposed integer-only label construction and target-mask loss cause the reported VQA improvement—is unsupported by the evidence accessible in the manuscript. The abstract provides only a single third-place challenge ranking with no numerical comparison to a baseline fine-tune of Qwen2.5-VL without the proposed components; ranking position is confounded by dataset, training schedule, hyperparameters, and base-model choice. The target-mask strategy is under-specified: since labels are already integers in [10,50], 'masking all but the first two-digit integer' appears to be either a no-op or an undefined token-level mask, and no ablation or equation is legible to clarify it. The supplied full text is severely corrupted (many passages are mojibake, and the running header cites arXiv:2508.11176v1 rather than the paper's 2508.11170), so the loss definition, experimental setup, and results cannot be checked. Consequently the causal attribution—that IOVQA's specific loss design improves accuracy/consistency—remains an assertion, not a demonstrated result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes IOVQA, a fine-tuning method for vision-language models in video quality assessment. The claimed innovations are (i) converting decimal MOS labels to integers in the range [10,50] and constraining the model output to that integer range, and (ii) a target-mask loss that is said to unmask only the first two-digit integer of the label, forcing the model to focus on critical parts of the numerical evaluation. The method is applied to Qwen2.5-VL and the paper reports that it achieves 3rd place in the VQualA 2025 GenAI-Bench AIGC Video Quality Assessment Challenge Track I. The abstract asserts significant improvement in accuracy and consistency, but provides no quantitative results, no baseline comparison, no ablations, and no statistical analysis. The supplied full text is heavily corrupted, so the loss definition and experimental tables cannot be verified.","tokens_in":10952,"tokens_out":4135,"duration_ms":49796,"significance":"If the central claim were substantiated, the paper would offer a simple, computationally cheap fine-tuning trick for scalar-output VQA—integer-only labels and a masked loss—with potential value for video quality and aesthetic assessment. The use of an external challenge leaderboard is a strength insofar as it provides an independent evaluation setting, and the proposed design is concrete and falsifiable. However, the significance is currently unestablished: the evidence in the manuscript is limited to a single leaderboard rank, which is confounded by dataset, schedule, and base-model choices, and the target-mask operation appears possibly vacuous given the integer label range. The paper therefore needs substantial additional evidence before its claims can be assessed.","major_comments":[{"comment":"The central claim that the method 'significantly improves' accuracy and consistency is not supported by any reported numerical result. The only evidence cited is a 3rd-place rank in an external challenge, with no comparison to a baseline fine-tune of Qwen2.5-VL, no metric values, no error bars, and no statistical tests. A single leaderboard position is confounded by dataset composition, training schedule, hyperparameters, and base-model initialization, so it cannot isolate the contribution of the proposed components. A direct controlled comparison (same data and schedule, with and without the integer-label and target-mask design) is required.","section":"Abstract and §4 (Experimental results)"},{"comment":"The masking rule is under-specified and possibly a no-op. Labels are converted to integers in [10,50]; every integer in that range has exactly two decimal digits. 'Only the first two-digit-integer of the label is unmasked' therefore either leaves the entire label unmasked or requires a token-level mask definition that is not given in the legible portions of the text. Please state precisely which tokens/positions are masked in the loss and provide an ablation that isolates the mask from the integer-label construction; otherwise the claimed novelty of the target mask cannot be evaluated.","section":"§3, target-mask strategy"},{"comment":"The choice of the integer range [10,50] and the two-digit mask are free parameters. No sensitivity analysis or ablation is provided to justify these choices or to show they are responsible for any improvement. In particular, the paper does not compare integer labels with decimal labels, nor does it compare masked and unmasked losses. Without such evidence, the observed ranking could be driven by unrelated fine-tuning details rather than by the proposed design.","section":"§3, label construction and loss"},{"comment":"The supplied manuscript is heavily corrupted: many passages are mojibake, and the running header cites arXiv:2508.11176v1 rather than arXiv:2508.11170. The loss equations, dataset description, training details, and result tables are not legible enough to be checked. This makes verification impossible in the current form. The authors should provide a clean, complete version before the technical content can be reviewed.","section":"Full text, equations and results"}],"minor_comments":[{"comment":"The running header of the supplied PDF/plain text cites a different arXiv identifier (2508.11176v1) than the paper under review (2508.11170). This should be corrected.","section":"Header"},{"comment":"The abstract claims the contribution is 'merely leaving integer labels during fine-tuning,' but the method also introduces a target-mask loss; the phrasing is inconsistent and should be clarified.","section":"Abstract"},{"comment":"The notation around 'Overall_MOS', 'MOS', and 'two-digit-integer' is inconsistent. Please define all symbols and use a consistent hyphenation and capitalization.","section":"Notation"},{"comment":"Provide a citable reference or URL for the VQualA 2025 challenge leaderboard, including the date of access, so the claimed 3rd-place result can be independently verified.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is currently closer to a challenge report than a self-contained technical contribution. The most urgent issue is the absence of any direct experimental evidence; I would recommend asking for a clean manuscript with full equations, tables, and at least one controlled ablation of the two proposed components. The corrupted text makes it impossible to tell whether the authors already have such results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You need to know one thing upfront: the version I have is unreadable. The full text is mostly mojibake, the running header cites a different arXiv ID, and the equations are garbage. So anything I say about the method is based on the abstract alone. The abstract may be the only reliable content here.\n\nWhat's actually worth credit: the idea is simple and not silly. Constraining the output to integer MOS labels in [10,50] and computing loss only on the first two digit positions is a cheap, plausible fine-tuning trick for VQA with VLMs. The challenge result (3rd in VQualA 2025 Track I) is a real external data point, though a single ranking with no numbers attached tells us almost nothing.\n\nNow the soft spots. First, there are no quantitative results in the abstract—no MOS/PLCC/SROCC, no comparison to a baseline fine-tune of Qwen2.5-VL, no ablations. The causal claim \"integer-only loss causes the improvement\" is unsupported by what's visible; the ranking is confounded by dataset construction, schedule, and hyperparameters. Second, the target-mask strategy is confusing as described. If labels are already integers in [10,50], then \"masking all but the first two-digit integer\" reads like either a no-op or an undefined token-level mask. That needs a clear equation or a worked example. Third, the text corruption is a practical blocker—no referee can check the loss, the training setup, or even the section ordering.\n\nThe stress-test note matches my reading: the central claim is an assertion, not a demonstrated result. I don't think the argument is incoherent—just unverifiable in this artifact. The idea is the kind of thing someone might try in five minutes and find useful, but that's a hunch, not evidence.\n\nWho is this for? Practitioners fine-tuning VLMs for quality scoring. They might extract the gist from the abstract and try it. But as a paper, it needs a clean, complete text and at least one controlled comparison before it deserves referee time.\n\nMy recommendation: desk reject this version, not because the idea is bad, but because the submitted manuscript is unreadable. If the authors post a clean PDF with numbers, then it deserves a genuine peer review.","headline":"A plausible little VQA fine-tuning trick buried in an unreadable PDF; no numbers, no ablations, can't verify anything yet.","tokens_in":11401,"tokens_out":1435,"would_cite":false,"duration_ms":19889,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a fine-tuning method called IOVQA can improve video quality assessment by rounding decimal MOS scores to integers and computing loss only on the first two digits of each label.","keywords":["vision-language models","video quality assessment","MOS regression","integer labels","target-mask loss","supervised fine-tuning","VQualA 2025","Qwen2.5-VL"],"falsifier":"Fine-tune the same base model on the same dataset with the same schedule, but replace integer labels and the two-digit mask with original decimal MOS labels and full-label loss; if this control meets or beats the integer-mask variant on a held-out video quality benchmark, the central claim is falsified.","tokens_in":10632,"feed_emoji":"🎬","tokens_out":4538,"duration_ms":49710,"temperature":0.7,"pith_summary":"The paper tries to show that a simpler labeling scheme makes vision-language models better at scoring video quality. Instead of regressing on decimal mean opinion scores (MOS), it rounds each score to an integer and restricts outputs to the range [10, 50]. When computing the fine-tuning loss, it masks every digit of the label except the first two. Fine-tuning Qwen2.5-VL this way, the method placed third in the VQualA 2025 AIGC video quality assessment challenge, and the authors report improved accuracy and consistency. A sympathetic reader would take the central bet to be that integer-only labels carry enough signal to outperform fine-grained decimal supervision.","feed_headline":"Integer-only labels sharpen video quality scoring","feed_subtitle":"Rounding MOS scores to integers and masking all but two digits improved a vision-language model's video quality ratings.","key_machinery":"The central object is the target-masked integer-only loss: labels are integer-rounded MOS scores in the range [10, 50], and the loss selectively unmask keeps only the first two digits of each label while masking the rest. This forces the gradient to concentrate on the leading digits of the score, while the range constraint keeps outputs numerically stable.","core_discovery":"IOVQA (Integer-only VQA) is a supervised fine-tuning recipe for video quality assessment. During dataset curation, Overall_MOS values are converted from decimals to integers, and the model's answer space is constrained to integers in [10, 50] to avoid numerical instability. At training time a target-mask is applied so the loss sees only the first two digits of the integer label; the rest are masked out. The paper reports that this focuses learning on the critical components of the numerical evaluation. Applied to Qwen2.5-VL, the method improves accuracy and consistency in video quality assessment and achieved 3rd place in Track I of the VQualA 2025 GenAI-Bench AIGC Video Quality Assessment C","pith_inferences":["Editorial inference: rounding and masking act as a form of output regularization, and it would be worth testing whether the benefit persists when the same data and training schedule are kept and only the masking pattern changes.","Editorial inference: the same target-mask idea could be applied to other continuous regression outputs of vision-language models, such as aesthetic scores or relevance ratings, where leading digits carry most of the decision-relevant signal.","Editorial inference: a direct ablation varying mask width (one digit, two digits, full integer, decimal) would separate how much of the gain comes from masking versus integer rounding; the paper's reported comparison does not by itself isolate these two factors."],"forward_implications":["If integer-rounded labels suffice, future video quality datasets need not preserve decimal MOS precision for fine-tuning.","Masking all but the first two digits implies the model learns the coarse score band before fine detail, which may simplify loss design for other quantitative evaluation tasks.","Constraining output to [10, 50] avoids the numerical instability of unbounded regression, so the recipe may transfer to other vision-language models beyond Qwen2.5-VL.","The third-place challenge result is a field-level demonstration that label construction alone can move performance in AIGC video quality assessment."],"supporting_citations":[],"fun_headline_variants":["Integer-only loss sharpens video quality scoring","Rounding MOS to integers boosts VQA accuracy","Masking all but two digits improves video QA","Simple integer trick lifts VQA model to 3rd place","Integer labels: a cheap win for video quality VQA"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The claim collapses if the improvement seen in the challenge comes from dataset composition, training schedule, or the base model rather than from the integer labels and the two-digit target mask, or if rounding decimal MOS scores to integers discards information the model actually needs.","fun_headline_variants_meta":{"raw":{"variants":["Integer-only loss sharpens video quality scoring","Rounding MOS to integers boosts VQA accuracy","Masking all but two digits improves video QA","Simple integer trick lifts VQA model to 3rd place","Integer labels: a cheap win for video quality VQA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1081,"prompt_tokens":796,"completion_tokens":285,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":211}},"tokens_in":540,"tokens_out":285,"duration_ms":3775,"temperature":1.0,"reasoning_tokens":211,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:04:01.516748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fine-tune the same base model on the same dataset with the same schedule, but replace integer labels and the two-digit mask with original decimal MOS labels and full-label loss; if this control meets or beats the integer-mask variant on a held-out video quality benchmark, the central claim is falsified.","supporting_citations":[],"review_version":1}