{"id":"5ee3e8dc-5ced-4ff7-b041-5d97c8dacc7d","arxiv_id":"2508.12857","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"REACH is claimed to improve task completion by up to 17%, double high-priority success, and cut bandwidth penalties by over 80% in simulations of community GPU scheduling.","lead":"A scheduling framework called REACH applies Transformer-based reinforcement learning to allocate AI tasks across community GPU platforms. The abstract reports large simulation gains, but the paper's full text is actually a different manuscript, so these claims cannot be verified here.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full text is a different paper (E3RG), so REACH's claimed results have no supporting evidence in the manuscript.","rationale":"The reader's weakest_assumption concerned whether the simulation environment faithfully represents real community GPU platforms. That assumes the full text actually contains REACH's simulation setup and experiments. But the full text provided is a completely different paper on multimodal empathetic response generation (E3RG). Therefore the most load-bearing issue is not a subtle modeling assumption but the total absence of any REACH content: no method, no environment, no baselines, no results. The central claim is unsupported by the manuscript as submitted. This is a decisive manuscript integrity problem, not a scientific uncertainty that further experiments could resolve within the current text. I do not question the authors' integrity; I note an objective mismatch between the abstract and the body. The appropriate verdict is REJECT because the submission, in its current form, cannot support its stated contribution. This differs from the reader's UNVERDICTED verdict: we have enough information to know the evidence is missing, whereas UNVERDICTED suggests we simply cannot decide. The mismatch makes acceptance impossible without the actual REACH paper.","tokens_in":1809,"tokens_out":3479,"duration_ms":34610,"concrete_test":"Retrieve the arXiv record for ID 2508.12857 and inspect the actual full text. If it is the E3RG paper (as provided here), then the REACH abstract is not attached to its own manuscript and the claimed numbers have no evidentiary basis. Additionally, search for any separate REACH manuscript by the same or different authors; if found, run its reported experiments with released code and compare the claimed 17% improvement and >80% bandwidth penalty reduction against the stated baselines.","verdict_should_be":"REJECT","load_bearing_attack":"The abstract presents REACH as a Transformer-based RL scheduler that improves task completion by up to 17%, more than doubles high-priority success, and reduces bandwidth penalties by over 80% in extensive simulations. However, the supplied full text is an unrelated ACM MM 2025 paper titled 'E3RG: Building Explicit Emotion-driven Empathetic Response Generation System' by Lin et al. It contains no mention of REACH, community GPU platforms, scheduling, simulation environments, baselines, or stress tests. Consequently, every component of the central claim is unsupported by the manuscript as submitted. This is not a question of simulation fidelity or methodological detail; there is no simulation or method described at all. The mismatch between abstract and body makes the reported results unverifiable. Unless the actual REACH manuscript is supplied—with environment specifications, hyperparameters, baselines, and code—the central claim cannot be evaluated. The E3RG paper may itself be a legitimate contribution, but it does not address REACH.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission consists of an abstract describing REACH, a Transformer-based reinforcement learning scheduler for community GPU platforms, alongside a full text that is entirely a different paper titled 'E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model' by Lin et al. The abstract claims that REACH improves task completion rates by up to 17%, more than doubles high-priority task success, and reduces bandwidth penalties by over 80% in extensive simulations, with robustness to GPU churn and network congestion. The full text instead discusses multimodal empathetic response generation and an ACM MM 2025 competition entry, containing no mention of REACH, scheduling, community GPU platforms, simulations, baselines, or any related technical content.","tokens_in":1926,"tokens_out":1562,"duration_ms":16572,"significance":"If the REACH results were presented with adequate methodology, the claims of substantial gains in task completion, high-priority success, and bandwidth efficiency would be of practical interest to the community GPU scheduling literature. However, the manuscript as submitted provides no evidence for these claims: the abstract is not supported by any corresponding body text, methodology, experiments, or comparisons. The E3RG paper may itself be a valid contribution to multimodal empathetic response generation, but it is irrelevant to the REACH claims. The novelty and significance of REACH cannot be assessed from the submitted material.","major_comments":[{"comment":"The full text of the submission is an unrelated paper titled 'E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model' (ACM MM 2025), which contains no reference to REACH, community GPU platforms, task scheduling, simulation environments, or any of the quantitative results in the abstract. The central claims of the abstract are therefore entirely unsupported by the manuscript as submitted.","section":"Abstract vs. full text"},{"comment":"No description of the REACH framework, its Transformer-based sequence scoring formulation, state representation, reward design, training procedure, or hyperparameters appears anywhere in the submitted text. Similarly, the simulation environment, baseline schedulers, evaluation metrics, and stress-test protocols referenced in the abstract are all absent. Without these components, the reported improvements of 17% task completion, doubled high-priority success, and 80% bandwidth penalty reduction cannot be verified or reproduced.","section":"Methodology"},{"comment":"The abstract reports 'extensive simulation results' and 'stress tests' but the submitted document contains no figures, tables, error bars, statistical analyses, or code artifacts for these experiments. The complete detachment between the abstract and the body text means that the empirical claims are unverifiable, and no amount of revision to the current text can remedy this without essentially rewriting the paper.","section":"Results and reproducibility"}],"minor_comments":[{"comment":"The arXiv identifier in the header footer (arXiv:2508.12854v1 [cs.AI]) differs from the abstract's stated identifier (arXiv:2508.12857), which is confusing and may indicate a compilation or submission error; this should be resolved if a corrected manuscript is submitted.","section":"Manuscript metadata"},{"comment":"The title, keywords, and abstract of the manuscript describe REACH, while the body reports E3RG with different keywords and CCS concepts; this inconsistency should be corrected by submitting the intended REACH manuscript instead of the current mismatched content.","section":"Title and scope"}],"recommendation":"reject","confidential_remarks":"The submitted file appears to contain a different paper than the one described in the abstract. This is not a minor fixable issue but a fundamental mismatch that makes the manuscript unpublishable in its current form. I would suggest the editor check whether this is a submission workflow error, but based on the present content, the REACH claims are entirely unsubstantiated and the manuscript cannot be evaluated for soundness. I also note that the E3RG paper is not the subject of this submission and should not be counted toward any assessment of the REACH contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe headline: this arXiv submission is not the paper it claims to be. The abstract describes REACH, a Transformer-based reinforcement learning scheduler for community GPU platforms, and reports large simulation gains—17% better task completion, doubled high-priority success, over 80% reduction in bandwidth penalties. The full text, however, is an unrelated ACM Multimedia 2025 paper titled E3RG, about multimodal empathetic response generation. Different title, different authors, different subject. There is no REACH method, no simulation, no baselines, no code in the manuscript. Nothing in the body supports a single sentence of the abstract.\n\nI don't want to be theatrical: this is most likely a submission mix-up. The E3RG paper itself looks like a legitimate piece of work—it claims a top-1 spot in the Avatar-based Multimodal Empathy Challenge and ships code on GitHub. But that doesn't help. For REACH, the submitted manuscript contains only an abstract appended to someone else's paper. The reader's unverified status and low confidence are exactly right; the stress-test note lands squarely.\n\nWhat is genuinely new? Hard to say. The abstract's idea—framing scheduling as sequence scoring, letting a Transformer policy balance performance, reliability, cost, and network efficiency, and co-locating compute with data—is a sensible direction for community GPU pools with churn and heterogeneity. But there is no methodology here to inspect. No environment description, no baseline definitions, no error bars, no ablation. The claims are pure numbers backed by no text.\n\nThe soft spot is not a subtle flaw in the analysis. It is the absence of the analysis. The mismatch between abstract and body is load-bearing and total. I cannot review a different paper's body and pretend it supports REACH's conclusions. If the authors meant to submit the real REACH manuscript, they should resubmit it; then it can be evaluated on its merits. As this version stands, a serious referee has nothing to read.\n\nMy recommendation: desk-reject this version, or at minimum return it to the authors for a valid full text. Do not send it to peer review until the actual REACH paper is present. The E3RG paper should be cited via its own ACM publication if you're interested in that thread; it doesn't belong in a cs.NI review loop.","headline":"This arXiv listing's abstract describes REACH, a GPU-scheduling RL system, but the full text is an unrelated empathetic-response paper—so there is no supporting content for any of REACH's claims.","tokens_in":2367,"tokens_out":2482,"would_cite":false,"duration_ms":23883,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"REACH is a Transformer-based reinforcement learning scheduler that claims to improve task completion rates on community GPU platforms by up to 17%, more than double the success rate for high-priority tasks, and reduce bandwidth penalties…","keywords":["reinforcement learning","task scheduling","community GPU","heterogeneous networks","Transformer","sequence scoring","bandwidth penalty","resource reliability"],"falsifier":"Deploy REACH and a strong baseline scheduler on a live community GPU testbed with dozens of heterogeneous consumer GPUs across multiple network domains, run the same workload mix for a week, and compare task completion rates and high-priority success rates; if REACH fails to improve upon the baseline or to more than double high-priority success, the central claim is falsified.","tokens_in":1650,"feed_emoji":"⚙️","tokens_out":4072,"duration_ms":40080,"temperature":0.7,"pith_summary":"Community GPU platforms pool idle consumer GPUs from diverse, volatile, and network-constrained environments, but traditional schedulers struggle with such heterogeneity. REACH addresses this by recasting task scheduling as a sequence scoring problem and learning a Transformer-based reinforcement learning policy. The paper claims that, in extensive simulations, REACH outperforms state-of-the-art baselines by up to 17% in task completion rate, more than doubles the success rate for high-priority tasks, and reduces bandwidth penalties by over 80%. A sympathetic reader would care because an effective scheduler could make community GPU platforms a practical, cost-efficient alternative to centralized clusters for AI workloads.","feed_headline":"RL scheduler boosts community GPU completion by 17%","feed_subtitle":"Transformer-based scheduling also doubles high-priority task success and cuts bandwidth penalties by 80%.","key_machinery":"The central object is the Transformer-based reinforcement learning policy that scores candidate task-to-GPU allocation sequences. The scheduling problem is cast as sequence scoring: given the current global GPU states, task requirements, and network topology, the model produces a score for every feasible assignment sequence and selects the one with the highest score. This single learned decision mechanism incorporates data-computation co-location, priority-aware task handling, and mitigation of unreliable resources, replacing the multi-objective hand-tuning of traditional schedulers.","core_discovery":"The central claim is that scheduling on community GPU platforms can be effectively solved by a Transformer-based reinforcement learning agent that views task allocation as a sequence scoring problem. Instead of assigning tasks with hand-crafted heuristics, REACH learns a policy that maps the global GPU state, task requirements, and network conditions to a score for each candidate allocation sequence, then selects the highest-scoring sequence. This learned policy balances performance, reliability, cost, and network efficiency, and it adaptively co-locates computation with data, prioritizes critical jobs, and avoids unreliable resources. In simulation, REACH achieves up to 17% higher task completion rates, more than doubles the success rate for high-priority tasks, and cuts bandwidth penalties by over 80% relative to baseline schedulers, with stress tests showing resilience to GPU churn and network congestion and scalability experiments confirming gains in large-scale, high-contention settings.","pith_inferences":["Because the reported gains come from simulation, a real-world deployment on a live community GPU testbed would be the natural next test; the policy's ability to transfer to unseen hardware and network behavior remains open.","REACH's sequence scoring could also be used as a re-ranking layer on top of existing heuristic schedulers, reducing adoption friction while still yielding improvements.","The paper's balancing of performance, reliability, cost, and network efficiency suggests a tunable scheduler could let users specify preferences (e.g., cheapest-first vs. fastest-first), an extension not explicitly explored.","Ablations that isolate the effects of data co-location versus reliability avoidance would clarify which component drives the measured gains and help generalize the approach to other platforms."],"forward_implications":["Existing schedulers for geo-distributed or volunteer GPU platforms could be replaced or augmented with learned sequence-scoring policies that adapt to heterogeneity and volatility.","Reliable scheduling would make community GPU platforms attractive for cost-sensitive AI training and inference, broadening access to computational resources.","The formulation could extend to other resource-constrained distributed systems such as edge computing, federated learning, or volunteer computing, where data location and resource reliability vary.","Reducing bandwidth penalties through learned co-location could lower the communication bottleneck in distributed training across wide-area networks.","The learned policy could be retrained or fine-tuned on new platform profiles without redesigning the scheduler, easing deployment across different community networks."],"supporting_citations":[],"fun_headline_variants":["RL scheduler boosts community GPU completion by 17%","Transformer RL scheduler doubles high-priority task success","RL scheduler cuts bandwidth penalties by 80% on shared GPUs","REACH: RL scheduler for community GPU pools lifts completion 17%","RL scheduler for heterogeneous GPU networks improves completion 17%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulation environment faithfully represents the real-world diversity, volatility, and network conditions of community GPU platforms; if real platforms behave differently, the reported gains may not reproduce in deployment.","fun_headline_variants_meta":{"raw":{"variants":["RL scheduler boosts community GPU completion by 17%","Transformer RL scheduler doubles high-priority task success","RL scheduler cuts bandwidth penalties by 80% on shared GPUs","REACH: RL scheduler for community GPU pools lifts completion 17%","RL scheduler for heterogeneous GPU networks improves completion 17%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1587,"prompt_tokens":913,"completion_tokens":674,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":591}},"tokens_in":529,"tokens_out":674,"duration_ms":6218,"temperature":1.0,"reasoning_tokens":591,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:16:58.133742+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deploy REACH and a strong baseline scheduler on a live community GPU testbed with dozens of heterogeneous consumer GPUs across multiple network domains, run the same workload mix for a week, and compare task completion rates and high-priority success rates; if REACH fails to improve upon the baseline or to more than double high-priority success, the central claim is falsified.","supporting_citations":[],"review_version":2}