{"id":"6a686c89-226b-4911-9259-0acb602f2532","arxiv_id":"2607.12856","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Reinforcement fine-tuning with DP-derived verifiable rewards cuts an open-weight reasoning model's TES emissions from 70.5 to 61.2 kg-CO2, near the 60.8 kg-CO2 optimum.","lead":"Researchers fine-tuned an open-weight reasoning model with reinforcement learning using exact dynamic-programming rewards so it can schedule building thermal storage. On a simple office benchmark the model nearly matches the known optimum, suggesting verifiable rewards can adapt language models to energy control.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"DP-derived dense rewards and DP optimum share the same oracle on a toy plant, so near-optimality is partly by construction; transfer beyond that setting is the untested load-bearing claim.","rationale":"The Reader already isolates the same soft spot: reward and evaluation both come from the same DP oracle on a toy plant, so closeness is partly by construction, and generalization beyond that plant is only lightly probed. Full text is unavailable, so no new internal contradiction or stronger evidence can be checked; the abstract-level circularity remains the load-bearing concern. The concrete test above is a direct, low-cost way to settle whether the near-optimality survives an independent reward signal. Because the Reader already conditions acceptance on higher-fidelity verifiers and artifact release, no verdict shift is warranted; CONDITIONAL with low confidence is the right posture.","tokens_in":2235,"tokens_out":570,"duration_ms":4861,"concrete_test":"Hold out a non-DP reward (e.g., rule-based or MPC-derived dense scores, or a simple physics-based penalty on storage violation + emissions) and re-run the same 30-prompt RFT; if the resulting policy still reaches within ~1 kg-CO2 of the DP optimum on the original plant and retains the reported planning patterns under forecast error, the circularity concern is weakened. If performance collapses or patterns degrade, the headline near-optimality is largely by construction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that RFT with dense rewards converted from exact offline DP action values brings an open-weight reasoner from 70.5 to 61.2 kg-CO2, near the known DP optimum of 60.8, mainly by stabilizing planning patterns. Because both the training reward and the evaluation optimum come from the same DP oracle on a deliberately simple office TES plant (chosen so DP is tractable), the headline closeness is partly by construction: the model is being steered toward the same action-value surface it is later scored against. The abstract itself flags that higher-fidelity whole-building and city-scale settings will lack exact DP verifiers, yet the reported robustness checks (forecast error, one unseen TES condition, limited battery transfer) remain inside or adjacent to the same simple plant. Without an independent non-DP reward or a higher-fidelity plant where the DP optimum is unavailable, it is unclear whether the stabilized patterns are genuine transferable planning skill or reward-hacking of the DP surface. That circularity is the single most load-bearing soft spot for the claim that DP-based RLVR is a practical route to scalable TES control.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper adapts an open-weight reasoning model for thermal energy storage (TES) scheduling via reinforcement fine-tuning (RFT) with dense rewards converted from exact offline dynamic-programming (DP) action values. On a deliberately simple office-building TES benchmark where exact DP is tractable and the optimum is known, RFT reduces emissions from 70.5 to 61.2 kg-CO2, near the DP optimum of 60.8 kg-CO2, using only 30 training prompts. Trace analysis attributes the gain mainly to stabilization of planning patterns (candidate comparison, look-ahead, feasibility checking) rather than a new strategy. GPT-5 nearly matches DP/MPC without task-specific training, while non-reasoning GPT-4o underperforms the no-storage baseline. Robustness checks cover forecast error, one unseen TES condition, and limited transfer to a battery task; the abstract motivates higher-fidelity whole-building tests and scalable verifiers.","tokens_in":2436,"tokens_out":1142,"duration_ms":22396,"significance":"If the central results hold under independent evaluation, the work is a concrete demonstration that DP-derived dense rewards can steer an open-weight reasoner toward near-optimal TES control with very few prompts, and that the mechanism is pattern stabilization rather than strategy invention. Explicit scoring against a known DP optimum, proprietary reasoning/non-reasoning baselines, and ablations (forecast error, unseen TES, battery transfer) are methodological strengths relative to typical RL-for-buildings claims that lack a known optimum. Framing DP-based RLVR as a practical adaptation route, while openly motivating higher-fidelity plants and scalable verifiers, is a useful contribution to LLM-for-control and building energy management.","major_comments":[{"comment":"Both the dense training rewards and the headline evaluation metric (61.2 vs DP optimum 60.8 kg-CO2) are derived from the same offline DP action-value surface on the same simple plant. The reported closeness is therefore partly by construction of the verifier. The manuscript needs either an independent non-DP evaluation metric (e.g., rule-based or MPC cost under identical dynamics, or a held-out plant model) or a quantitative decomposition separating reward-surface fitting from transferable planning skill. Without that, the claim that DP-based RLVR is a practical route to scalable TES control remains under-supported.","section":"Abstract (reward design and headline metric)"},{"comment":"The abstract itself notes that higher-fidelity whole-building and city-scale settings will lack exact DP verifiers, yet the reported robustness checks (forecast error, one unseen TES condition, limited battery transfer) remain inside or adjacent to the same DP-tractable office plant. Battery transfer is acknowledged to be limited by structural difference. For the load-bearing transfer claim, provide either a higher-fidelity plant evaluation where the DP optimum is unavailable, or independent (non-DP-referenced) metrics of the stabilized planning patterns that can be scored outside the training oracle.","section":"Abstract (robustness and generalization; closing motivation)"},{"comment":"The mapping from DP action values to dense rewards is a free, load-bearing design choice, as is the use of only 30 training prompts. The manuscript must fully specify this mapping, report sensitivity of the 61.2 kg-CO2 result to alternative mappings and prompt counts, and justify the design. Without that specification and sensitivity, reproducibility and the claim of a practical adaptation method cannot be assessed.","section":"Abstract (\"We convert exact offline dynamic-programming (DP) action values into dense rewards\"; \"Using only 30 training"}],"minor_comments":[{"comment":"Report number of evaluation episodes/seeds and variability (std or CI) for 70.5, 61.2, and 60.8 kg-CO2 so the 0.4 kg-CO2 gap to DP is interpretable.","section":"Abstract (reported emissions numbers)"},{"comment":"Operationally define the trace-analysis categories (candidate comparison, look-ahead, feasibility checking) so the claim that RFT mainly stabilizes these patterns is falsifiable.","section":"Abstract (trace analysis claim)"},{"comment":"State the MPC baseline emissions on the same benchmark for the fine-tuned open-weight model, not only for GPT-5, to complete the evaluation design.","section":"Abstract (GPT-5 vs DP and MPC)"},{"comment":"Clarify the no-storage baseline emissions and whether the pre-RFT open-weight model was prompted under the same protocol as GPT-5/GPT-4o.","section":"Abstract (baselines)"}],"recommendation":"major_revision","confidential_remarks":"This review is based solely on the abstract (full text was not available in the review package). Severity of the oracle-circularity and transfer concerns may change if the body contains independent metrics or higher-fidelity experiments not summarized in the abstract. Reasonable fit for cs.LG / applied RL venues interested in verifiable rewards and LLM control, provided the shared-oracle issue is addressed head-on rather than only deferred to future work."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know is that this is a practical recipe paper, not a control breakthrough: they convert exact offline DP action values into dense per-action rewards and RFT an open-weight reasoner as a text-based upper-level TES scheduler. With only 30 prompts on a deliberately simple office benchmark where DP is tractable, emissions drop from 70.5 to 61.2 kg-CO2, near the known DP optimum of 60.8. That is the concrete new empirical result.\n\nWhat they do well is keep the evaluation honest relative to the setup. They report a hard optimum baseline, compare GPT-5 (near-DP without task training) against GPT-4o (worse than no-storage), and use trace analysis to argue that RFT mainly stabilizes existing planning patterns—candidate comparison, look-ahead, feasibility checking—rather than inventing a strategy. The robustness checks (forecast error, one unseen TES condition, limited battery transfer) are modest but present, and the abstract itself flags the need for higher-fidelity plants and scalable verifiers. That self-awareness is useful.\n\nThe soft spot is real but proportionate: both the training reward and the headline metric come from the same DP oracle on the toy plant chosen so DP is exact. Closeness to 60.8 is therefore partly by construction of the verifier; that is intentional for RLVR, yet it raises the circularity burden the stress-test flags. The load-bearing claim is that the stabilized patterns transfer to settings where exact DP verifiers disappear. The reported probes stay inside or adjacent to the simple plant, so that claim is still open. No code, data, or prompts are visible from the abstract, which limits independent check.\n\nThis is for people working on LLM-for-control or building energy management who want a concrete RLVR recipe with a known-optimum baseline. It is not for someone looking for a general control method or city-scale result. The math and citation pattern look coherent on the available text; free parameters (prompt count, reward mapping) are at least named. I would send it to a serious referee as a methods paper on a known-optimum toy task. Expect revision pressure on independent grounding and higher-fidelity tests, but it is not desk-reject material.","headline":"Useful methods recipe for DP-verifier RFT of open-weight reasoners on a toy TES scheduler; near-DP numbers are partly by construction and transfer remains the open claim.","tokens_in":3148,"tokens_out":568,"would_cite":false,"duration_ms":8734,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Verifier-based reinforcement fine-tuning of an open-weight reasoning model cuts TES control emissions nearly to the known DP optimum.","keywords":["thermal energy storage","reinforcement fine-tuning","reasoning models","verifiable rewards","dynamic programming","building energy control","load shifting","RLVR"],"falsifier":"Apply the same RFT procedure to a higher-fidelity whole-building TES simulation (or a real building) where the DP optimum is no longer known; if the reinforced model no longer reduces emissions relative to the untrained baseline or loses its candidate-comparison and look-ahead patterns under realistic forecast error, the central claim fails.","tokens_in":3023,"feed_emoji":"🏢","tokens_out":610,"duration_ms":4276,"temperature":0.7,"pith_summary":"This paper tries to show that an open-weight reasoning model can be turned into a near-optimal thermal energy storage scheduler by reinforcement fine-tuning that uses dense rewards taken from exact offline dynamic-programming action values. The model acts as an upper-level controller that reads text-based building states and forecasts and outputs hourly heat-pump setpoints. On a deliberately simple office-building TES benchmark where the true optimum is known, the fine-tuned model lowers emissions from 70.5 to 61.2 kg-CO2, almost matching the DP optimum of 60.8 kg-CO2. Trace analysis indicates the gain comes mainly from stabilizing existing planning habits—candidate comparison, look-ahead, and feasibility checking—rather than inventing a new strategy. The result matters because it offers a practical route for adapting general reasoning models to building storage control without needing per-building model-predictive controllers or large amounts of online interaction data.","feed_headline":"DP rewards fine-tune a reasoning model to near-optimal TES control","feed_subtitle":"Emissions fall from 70.5 to 61.2 kg-CO2 on a simple office benchmark, close to the known optimum","key_machinery":"Dense DP-derived rewards for every candidate action, used inside reinforcement fine-tuning of an open-weight reasoning model so that the model can be trained as an upper-level TES scheduler from only 30 text prompts.","core_discovery":"Reinforcement fine-tuning that converts exact offline DP action values into dense rewards for every candidate action lets an open-weight reasoning model reduce TES-control emissions from 70.5 to 61.2 kg-CO2 on a simple office benchmark, approaching the known DP optimum of 60.8 kg-CO2, principally by stabilizing observable planning patterns rather than inventing a new strategy.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["DP rewards fine-tune open-weight model near TES optimum","RFT with DP action values cuts TES emissions to 61.2 kg-CO2","Verifiable DP rewards adapt reasoning model for TES scheduling","Open-weight model hits 61.2 kg-CO2 TES control after DP-RFT","DP-dense rewards stabilize LLM planning near DP TES emissions"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That the planning patterns learned on a deliberately simple office-building TES benchmark, chosen so exact DP is tractable, will transfer to higher-fidelity whole-building and city-scale settings where exact DP verifiers are no longer available.","fun_headline_variants_meta":{"raw":{"variants":["DP rewards fine-tune open-weight model near TES optimum","RFT with DP action values cuts TES emissions to 61.2 kg-CO2","Verifiable DP rewards adapt reasoning model for TES scheduling","Open-weight model hits 61.2 kg-CO2 TES control after DP-RFT","DP-dense rewards stabilize LLM planning near DP TES emissions"]},"model":"grok-4.5","effort":"low","cost_usd":0.00504,"raw_usage":{"total_tokens":1500,"prompt_tokens":894,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":50400000,"prompt_tokens_details":{"text_tokens":894,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":522,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":894,"tokens_out":84,"duration_ms":4258,"temperature":1.0,"reasoning_tokens":522,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T02:49:24.064950+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Apply the same RFT procedure to a higher-fidelity whole-building TES simulation (or a real building) where the DP optimum is no longer known; if the reinforced model no longer reduces emissions relative to the untrained baseline or loses its candidate-comparison and look-ahead patterns under realistic forecast error, the central claim fails.","supporting_citations":[],"review_version":1}