{"id":"271c0c89-e946-4017-9c3e-386f6f945343","arxiv_id":"2601.22901","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"For an ISAC status-update problem, the paper claims the optimal policy has a monotone threshold structure, but the proof of the required monotonicity lemma is flawed.","lead":"This paper models a wireless base station that each time step chooses between sensing a remote source or communicating previously sensed data, to keep the source's information fresh at minimum cost. The paper claims the optimal policy is a monotone threshold rule, but a key lemma in the proof is applied incorrectly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 4 applies submodularity in the wrong order; the monotone-threshold proof is invalid.","rationale":"I examined the proof of Theorem 1 in Section III. The central mechanism is to show the value function lies in F (coordinatewise nondecreasing and submodular) and then use Lemma 4 to establish that the action difference ∆ is nonincreasing in αb, yielding the nondecreasing threshold τ(αb). The reader's critique of Lemma 4 is accurate. Submodularity gives decreasing differences: for a≤a′, V(a,b′)−V(a,b) ≥ V(a′,b′)−V(a′,b). Lemma 4's proof asserts the reverse inequality without conditioning on the order of αs+1 and αb+1. In the regime αs<αb, the asserted inequality is reversed, so the proof fails. I confirmed that the lemma's claim is actually false for a valid submodular function: for V=−e^{−(x+y)}, at (αs,αb)=(0,10), Db>0, so ∆ is not nonincreasing in αb for all V∈F. Since Lemma 4 is essential for the monotone switching curve, the central claim is unproven. Lemma 5 is also only sketched. The numerical experiment in Section IV uses a single parameter set and, as a finite truncation, cannot establish the infinite-horizon structural result. Therefore the paper should be rejected unless a corrected proof is supplied. This agrees with the reader's verdict.","tokens_in":9478,"tokens_out":13395,"duration_ms":116564,"concrete_test":"Analytically verify the counterexample: set V(x,y)=−e^{−(x+y)}, λs=0.6, λc=0.9, cs=cc=0, and use Eq. (12) to evaluate Db at (αs,αb)=(0,10). A positive Db (≈1.17×10^{-6}) falsifies Lemma 4 as stated for functions in F, confirming the proof of Theorem 1 is invalid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 4 (Section III.C) is the linchpin for the nondecreasing switching curve τ(αb). Its proof asserts that submodularity of V gives V(αs+1,αb+2)−V(αs+1,αb+1) ≤ V(αb+1,αb+2)−V(αb+1,αb+1) for all αs,αb. But submodularity on N0^2 yields decreasing differences: for a≤a′, V(a,b′)−V(a,b) ≥ V(a′,b′)−V(a′,b). The claimed inequality holds only when αs+1 ≥ αb+1; for αs+1 < αb+1 the reverse inequality follows. The proof uses this bound to conclude Db = ∆(αs,αb+1)−∆(αs,αb) ≤ 0, but in the regime αs<αb the bound cannot be obtained. This is not merely a gap: for V(x,y)=−e^{−(x+y)} (coordinatewise nondecreasing and submodular), with λs=0.6, λc=0.9, cs=cc=0, and (αs,αb)=(0,10), one computes Db ≈ 1.17×10^{-6} > 0, so ∆ is not nonincreasing in αb for every V in F. Hence Lemma 4 is false as stated, and the proof of Theorem 1 collapses. Lemma 5's sketch ('precludes oscillatory action patterns') is also insufficient, but Lemma 4 is the decisive failure. The numerical experiment (Fig. 3) only covers one parameter set and cannot substitute for a proof.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies an ISAC-enabled base station that, at each time slot, chooses between sensing the state of a remote source and communicating previously sensed state information back to the source. Sensing and communication are modelled as independent Bernoulli trials with success probabilities λs and λc, and incur costs cs and cc. The state is the AoI pair (αs, αb); the stage cost is αs plus the action cost. The main claim (Theorem 1) is that the optimal stationary deterministic policy for the discounted infinite-horizon MDP has a monotone switching structure: for a nondecreasing function τ(αb), sense when αs ≤ τ(αb) and communicate otherwise. The proof proceeds by showing that the value function lies in the class F of coordinatewise nondecreasing submodular functions and that the action-difference function is increasing in αs and decreasing in αb. A numerical value-iteration experiment illustrates the claimed structure.","tokens_in":9789,"tokens_out":9222,"duration_ms":92558,"significance":"The problem is well motivated and the MDP formulation is clean. A monotone switching curve, if rigorously established, would be an interpretable and implementable scheduling rule and a useful contribution to the ISAC/AoI literature. The paper does not fit parameters or derive predictions from its conclusion, and its numerical study is appropriately presented as illustrative. However, the central proof rests on Lemma 4, and that lemma is false as stated. Because Theorem 1 is not established by any other argument in the manuscript, the main contribution cannot be accepted as proven in the current form.","major_comments":[{"comment":"The proof of Lemma 4 asserts the inequality V(αs+1, αb+2) − V(αs+1, αb+1) ≤ V(αb+1, αb+2) − V(αb+1, αb+1) as a consequence of submodularity. For a submodular V, the vertical increment V(a, b+1) − V(a, b) is nonincreasing in a, so the displayed inequality is valid only when αs+1 ≥ αb+1. When αs < αb, the reverse inequality holds. Since Lemma 4 is stated for all (αs, αb) ∈ N0², the proof fails in the regime αs < αb. This is not a minor gap: for V(x,y) = −e^{−(x+y)} ∈ F with λs=0.6, λc=0.9, cs=cc=0, and (αs, αb)=(0,10), equation (12) gives Db = Δ(αs, αb+1) − Δ(αs, αb) ≈ 1.17×10⁻⁶ > 0. Thus Δ is not nonincreasing in αb for every V ∈ F, and Lemma 4 is false as stated. Section III.E invokes Lemma 4 to prove that τ(αb) is nondecreasing, so the proof of Theorem 1 does not go through.","section":"Section III.C, Lemma 4"},{"comment":"The proof of Lemma 5 is only a one-sentence sketch: the single-crossing property of Δ is said to 'preclude oscillatory action patterns' and thereby preserve submodularity of min{Qsense, Qcomm}. No 2×2-lattice verification is given. Since Lemma 5 is needed for Lemma 6 (V* ∈ F), this is another load-bearing step that requires a rigorous proof even if Lemma 4 were repaired.","section":"Section III.D, Lemma 5"},{"comment":"The deduction of the monotone threshold relies entirely on Lemma 4, which is false as stated. The numerical experiment in Section IV shows only one parameter set; the claim that 'all results reported below are robust to variations' is not supported by any displayed parameter sweep. The numerics illustrate the desired geometry but cannot substitute for the missing proof.","section":"Section III.E and Section IV"}],"minor_comments":[{"comment":"The phrase 'with truncation at αi ≤ Amax' is introduced parenthetically but the theorem is stated for the infinite state space N0². Please clarify whether Theorem 1 concerns the infinite-state MDP or the truncated one, and explain how truncation affects the optimality argument.","section":"Section III.A"},{"comment":"The term 'single-crossing property' is used without a formal definition. Since this property is central to the switching argument, it should be defined explicitly and stated as a lemma, including the exact sense in which the sign of Δ changes.","section":"Section III.D"},{"comment":"Figure 3 reports a single parameter realization (Amax=30, γ=0.95, λs=0.6, λc=0.9, cs=0.2, cc=0.1). The statement that the results are robust to parameter variations should be substantiated with additional runs or tempered to what the displayed data support.","section":"Section IV"}],"recommendation":"reject","confidential_remarks":"The central proof is invalid: the counterexample to Lemma 4 is concrete and decisive. Because Theorem 1 depends on that lemma and no alternative argument is supplied, the manuscript's main contribution is unproven. Repairing this would require either a substantially different structural proof or additional assumptions on the value function, not a local fix. I therefore recommend rejection. The underlying problem and the claimed policy geometry are interesting, so a future substantially revised version might be worth considering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the thing: the central claim is unproven. The paper formulates a sensible ISAC status-updating problem as an MDP with two-dimensional AoI, and the claimed monotone switching curve is exactly the kind of structural result that would make optimal scheduling implementable. But the proof of that theorem goes through Lemma 4, and Lemma 4 is wrong.\n\nWhat is new: the problem setup — an ISAC base station that chooses each slot between sensing the source and transmitting the previously sensed state, with both links lossy, and a freshness objective at the source — is a legitimate new extension. The action-difference expression in Eq. (12) is clean, and the Bellman-operator approach is standard but appropriate. The numerical experiment (only one parameter set) does show the expected threshold shape, but that is illustration, not evidence.\n\nThe soft spot is specific: Lemma 4 tries to show that the action difference Δ(αs,αb) is nonincreasing in αb for any V in the class F of coordinatewise nondecreasing submodular functions. The proof uses submodularity to assert that the vertical increment of V at first coordinate αs+1 is no larger than at first coordinate αb+1. Submodularity gives that inequality only when αs+1 ≥ αb+1; for αs+1 < αb+1 the inequality is reversed, and the bound used in the proof is unavailable. The stress-test counterexample V(x,y) = -e^{-(x+y)} is in F, and with λs=0.6, λc=0.9, and (αs,αb)=(0,10), the forward difference in αb is positive, so Lemma 4 is false for the stated class. Since the nondecreasing threshold τ(αb) relies on this lemma, Theorem 1 is not established. Lemma 5's proof is also more of a sketch than a proof, but the Lemma 4 flaw is decisive.\n\nThis is not a circular or sloppy paper. The error is a genuine mathematical slip in a standard argument, and the problem is well posed. It belongs in front of a referee who can ask the authors to either add structural assumptions on V* that restore the inequality, or give a counterexample to the theorem. I would not desk-reject it.\n\nThe paper is for researchers in AoI scheduling and ISAC resource management who want to see threshold-policy machinery applied to a coupled sense/communicate problem. I would send it to peer review, expecting heavy revision on the proof section.","headline":"Nice ISAC status-updating formulation, but the main monotone-threshold theorem rests on a broken Lemma 4; the theorem is unproven as it stands.","tokens_in":10311,"tokens_out":4155,"would_cite":true,"duration_ms":41852,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C40","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"In an ISAC status-update system, the optimal sensing-vs-communication policy is a monotone threshold curve in the two-dimensional age-of-information state space.","keywords":["age of information","integrated sensing and communication","status updating","Markov decision process","threshold policy","submodularity","remote navigation","freshness"],"falsifier":"Compute the increment inequality used in Lemma 4 with a submodular value function such as V(αs, αb) = −e^{−(αs+αb)} at a state with αs < αb; the inequality reverses, so if value iteration ever produces such a function the claimed monotone threshold structure would not follow from the given proof.","tokens_in":9334,"feed_emoji":"📡","tokens_out":4794,"duration_ms":42695,"temperature":0.7,"pith_summary":"The paper studies a base station that can either sense a remote source's state or communicate previously sensed information back to that source, where both actions are unreliable and costly. Its aim is to characterize the policy that minimizes long-term cost including the source's age of information (AoI) plus sensing/communication overheads. The central result is that an optimal stationary deterministic policy has a monotone switching structure: for each base-station AoI, there is a threshold on the source AoI below which sensing is optimal and above which communication is optimal, and this threshold rises as the base station's AoI grows. If correct, this means a provably optimal scheduler can be implemented as a single interpretable curve rather than a lookup table, which matters for real-time operation.","feed_headline":"Optimal ISAC sensing policy is a single monotone switching curve","feed_subtitle":"The paper proves that the best sensing/communication tradeoff follows one threshold curve that rises as the base station's data ages.","key_machinery":"The central object is the action-value difference Δ(αs, αb) = Q_sense(αs, αb) − Q_comm(αs, αb). The paper argues that the Bellman operator preserves the class of coordinatewise nondecreasing submodular value functions, so the optimal value function is in this class; then Δ is nondecreasing in αs and nonincreasing in αb, producing a single-crossing property that yields the nondecreasing threshold curve τ(αb).","core_discovery":"The paper formulates the joint sensing-communication scheduling problem as a discounted infinite-horizon MDP with state (αs, αb) tracking the AoI at the source and at the base station. It proves (Theorem 1) that the optimal stationary deterministic policy is of threshold form: there exists a nondecreasing integer-valued function τ such that the optimal action is 'sense' when αs ≤ τ(αb) and 'communicate' otherwise. The nondecreasing property means that as the base station's information becomes staler, the system optimally favors sensing over communication for a weakly larger set of source-AoI values. The proof proceeds via the Bellman operator preserving coordinatewise monotonicity and submod","pith_inferences":["If the threshold structure is correct, it suggests that in practice only coarse AoI information (e.g., thresholds on age) is needed to implement near-optimal sensing/communication arbitration, which could be encoded in lightweight hardware.","The same monotone-structure argument might transfer to other two-state semantic metrics such as value of information or age of incorrect information, as long as the stage cost is coordinatewise nondecreasing and submodular.","One testable extension is to allow randomized policies or a third 'idle' action; the threshold form may then become a randomized switching region, and quantifying the loss relative to deterministic thresholds would bound the cost of simplicity."],"forward_implications":["The optimal policy can be encoded by a single nondecreasing curve, reducing implementation to comparing αs with τ(αb).","As the base-station AoI increases, the sense region expands: the system becomes more willing to renew its own observation even though communication is also available.","The structure holds for any costs and success probabilities satisfying λc ≥ λs, so the qualitative geometry is robust to the actual reliability and cost values.","The result gives a concrete performance guarantee for a freshness-based objective in ISAC, linking semantic metrics to structured decision rules."],"fun_headline_variants":["Sensing wins as base-station data ages: threshold proof","Optimal ISAC switching: one monotone AoI curve","Freshness policy: sense more when base info is stale","Proven: monotone sensing/communication threshold in ISAC","When to sense vs send: a single rising curve rule"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The proof of monotonicity in the base-station AoI assumes that for a submodular, coordinatewise nondecreasing value function, the vertical increment at the source coordinate is bounded by the vertical increment at the base-station coordinate, an inequality that submodularity only provides when the source coordinate is at least the base-station coordinate.","fun_headline_variants_meta":{"raw":{"variants":["Sensing wins as base-station data ages: threshold proof","Optimal ISAC switching: one monotone AoI curve","Freshness policy: sense more when base info is stale","Proven: monotone sensing/communication threshold in ISAC","When to sense vs send: a single rising curve rule"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1220,"prompt_tokens":730,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":420}},"tokens_in":474,"tokens_out":490,"duration_ms":5273,"temperature":1.0,"reasoning_tokens":420,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T06:19:00.487754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the increment inequality used in Lemma 4 with a submodular value function such as V(αs, αb) = −e^{−(αs+αb)} at a state with αs < αb; the inequality reverses, so if value iteration ever produces such a function the claimed monotone threshold structure would not follow from the given proof.","supporting_citations":[],"review_version":1}