{"id":"974dad2d-a57c-42ef-9910-7a34d1ce3a28","arxiv_id":"2602.09191","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A two-stage digital-twin scheduler for satellite-terrestrial spectrum sharing cuts simulated queue congestion by roughly 5-17 MB versus benchmarks and lands within 0.27 MB of an ideal full-information scheduler.","lead":"This paper builds a software 'digital twin' of London, its users, and a LEO satellite to predict traffic and radio channels, then uses those predictions to schedule spectrum and power in a shared satellite-terrestrial network. The two-stage scheduler cuts simulated buffer congestion far below standard benchmarks and comes within about 0.27 MB of an omniscient full-information scheduler.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DT-fidelity assumption is self-referential: real channels are generated from the same twin model (Eq. 8, ξ=0.5), so the 0.25 MB gap likely reflects the assumed error model rather than measured reality.","rationale":"The reader's weakest assumption correctly identifies the digital-twin replica fidelity as the load-bearing point. The paper's headline numerical result is a function of the assumed channel-error model in Eq. (8), with ξ=0.5 used to generate the simulated 'real' channels and ξ=1 used for prediction. This is self-referential: the synthetic real world is constructed from the same twin plus the same noise model, so the simulation cannot validate the twin's predictive accuracy in any physically grounded sense. A secondary issue is the convergence claims in Propositions 4–5, which appear to overclaim local optimality of the original MINLP from an SCA-relaxation argument; however, the empirical results do not depend on this proof, so the fidelity concern is more directly connected to the practical-feasibility conclusion. The proposed concrete test—calibrating ξ with the cited experimental data—would directly determine whether the claimed gap survives real-world error statistics. Since the reader already assigned CONDITIONAL with this same concern, my stress-test does not change the verdict; it strengthens the conditionality.","tokens_in":40843,"tokens_out":5063,"duration_ms":49083,"concrete_test":"Using the experimental C-band measurements from ref. [23], build twin predictions with the same 3D map and ray-tracing as in Section VI, pair them with measured channels, and estimate ξ (or the full empirical error distribution) in Eq. (8). Re-run Figs. 10–12 with this empirically calibrated ξ instead of ξ=0.5. If the RT-Reffine-to-FIA gap grows beyond ~1 MB, or RT-Reffine no longer beats the reference algorithm, the practical-feasibility claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The practical-feasibility claim—RT-Reffine within ~0.25 MB of full information—rests on the twin-replica error model in Eq. (8), where real NLoS channels equal sqrt(ξ) times the virtual channel plus a complex-normal error. Algorithm 1 predicts channels with ξ=1 (virtual channels exactly), while Section VI.A generates the 'real' environment with ξ=0.5 using the same error distribution. Thus the simulated prediction error is exactly the assumed error; there is no external ground truth. The gap between RT-Reffine and FIA (Fig. 12) and the 15.7% phase-2 refinement gain (Fig. 10) are direct functions of ξ. The authors cite an experimental C-band study [23] but do not use it to calibrate ξ or the error statistics. If actual ray-tracing prediction errors are larger, correlated, or biased, the performance gap widens and RT-Reffine may no longer beat the reference algorithm. The manuscript does not flag this missing calibration, making the central 'practical feasibility' conclusion self-referential until the error model is validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a digital-twin-aided two-stage resource-management framework for an integrated satellite-terrestrial network sharing the 5G-NR C-band. A DT model combining a 3D map, ray tracing, and mobility/traffic prediction is used to obtain channel and traffic forecasts for the next cycle; the DT-JointRA algorithm solves an MINLP for bandwidth allocation, traffic steering, association, RB assignment, and power control using predicted information, and the RT-Reffine stage re-optimizes the TN short-term decisions at sub-frame granularity using real-time feedback. The objective is mean queue length. Simulations using a London 3D map and real traffic traces show that RT-Reffine comes within about 0.25–0.27 MB of the Full-Information Algorithm and outperforms greedy, heuristic, and reference benchmarks. Proposition 4 claims that Algorithm 2 converges to a local optimum of the original MINLP.","tokens_in":41080,"tokens_out":6571,"duration_ms":63050,"significance":"If the DT fidelity model is accepted as realistic, the optimization architecture is valuable: the Full-Information Algorithm is an honest upper-bound control, the kappa trade-off is mechanistically explained, and the SCA convexification steps in Propositions 1–3 are worked out in appendices. The paper is also among the first to couple 3D-map/ray-tracing DT channel prediction with dynamic spectrum sharing and a real-time refinement stage in ISTNs. However, the headline practical-feasibility claim rests on the assumed channel-error model in Eq. (8), and the current evaluation does not validate that model against independent measurements. The convergence guarantee for the MINLP is also not rigorously established. These are substantial but fixable issues; they do not undermine the optimization machinery itself.","major_comments":[{"comment":"The central practical-feasibility result is not validated against an independent ground truth. In Eq. (8), real NLoS channels are generated from DT channels as sqrt(xi) times the DT value plus a complex-normal error; Section VI.A states that the simulation environment is generated from the DT model with xi=0.5, while Algorithm 1 (line 4) predicts channels with xi=1, i.e., the predicted channel is exactly the DT channel. The prediction error in the simulation is therefore exactly the error distribution assumed in Eq. (8). The 0.25–0.27 MB gap between RT-Reffine and the Full-Information Algorithm (Figs. 11 and 12) is a direct function of this assumed model; it does not by itself demonstrate practical feasibility. The experimental C-band study [23] is cited but not used to calibrate xi or the error statistics. The authors should either calibrate Eq. (8) with measured ray-tracing/channel dat","section":"Section II-D4 and Section VI.A"},{"comment":"Proposition 4 states that Algorithm 2 converges to a local optimum of the original MINLP (P0)_c, but the proof is not supplied in this manuscript and the argument given is insufficient. Deferring to Proposition 4 of [16] is not acceptable for a new, central claim. Moreover, the statement that 'the feasible set of (P2)_c is a subset of that of (P0)_c' is not meaningful as written: (P2)_c is a continuous SCA surrogate with l0-norm upper bounds and slack variables, while (P0)_c contains binary variables; the binary variables are recovered only after convergence by thresholding (36). A limit point of the continuous iterates need not be locally optimal for the mixed-integer problem, and the thresholding step can alter feasibility of the original constraints. The authors should either provide a rigorous mixed-integer local-optimality proof or weaken Proposition 4 to a statement about convergen","section":"Section IV.B, Proposition 4"},{"comment":"The traffic prediction used by the DT is simply the previous cycle's average (Eq. (28)), and no measure of traffic prediction error or its effect on QL is reported. Because the two-stage framework is specifically motivated by predicting future traffic and channels, the paper should quantify how the QL gap grows when the actual traffic differs from the previous-cycle average. This is load-bearing for the practical-feasibility claim, although less critical than the channel-fidelity issue in Major Comment 1.","section":"Section III.E and Eq. (28)"}],"minor_comments":[{"comment":"Only a single simulation scenario/trajectory is presented, with no error bars or multiple runs. Since the central differences are on the order of 0.25 MB, the authors should report run-to-run variability or state that the results are one representative realization.","section":"Section VI.A"},{"comment":"The manuscript contains numerous typos and OCR-like artifacts: 'digitial-twin' in the abstract, 'bechmarks' in Section V, 'heusistic' and 'algorihm' in Section VI, and 'Alg. 2 and ,' in Section V.D. A careful copyedit is needed.","section":"Throughout"},{"comment":"The symbol S is used both for the set of services and for the SatCom service, causing confusion in expressions such as K = K_D ∪ K_S ∪ K_M, L, S. Consider renaming the service set to avoid the clash.","section":"Section III"},{"comment":"The column 'Remaining D traffic (%)' is not defined. Please specify whether it is the percentage of unserved D-service bits over total D arrivals, per cycle or averaged, and state how N_SC^D is selected for the comparisons.","section":"Table II"},{"comment":"The y-axis label 'Average percentage' is vague. Clarify that the two plotted quantities are the mean QL reduction by phase 2 relative to phase 1 and the mean QL gap relative to the Full-Information Algorithm, respectively.","section":"Fig. 10"},{"comment":"The definition of xi_D after Eq. (13) appears garbled; check the formula for the finite-blocklength penalty term to ensure the notation is consistent.","section":"Eq. (13)"}],"recommendation":"major_revision","confidential_remarks":"The DT-fidelity circularity is the main risk to the paper's central claim. If the authors can calibrate Eq. (8) against measured channels or clearly re-scope the practical-feasibility claim as conditional on the assumed error model, I would support publication. The convergence proof in Proposition 4 also needs to be completed rather than deferred to [16]. I do not see grounds for rejection: the optimization cores, benchmarks, and the Full-Information-Information control are sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one-liner: this is a competent systems-engineering paper with a genuinely new joint scope, and it deserves referee time. But the headline feasibility claim should not be accepted until the digital-twin error model is calibrated and the convergence claims are tightened.\n\nWhat is actually new: no prior work in the cited set jointly optimizes bandwidth allocation, traffic steering, TAP/LSat-UE association, RB assignment, and power control over time windows with a 3D-map/ray-tracing digital twin. The Table I comparison is accurate. The Full-Information-Algorithm control is an honest upper bound, and the two-stage architecture with one satellite-ground round-trip per cycle is concrete and relevant to the FCC/ESA DSS agenda. The London 3D map and real traffic traces are real assets, and the SCA machinery is standard and internally consistent.\n\nWhere it gets soft: the DT-fidelity model is self-referential in exactly the way the stress-test note says. Eq. (8) defines the real NLoS channel as sqrt(ξ) times the twin channel plus complex-normal error. Algorithm 1 predicts channels with ξ=1, while the simulator generates the \"real\" environment with ξ=0.5 using the same error distribution. The 0.25–0.27 MB gap between RT-Reffine and FIA is therefore a direct function of an assumed error model, not measured reality. The authors cite their experimental C-band study [23] but never use it to set ξ. That missing calibration is not flagged as a limitation in the manuscript. Traffic prediction being the previous cycle's average is also a weak assumption.\n\nSecond, Proposition 4 asserts guaranteed convergence to a local optimum of MINLP (P0), with the proof deferred to the authors' [16]. The SCA argument gives monotone convergence of the relaxed problem, but the feasible-set-subset argument does not establish local optimality in the original MINLP. Same issue applies to Proposition 5. I would ask the authors to prove a proper stationarity result or soften the language.\n\nMinor concerns: all results are single-run point estimates with no error bars; κ=1.1 is tuned on the evaluation scenario (at least this is transparent in the text); and the D-service SC count is tuned via Table II. These are minor compared to the DT issue.\n\nOverall: the optimization core is credible and the architecture is practically relevant, but the feasibility conclusion is conditional on an unvalidated twin-error model. This paper deserves a serious referee; a conditional accept after calibration or a major revision would be reasonable.","headline":"Solid systems-engineering scheduler with an honest FIA control, but the DT-feasibility claim rests on an assumed error model and the MINLP local-optimality proof is deferred.","tokens_in":41662,"tokens_out":2171,"would_cite":false,"duration_ms":21048,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a digital-twin-aided two-stage scheduler can bring mean queue length in an integrated satellite-terrestrial network to within about 0.25 MB of the full-information optimum while exchanging data with the satellite only","keywords":["digital twin","integrated satellite-terrestrial networks","dynamic spectrum sharing","resource allocation","LEO satellites","queue length minimization","successive convex approximation","compressed sensing"],"falsifier":"Measure, in a real urban deployment, the actual correlation between digital-twin ray-traced channels and measured channels (e.g., with the authors' C-band setup). If the measured ξ is substantially below 0.5, or if traffic prediction error exceeds a previous-cycle average by a large margin, then running RT-Refine with κ=1.1 will not reproduce the claimed ~0.25 MB gap to FIA; the gap will exceed the reported value and may approach the benchmark gap.","tokens_in":40622,"feed_emoji":"🛰️","tokens_out":4029,"duration_ms":38770,"temperature":0.7,"pith_summary":"The paper aims to show that digital twin predicted information can drive dynamic spectrum sharing and resource allocation in integrated satellite-terrestrial networks without the heavy signaling of real-time channel estimation from LEO satellites. It formulates two problems: a DT-based joint RA (DT-JointRA) that optimizes long-term traffic steering, bandwidth allocation, and short-term association/power/RB assignments using predicted information, and a real-time refinement (RT-Refine) that re-optimizes terrestrial short-term decisions with actual CSI and traffic each subframe. The central result is that the refined scheduler lands within about 0.25 MB of an omniscient full-information algorithm while using only one satellite-ground round-trip per cycle, and beats benchmark algorithms by 5–17 MB in mean queue length. If true, this would make digital twins a practical mechanism for reducing congestion and signaling overhead in 6G NTN integration.","feed_headline":"Predicted-only scheduling lands within 0.25 MB of omniscient","feed_subtitle":"A digital-twin scheduler cuts congestion to within 0.25 MB of the full-information bound, with one satellite round-trip per cycle.","key_machinery":"The digital twin is the enabler: it combines a 3D map, ray tracing, TLE orbit data, and predicted UE positions/traffic to forecast channels and load. Binary association variables are relaxed through a compressed-sensing-style ℓ0-norm approximation (F_apx), and non-convex SINR/rate constraints are convexified via successive convex approximation with slack variables. An interference margin κ multiplies predicted inter-system interference terms in the refinement stage to guard against twin fidelity errors, which is what allows the refinement to operate with only actual terrestrial UE channels while keeping satellite channels predicted.","core_discovery":"The paper's central claim is that a two-stage optimization framework, DT-JointRA followed by RT-Refine, can nearly match the performance of a full-information omniscient scheduler (FIA) using only predicted digital-twin information plus subframe-level refinement of terrestrial decisions. The load-bearing numeric result is that with an interference margin κ=1.1, the gap between RT-Refine and FIA is about 0.25 MB in mean queue length (0.27 MB across power-budget sweeps), while the gain over the reference algorithm is about 5–5.7 MB and over greedy/heuristic schemes about 17 MB. The paper also claims Algorithm 2 converges to a local optimum of the joint RA problem (Proposition 4) and Algorithm","pith_inferences":["This is an editorial inference: the paper's assumed twin fidelity (ξ=0.5 for NLoS errors, and ξ=1 for the DT channel used in optimization) is not calibrated against the cited C-band measurement study; if real-world twin error is larger or less stationary than this model, the 0.25 MB gap to FIA would widen.","This is an editorial inference: the traffic predictor used in Algorithm 1 is a simple previous-cycle average; the framework's practical advantage may be sensitive to traffic non-stationarity, and a more sophisticated predictor could shrink or enlarge the gap depending on environment.","This is an editorial inference: the interference margin κ=1.1 is tuned empirically; in deployment, κ would need to be adaptively set from live error statistics to avoid either under- or over-protection.","This is an editorial inference: the compressed-sensing ℓ0 relaxation and SCA machinery are general; the same two-stage DT-plus-refine pattern could extend to other NTN/TN resource management problems, e.g., uplink or multi-satellite coordination."],"forward_implications":["A network operator could run the joint RA on DT predictions and only refine terrestrial decisions at subframe level, cutting LEO signaling to one round-trip per cycle.","The near-zero gap to FIA suggests that, under the assumed twin fidelity, prediction error is almost fully compensated by the refinement stage; hence DT-based scheduling need not wait for perfect channel knowledge.","The algorithms satisfy delay-sensitive (D) service constraints in all tested cases, while heuristic and reference schemes leave 17.5–22.2% of D traffic unserved.","Because phase-2 refinement converges in about 3 iterations, the approach is compatible with subframe-level (1 ms) real-time operation."],"fun_headline_variants":["Digital-twin scheduler closes 0.25 MB gap to omniscient","Predicted info alone hits within 0.25 MB of full knowledge","Two-stage scheduling beats greedy by 17 MB","DT-aided scheduling lands 0.25 MB from ideal bound"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole near-optimality result rests on the assumption that the digital twin's predicted channels and traffic are accurate enough — specifically, that real NLoS channels follow Eq. (8) with correlation ξ=0.5 to the twin's channels, and that next-cycle traffic equals the previous cycle's average; if actual twin error is worse than this, the 0.25 MB gap to the full-information optimum will grow.","fun_headline_variants_meta":{"raw":{"variants":["Digital-twin scheduler closes 0.25 MB gap to omniscient","Predicted info alone hits within 0.25 MB of full knowledge","Two-stage scheduling beats greedy by 17 MB","DT-aided scheduling lands 0.25 MB from ideal bound"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00027,"raw_usage":{"total_tokens":1481,"prompt_tokens":785,"completion_tokens":696,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":622}},"tokens_in":529,"tokens_out":696,"duration_ms":6586,"temperature":1.0,"reasoning_tokens":622,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:00:50.684746+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, in a real urban deployment, the actual correlation between digital-twin ray-traced channels and measured channels (e.g., with the authors' C-band setup). If the measured ξ is substantially below 0.5, or if traffic prediction error exceeds a previous-cycle average by a large margin, then running RT-Refine with κ=1.1 will not reproduce the claimed ~0.25 MB gap to FIA; the gap will exceed the reported value and may approach the benchmark gap.","supporting_citations":[],"review_version":1}