{"id":"a2dbbad6-66f5-4026-a839-affd46bc8b56","arxiv_id":"2507.00928","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Measurements of O-RAN component power consumption under varying PRB load show that power draw is largely load-invariant across Split 8 and Split 7.2b hardware.","lead":"This paper measures how much electrical power O-RAN radio, distributed unit, and central unit hardware uses under different network loads, for two functional split configurations. It finds that most of the power draw stays constant even as traffic rises, which matters for designing energy-saving and digital twin models for 5G networks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DU/CU power values conflate full-server RAPL estimates with component-level draws, so the load-invariance claim is not yet attributable to DUs/CUs.","rationale":"The reader's verdict of CONDITIONAL is appropriate: the paper contains a direct, reproducible-looking measurement campaign and the load-invariance pattern is visible even in the raw RU data obtained from a metered PDU, which does not depend on powerstat. However, the single most load-bearing weakness is not just that powerstat estimates rather than measures; it is that the estimates are host-level CPU energy values for servers that simultaneously run several containers, so the tables labelled 'CU' and 'DU' do not isolate those components. This matters because the headline recommendation is specifically about switching RAN components on and off: if the constant power is actually RIC/Core/container overhead, the component-level savings from switching off a DU or CU may be different from the reported numbers. The RAPL measurement technology can be calibrated and the component attribution can be fixed by re-running with separated hosts and a wall-power meter, so the paper is not fundamentally unsound; the issues are addressable in a revision. The inconsistent quadratic fits in Section V reinforce the need for correction but do not overturn the qualitative conclusion from the tables. I therefore keep the reader's CONDITIONAL verdict rather than escalating to REJECT, while noting that the requested revision should explicitly separate DU/CU power from RIC/Core/background power and provide a wall-power calibration check.","tokens_in":8249,"tokens_out":7696,"duration_ms":160369,"concrete_test":"Repeat a subset of the Split 7.2b measurements with a direct wattmeter at the PDU/PSU for each server, run the DU and CU on separate servers, and move RIC and Core to a third host or disable them; then recompute the load sensitivity of DU and CU separately. If the per-component wall-power deltas from 0% to 100% PRB are within the same roughly 3-8 W range seen in Tables VI and VII, the load-invariance claim survives; if the deltas deviate substantially, the current tables conflate server overhead with DU/CU energy. Additionally compare powerstat values against the wattmeter to quantify RAPL bias in absolute terms.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that DU/CU power is load-invariant rests on measurements that are neither component-level nor wall-power. Section III states that server power was obtained with powerstat, 'based on intel RAPL' and 'estimate the power consumed based on a power consumption model'; the Split 8 server runs a Docker image of CU, DU, RIC and Core on one host, and the Split 7.2b 'CU' server hosts RIC and Core alongside the CU. RAPL measures CPU package energy (on some platforms also DRAM), not whole-server draw, and it cannot attribute energy to individual containers. Therefore Tables IV, VI and VII report host/software-stack power, not DU/CU power. The observed flatness may be dominated by RIC/Core/container overhead and Linux idle behaviour; the DU/CU-specific processing energy could scale differently. The recommendation to switch components off rather than scale active power is a reasonable hypothesis, but the component-level evidence needed to support it is not in the paper. The quadratic fits in Section V are also presented in a way inconsistent with Table IV (intercepts near 61-62 W rather than 119.5 W at 0% PRB), consistent with equations that omit the idle baseline or use unstated normalization; this must be corrected for the model to be usable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an experimental measurement study of power consumption in O-RAN functional splits, specifically Split 8 and Split 7.2b, under varying PRB utilization. RU power is measured with a metered PDU for a USRP-based RU and for commercial Benetel RAN550/RAN650 RUs; DU and CU power is estimated with powerstat, which the text describes as an RAPL-based power consumption model. The measurements show RU power nearly constant and DU/CU power increasing only slightly with PRB utilization. The authors fit quadratic polynomials to the combined Split 8 DU+CU data and conclude that power does not scale significantly with load, recommending that energy optimization focus on switching off radio chains or components rather than on load-adaptive scaling. They further claim that these measurements provide foundations for developing accurate digital twins of O-RAN deployments.","tokens_in":8620,"tokens_out":5548,"duration_ms":62498,"significance":"The RU measurements are a genuine empirical contribution: they use a metered PDU, cover both functional splits and both uplink and downlink, and show a clear load-invariant pattern for the RU. If the DU/CU estimates were trustworthy, the conclusion would have practical implications for energy-saving strategies in O-RAN. The paper's recommendation to prioritize component-level on/off switching over fine-grained load-adaptive power scaling is a concrete, testable hypothesis. At present, however, the significance is substantially diminished by the unvalidated RAPL-based estimates for DU/CU and by the internally inconsistent quadratic model in Section V; the digital-twin claims are not yet substantiated.","major_comments":[{"comment":"The DU/CU power measurements are not component-level and are not wall-power. The text explicitly states that server power was obtained with powerstat, 'based on intel RAPL,' and that the tool 'estimate[s] the power consumed based on a power consumption model.' Figure 2 shows that in Split 8 a single server hosts the CU, DU, RIC, and Core in one Docker image, and in Split 7.2b the CU server also hosts RIC and Core. RAPL measures CPU package energy (and on some platforms DRAM), not per-container or per-VNF energy, and it cannot attribute energy to the DU or CU. Tables IV, VI, and VII therefore report host/software-stack power, not DU/CU power. The observed flatness may be dominated by always-on RIC/Core/container overhead and Linux idle behavior, rather than by DU/CU processing characteristics. The central claim that DU and CU power are load-invariant is not supported by these data; the authors must either isolate the DU and CU (for example, separate hosts, container-level energy accounting with calibration, or direct AC measurement), or explicitly reframe the conclusion as applying to the whole server-side software stack.","section":"Section III, Tables IV, VI, VII"},{"comment":"The fitted quadratic models are internally inconsistent with Table IV. At l=0, Eqs. (1) and (2) predict 61.06 W and 62.28 W, respectively, whereas Table IV reports 119.51 W at 0% PRB (and 58.34 W at idle). The fits appear to model the power increase over the idle baseline, but the text does not state this; as written, the model would badly mispredict the 0% PRB operating point that the digital twin would need. In addition, the coefficient 22.12 on l^2 in Eq. (2) implies that l is normalized to [0,1], while the text defines l as the PRB utilisation percentage; if l were a percentage, the predicted value at 100% PRB would be implausibly large. The equations must be corrected to include the baseline explicitly and the normalization of l must be defined unambiguously.","section":"Section V, Eqs. (1) and (2)"},{"comment":"The quadratic model is fit to the same data it is compared against, with no validation on held-out measurements, no cross-validation, and no goodness-of-fit or prediction-error statistics. Because the paper's stated purpose is to 'provide foundations for developing accurate digital twins' (Section I), the authors should demonstrate that the model can predict measurements not used in fitting, or at minimum report residual errors and a train/test split. As it stands, Fig. 3 is a curve fit to the training data; the predictive value of the model for digital-twin forecasting is untested.","section":"Section V, Fig. 3"}],"minor_comments":[{"comment":"The list of testbed components says 'including the RU, DU and DU'; this should read 'RU, DU and CU.'","section":"Section III-A"},{"comment":"The abbreviations 'fronthall,' 'midhall,' and 'backhall' should be corrected to 'fronthaul,' 'midhaul,' and 'backhaul.'","section":"Fig. 2 caption"},{"comment":"The text contains typographical errors: 'receptively' should be 'respectively' and 'utlisation' should be 'utilisation.'","section":"Section IV-B"},{"comment":"The symbol l is described as 'PRB utilisation percentage' but the fitted coefficients are consistent with l being a fraction in [0,1]; please define the variable and its range explicitly in the text and in the figure axis label.","section":"Section V"},{"comment":"Reference [17] is cited as a prior study measuring RU and DU power in an Open RAN-based LTE system, but the listed title refers to deep convolutional neural network based reinforcement learning for mobile network power saving; please verify that the citation matches the claimed content.","section":"References"},{"comment":"The paper emphasizes digital twins in the title, abstract, and conclusions, but it contains no digital twin simulation, use case, or demonstration of how the fitted power model improves digital twin accuracy; please either add a concrete digital twin experiment or temper the claims.","section":"Title, Section I, Section VI"},{"comment":"The captions describe the values as power consumption in Watts, but these values are powerstat estimates, not direct measurements; the captions should note the estimation method.","section":"Tables VI and VII"}],"recommendation":"major_revision","confidential_remarks":"The RU measurements are the strongest part of the paper and would be worth publishing if the DU/CU measurement methodology were made sound and the Section V model inconsistency corrected. The quadratic model's mismatch with Table IV is a red flag that the modeling was not carefully checked against the reported data; the authors should be asked to provide the raw measurement logs and the fitting procedure. If component-level DU/CU power cannot be obtained, the conclusions should be explicitly limited to the full processing stack, which would still be a useful empirical observation but would weaken the paper's stated contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful measurement dataset, not a breakthrough. The RU figures are credible; the DU/CU figures need a heavy caveat.\n\nWhat's genuinely new: first published power measurements for O-RAN splits 8 and 7.2b in both directions, covering an SDR USRP and two commercial Benetel RUs. The RU measurements come from a metered PDU and are simple, direct, and believable. The main result there—RU power barely moves with PRB utilisation—is a solid, reproducible observation, and it matters for anyone building O-RAN energy models. The testbed description is clear enough to reproduce the RU side.\n\nThe soft spots are concentrated in the DU/CU half. Power for the servers comes from powerstat, which is a RAPL-based estimate, not a wattmeter, and the hosts also run the RIC, Core network, and containers. So Tables IV, VI and VII report whole-server stack power, not DU/CU component power. The claim that DU/CU consumption is load-invariant therefore over-reaches. The data may still show the host is flat, but that's a different statement.\n\nAlso, the quadratic fits in Section V are internally inconsistent with Table IV. The intercepts (61–62 W) are about half the measured 0% PRB values (119.5 W), and if l is meant to be a percentage the coefficients produce absurd values at 100%. It looks like they're fitting an increment over idle or a normalized load, but they don't say. That needs correction, plus error estimates for the fits.\n\nThe citation list has a real problem: [17] is described as an Open RAN RU/DU measurement, but the cited paper is a deep-reinforcement-learning power-saving study. That's a mis-citation and should be fixed.\n\nNone of this is fatal to the empirical core. The RU data stands on its own; the DU/CU analysis needs to be reframed as host-level power with all its caveats. The paper deserves a serious referee because the measurements are new and potentially useful for digital twin calibration and energy-saving research. I'd send it out with expectations of major revision.\n\nFor who it's for: people working on O-RAN energy modeling, digital twins, and energy-saving algorithms. I'd cite the RU data but not the DU/CU fits.","headline":"Useful RU power measurements for O-RAN splits, but DU/CU claims rest on host-level RAPL estimates.","tokens_in":9174,"tokens_out":3227,"would_cite":true,"duration_ms":34099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Measured O-RAN power consumption stays nearly flat as network load rises, so energy savings should come from switching components off rather than scaling active ones.","keywords":["O-RAN","Open RAN","power consumption","functional split","PRB utilisation","energy efficiency","digital twin","software-defined radio"],"falsifier":"Measure the same DU and CU servers with a calibrated wattmeter while sweeping PRB utilisation from 0 to 100 percent; if total server power rises by much more than the few watts reported here, the claim that O-RAN power is largely load-independent would fail.","tokens_in":8082,"feed_emoji":"⚡","tokens_out":9902,"duration_ms":98550,"temperature":0.7,"pith_summary":"This paper reports measured power consumption for the components of two Open RAN (O-RAN) deployments—Split 8 and Split 7.2b—under varying Physical Resource Block (PRB) utilisation, the share of scheduled radio resource blocks, from 0 to 100 percent. The authors aim to establish how much energy actually scales with network load and to provide the empirical data needed to build a digital twin of an O-RAN deployment. Their central finding is that power consumption in Radio Units (RUs), Distributed Units (DUs), and Central Units (CUs) stays nearly flat as load increases: the RU barely moves and the DU/CU draw only a few extra watts from idle to full load. This matters because it suggests that load-adaptive power scaling inside active components has limited headroom, and that energy-efficiency strategies should instead target switching off radio chains or entire components when traffic is low. The authors also fit quadratic models to the Split 8 DU/CU measurements for use in digital-twin simulations.","feed_headline":"Measured O-RAN power stays flat as load rises","feed_subtitle":"Across two functional splits, most energy is constant base load, so savings come from powering components off.","key_machinery":"The load-bearing apparatus is a pair of physical O-RAN testbeds instrumented so that power can be read per component while PRB utilisation is swept. Split 8 uses a USRP software-defined radio as the RU with srsRAN-based DU and CU; Split 7.2b uses commercial Benetel RAN550 and RAN650 RUs with a commercial DU/CU stack. RU power is read from a metered power distribution unit, while server DU/CU power is estimated with the Linux powerstat tool, which uses Intel RAPL counters and a power model. The two functional splits—Split 8 separating RF from PHY, and Split 7.2b separating low-PHY from high-PHY—define where baseband processing is done and therefore where load-dependent energy could appear. The measured power-versus-PRB-utilisation curves, plus the quadratic fits for Split 8 DU/CU, are the machinery that carries the argument.","core_discovery":"The paper's central discovery is that O-RAN power consumption is dominated by a constant base load rather than by traffic-dependent processing. For the Split 8 USRP-based RU, mean power stays between 43.1 W and 45.0 W across PRB utilisation from 0 to 100 percent; the commercial Benetel RAN550 indoor RU holds at roughly 28–30 W and the RAN650 outdoor RU at 44–46 W. The server-hosted DU and CU show more movement but still modest: in Split 8 the combined DU/CU rises from 119.5 W at zero load to 125.2 W (uplink) and 141.6 W (downlink) at full load, and in Split 7.2b the DU and CU each vary by only a few watts across the whole sweep. Downlink power is consistently higher than uplink, matching the higher downlink throughput. The paper concludes that because the load-dependent portion is small, O-RAN energy optimisation should concentrate on switching radio chains or components off rather than on fine-grained scaling of active components, and it provides quadratic fits of the Split 8 DU/CU measurements as a basis for digital-twin energy modelling.","pith_inferences":["If this flat-power pattern holds on other O-RAN hardware, the practical energy lever shifts from improving baseband processing efficiency to cell-level sleep and radio-chain shutdown, and network energy models should be driven mainly by which components are active, not by instantaneous PRB utilisation.","Because the server-side figures come from a RAPL-based power model rather than a wattmeter, a direct metering pass on the same servers would show whether true DU/CU load scaling is even smaller or meaningfully larger; this is the cleanest test of the conclusion.","A natural next experiment is to measure power while actually switching off radio chains or placing RUs in sleep mode at low PRB utilisation, which would quantify the savings this paper's conclusion implies.","The quadratic fits are tied to these specific servers and RUs; using them in a digital twin for different hardware would require re-measuring the constant and linear coefficients rather than assuming they transfer."],"forward_implications":["Energy-saving algorithms for O-RAN should prioritise switching off radio chains or whole components during low traffic, because the load-dependent part of the power draw is small.","Digital-twin energy models should treat a large constant base-power term as the dominant contributor and add a small traffic-dependent correction, using the fitted curves as a starting point.","Downlink power exceeds uplink at full load on both testbeds, so downlink-heavy configurations have slightly more absolute power to recover, but the recoverable load-dependent fraction remains small.","Functional-split choice affects where the constant cost sits: the SDR-based Split 8 RU shows slightly more variation, while the commercial Split 7.2b RUs stay almost perfectly flat.","PRB utilisation alone appears sufficient as the load variable to capture the dominant energy behaviour on these testbeds, which simplifies measurement for digital-twin calibration."],"supporting_citations":[{"why":"O-RAN Alliance architecture specification that defines the CU, DU, RU and functional splits framing the testbeds.","marker":"[3]"},{"why":"O-RAN energy-saving use cases (sleep modes, radio chain switch-off) that this paper's conclusion targets.","marker":"[8]"},{"why":"Prior O-RAN energy-efficiency measurement platform that this study extends with split-specific and uplink/downlink measurements.","marker":"[15]"},{"why":"Prior CPU-level power measurement in virtualised RAN that supplies a comparison baseline for server-side power.","marker":"[16]"},{"why":"Earlier O-RAN-based LTE RU/DU power measurements limited to uplink, motivating the both-link and split-specific measurements here.","marker":"[17]"},{"why":"srsRAN software provides the DU and CU implementations used in the Split 8 testbed.","marker":"[18]"},{"why":"Benetel RAN550 and RAN650 are the commercial indoor and outdoor RUs used in the Split 7.2b testbed.","marker":"[19]"},{"why":"powerstat is the tool whose RAPL-based power model produces the DU and CU power estimates.","marker":"[24]"}],"fun_headline_variants":["Constant base load dominates O-RAN energy use","O-RAN power nearly load-independent, study finds","For O-RAN, power savings mean switching off, not scaling","Most O-RAN energy is base load, not traffic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the powerstat tool's estimate of server power (based on processor energy counters, not a wattmeter) is accurate; if background server activity biases that estimate, the reported DU/CU power numbers and the fitted quadratics would misrepresent real hardware draw.","fun_headline_variants_meta":{"raw":{"variants":["Constant base load dominates O-RAN energy use","O-RAN power nearly load-independent, study finds","For O-RAN, power savings mean switching off, not scaling","Most O-RAN energy is base load, not traffic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1436,"prompt_tokens":1086,"completion_tokens":350,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":702,"completion_tokens_details":{"reasoning_tokens":284}},"tokens_in":702,"tokens_out":350,"duration_ms":3924,"temperature":1.0,"reasoning_tokens":284,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:01:10.238753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the same DU and CU servers with a calibrated wattmeter while sweeping PRB utilisation from 0 to 100 percent; if total server power rises by much more than the few watts reported here, the claim that O-RAN power is largely load-independent would fail.","supporting_citations":[{"cited_title":"O-RAN Architecture Description v03.00,","cited_arxiv_id":null,"evidence_quote":"O-RAN Alliance architecture specification that defines the CU, DU, RU and functional splits framing the testbeds."},{"cited_title":"O-RAN Network Energy Saving Use Cases Technical Report v2.00,","cited_arxiv_id":null,"evidence_quote":"O-RAN energy-saving use cases (sleep modes, radio chain switch-off) that this paper's conclusion targets."},{"cited_title":"How much energy is needed to run a wireless network?","cited_arxiv_id":null,"evidence_quote":"Prior O-RAN energy-efficiency measurement platform that this study extends with split-specific and uplink/downlink measurements."},{"cited_title":"POET: A platform for O-RAN energy ef- ficiency testing,","cited_arxiv_id":null,"evidence_quote":"Prior CPU-level power measurement in virtualised RAN that supplies a comparison baseline for server-side power."},{"cited_title":"EARNEST: Experimental analysis of RAN energy with open-source software tools,","cited_arxiv_id":null,"evidence_quote":"Earlier O-RAN-based LTE RU/DU power measurements limited to uplink, motivating the both-link and split-specific measurements here."},{"cited_title":"Deep convolutional neural network assisted reinforcement learning based mobile network power saving,","cited_arxiv_id":null,"evidence_quote":"srsRAN software provides the DU and CU implementations used in the Split 8 testbed."},{"cited_title":"[Online]","cited_arxiv_id":null,"evidence_quote":"Benetel RAN550 and RAN650 are the commercial indoor and outdoor RUs used in the Split 7.2b testbed."},{"cited_title":"Available: https://www.rcrwireless.com/20210317/5g/ exploring-functional-splits-in-5g-ran-tradeoffs-and-use-cases-reader-forum","cited_arxiv_id":null,"evidence_quote":"powerstat is the tool whose RAPL-based power model produces the DU and CU power estimates."}],"review_version":1}