Pith. sign in

REVIEW 2 major objections 2 minor 16 references

Mixed Differences-in-Q estimators based on Little's Law reduce bias and variance in A/B tests for scheduling policies.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-29 05:52 UTC pith:XCISVRWL

load-bearing objection The mixed DiQ estimators add a Little's Law mixing step to Farias et al. for queue A/B tests, but the interference from policy switches may still undermine the claimed bias and variance gains. the 2 major comments →

arxiv 2605.29641 v1 pith:XCISVRWL submitted 2026-05-28 stat.ME cs.PFmath.PR

Experimentation for Different Scheduling Policies on Queues: Mixed Differences-in-Q Estimators Based on Little's Law

classification stat.ME cs.PFmath.PR
keywords A/B testingscheduling policiesLittle's LawDifferences-in-QMarkovian interferencequeueingdata centersbias and variance reduction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper develops A/B testing methods for scheduling policies in queues using mixed Differences-in-Q estimators derived from Little's Law. These methods address Markovian interference that disrupts standard A/B tests in data center environments. Simulations under non-stationary arrivals, heterogeneous services, and delays show the approach cuts bias and variance. A reader would care because accurate testing helps deploy better workload distribution algorithms without performance risks.

Core claim

The central claim is that mixed Differences-in-Q estimators grounded in Little's Law enable valid and efficient A/B testing of different scheduling policies on queues. By leveraging Little's Law, the estimators remain effective despite Markovian interference between test and control, leading to significantly lower bias and variance than straightforward A/B tests.

What carries the argument

Mixed Differences-in-Q estimators based on Little's Law, which combine queueing relationships with difference estimation to correct for interference in A/B tests.

Load-bearing premise

Little's Law can be applied to build mixed Differences-in-Q estimators that remain valid under the Markovian interference in A/B tests for scheduling policies.

What would settle it

A simulation or experiment where the mixed estimators do not show reduced bias and variance compared to standard methods under Markovian interference would challenge the central claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The methods work robustly under non-stationary arrival rates and heterogeneous service rates.
  • They handle scenarios with communication delays.
  • They allow more reliable assessment of new scheduling algorithms prior to full deployment.
  • Extensive simulations confirm the reduction in bias and variance across tested scenarios.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • These estimators might apply to A/B testing in other systems with interference, such as networks or service queues.
  • Little's Law could inspire similar mixed estimators for other performance metrics in dynamic environments.
  • Improved experimentation could lead to faster iteration on scheduling policies in production data centers.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper proposes mixed Differences-in-Q estimators constructed via Little's Law, extending the DiQ estimator of Farias et al. (2022), for A/B testing of scheduling policies on queues subject to Markovian interference. It claims these mixed estimators achieve significant reductions in bias and variance relative to standard approaches, with supporting evidence from simulations under non-stationary arrivals, heterogeneous service rates, and communication delays.

Significance. If the validity of the mixed estimators under policy-induced interference is established and the reported bias/variance reductions are reproducible, the work would provide a targeted methodological advance for online experimentation in queueing systems, directly addressing a practical challenge in data-center scheduling deployment. The grounding in Little's Law offers a potentially parameter-light way to mix estimators if the ergodicity conditions can be shown to hold.

major comments (2)
  1. [§3 (estimator construction)] The central claim that the mixed DiQ estimators remain valid and reduce bias/variance under Markovian interference (abstract and §3) rests on an unexamined application of Little's Law (L = λW) to time-varying regimes created by policy switches. The manuscript does not derive or bound the resulting bias term when the steady-state assumption is violated by the A/B test design itself.
  2. [Simulation section (likely §4 or §5)] Simulations are described as demonstrating robustness (abstract), but no quantitative metrics, baseline comparisons (e.g., plain DiQ or naive A/B), confidence intervals, or statistical significance tests are reported for the claimed bias/variance reductions. This leaves the magnitude of improvement unverified and prevents assessment of whether reductions are independent of the fitted parameters shared with Farias et al. (2022).
minor comments (2)
  1. Notation for the mixed estimator (e.g., how the Little's Law weighting is applied to the Q-function differences) should be made fully explicit with an equation, including any additional assumptions on arrival/service processes.
  2. The abstract and introduction should clarify the precise relationship to the Farias et al. (2022) estimator to avoid any appearance of circularity in the variance reduction claim.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the detailed and constructive report. We address the two major comments below and will incorporate revisions to strengthen the formal justification and empirical reporting in the manuscript.

read point-by-point responses
  1. Referee: [§3 (estimator construction)] The central claim that the mixed DiQ estimators remain valid and reduce bias/variance under Markovian interference (abstract and §3) rests on an unexamined application of Little's Law (L = λW) to time-varying regimes created by policy switches. The manuscript does not derive or bound the resulting bias term when the steady-state assumption is violated by the A/B test design itself.

    Authors: We agree that the current §3 applies Little's Law to the mixed estimators without an explicit derivation of the bias induced by policy-induced time-variation. The paper relies on the ergodicity conditions referenced in the referee summary and the original Farias et al. (2022) framework, but does not bound the deviation. In revision we will add a new subsection deriving the bias term under Markovian interference and intermittent policy switches, including a bound that vanishes under the stated mixing and ergodicity assumptions. This will also clarify the conditions under which the mixed estimator remains unbiased or asymptotically unbiased. revision: yes

  2. Referee: [Simulation section (likely §4 or §5)] Simulations are described as demonstrating robustness (abstract), but no quantitative metrics, baseline comparisons (e.g., plain DiQ or naive A/B), confidence intervals, or statistical significance tests are reported for the claimed bias/variance reductions. This leaves the magnitude of improvement unverified and prevents assessment of whether reductions are independent of the fitted parameters shared with Farias et al. (2022).

    Authors: The simulations in the current version illustrate qualitative robustness across non-stationary arrivals, heterogeneous rates, and delays, but we acknowledge the absence of tabulated quantitative comparisons, confidence intervals, and formal tests. In the revised manuscript we will expand the simulation section with tables reporting bias and variance for the mixed DiQ, plain DiQ, and naive A/B estimators; 95% confidence intervals computed over repeated runs; and statistical significance tests. We will also include sensitivity checks with respect to the shared parameters from Farias et al. (2022) to demonstrate that the reported gains are not artifacts of those choices. revision: yes

Circularity Check

0 steps flagged

No circularity: extension of external DiQ estimator via standard Little's Law with simulation validation

full rationale

The paper cites Farias et al. (2022) for the base Differences-in-Q estimator (external to the author list) and grounds its mixed variant in the classical Little's Law (L = λW), a standard result independent of this work. Claims of bias/variance reduction are presented as outcomes of simulations across non-stationary arrivals, heterogeneous services, and delays rather than any definitional equivalence, fitted-parameter renaming, or self-citation chain. No equations or steps in the abstract or description reduce the central result to its own inputs by construction; the derivation remains self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claim rests on the validity of Little's Law in the A/B test setting and on the extension of the Differences-in-Q estimator; no free parameters or invented entities are identifiable from the abstract alone.

axioms (1)
  • domain assumption Little's Law holds for the queueing systems under the tested scheduling policies and A/B test interference.
    Invoked to ground the mixed Differences-in-Q estimators.

pith-pipeline@v0.9.1-grok · 5676 in / 1068 out tokens · 23068 ms · 2026-06-29T05:52:12.836981+00:00 · methodology

0 comments
read the original abstract

In data centers, tasks are dispatched to various servers to evenly distribute the workload. When a data center considers implementing a new scheduling algorithm, it typically conducts an A/B test prior to deployment to assess the real-world impact of this new method. However, a straightforward A/B test might be interfered with so-called ``Markovian'' interference. We utilized the Differences-in-Q estimator, as developed by Farias et al. (2022), and introduced mixed Differences-in-Q estimators grounded in Little's Law. We show that our A/B testing methods significantly reduce bias and variance when testing various scheduling policies. Extensive simulations were conducted under scenarios like non-stationary arrival rates, heterogeneous service rates, and communication delays. These simulations highlight the robustness and efficacy of our A/B testing approach.

Figures

Figures reproduced from arXiv: 2605.29641 by Nanshan Jia, Nian Si, Ramesh Johari, Zeyu Zheng.

Figure 1
Figure 1. Figure 1: The parallel queue service system For the ease of exposition, we now introduce some notations. We define qi to be the queue length (including the task in service) in the i-th server and si to be the proportion of servers with at least i pending tasks. Equivalently, we denote si = PN j=1 1{qj≥i} N i = 0, 1, 2, · · · . (1) 4 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Estimator performance in the homogeneous service-time setting for the power-of-3 [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Estimator performance in the heterogeneous service-time setting for the power-of-3 [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Estimator performance in the non-stationary arrival rate setting for the power-of-3 [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Estimator performance in the communication delay setting for the power-of-3 [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Estimator performance in the homogeneous service-time setting for the power-of-5 [PITH_FULL_IMAGE:figures/full_fig_p028_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Estimator performance in the doubly robust experiments. The Figure compares the MSE [PITH_FULL_IMAGE:figures/full_fig_p031_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Estimator performance in the homogeneous service-time setting for the MJSQ-0 [PITH_FULL_IMAGE:figures/full_fig_p031_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Estimator performance in the homogeneous service-time setting for the MJSQ-0 [PITH_FULL_IMAGE:figures/full_fig_p034_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Estimator performance in the homogeneous service-time setting for the JIQ-2/MJSQ-0 [PITH_FULL_IMAGE:figures/full_fig_p034_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

16 extracted references · 8 canonical work pages

  1. [1]

    Dynamic pull-based load balancing for autonomic servers

    Remi Badonnel and Mark Burgess. Dynamic pull-based load balancing for autonomic servers. In NOMS 2008-2008 IEEE Network Operations and Management Symposium, pages 751–754. IEEE,

  2. [2]

    Load balancing in parallel queues and rank-based diffusions.arXiv preprint arXiv:2302.10317,

    Sayan Banerjee, Amarjit Budhiraja, and Benjamin Estevez. Load balancing in parallel queues and rank-based diffusions.arXiv preprint arXiv:2302.10317,

  3. [3]

    arXiv preprint arXiv:2209.00197 , year=

    Yuchen Hu and Stefan Wager. Switchback experiments under geometric mixing.arXiv preprint arXiv:2209.00197,

  4. [4]

    Clustered switchback experiments: Near-optimal rates under spatiotemporal interference.arXiv preprint arXiv:2312.15574,

    Su Jia, Nathan Kallus, and Christina Lee Yu. Clustered switchback experiments: Near-optimal rates under spatiotemporal interference.arXiv preprint arXiv:2312.15574,

  5. [5]

    arXiv preprint arXiv:2302.12093 , year=

    Shuangning Li, Ramesh Johari, Stefan Wager, and Kuang Xu. Experimenting under stochastic congestion.arXiv preprint arXiv:2302.12093,

  6. [6]

    doi: 10.1287/moor.2019.1042

    ISSN 1526-5471. doi: 10.1287/moor.2019.1042. URL http://dx.doi.org/10.1287/moor.2019.1042. Tu Ni, Iavor Bojinov, and Jinglong Zhao. Design of panel experiments with spatial and temporal interference.Available at SSRN 4466598,

  7. [7]

    Load balancing under strict compatibility constraints

    Daan Rutten and Debankur Mukherjee. Load balancing under strict compatibility constraints. Mathematics of Operations Research, 48(1):227–256, 2023a. Daan Rutten and Debankur Mukherjee. Mean-field analysis for load balancing on spatial graphs. In Abstract Proceedings of the 2023 ACM SIGMETRICS International Conference on Measurement and Modeling of Compute...

  8. [8]

    Optimal experimental design for staggered rollouts.arXiv preprint arXiv:1911.03764,

    Ruoxuan Xiong, Susan Athey, Mohsen Bayati, and Guido Imbens. Optimal experimental design for staggered rollouts.arXiv preprint arXiv:1911.03764,

  9. [9]

    Data-driven switchback design

    Ruoxuan Xiong, Alex Chin, Mohsen Bayati, and Sean Taylor. Data-driven switchback design. preprint, 2023a. URLhttps://www.ruoxuanxiong.com/data-driven-switchback-design.pdf. Ruoxuan Xiong, Alex Chin, and Sean Taylor. Bias-variance tradeoffs for designing simultaneous temporal experiments. InThe KDD’23 Workshop on Causal Discovery, Prediction and Decision, ...

  10. [10]

    Little’s law, first proposed in Cobham [1954], Morse

    22 A Little’s Law In this section, we briefly introduce and summarize Little’s law in queueing theory. Little’s law, first proposed in Cobham [1954], Morse

  11. [11]

    In general, Little’s Law studies a parallel-server system over finite time period

    and proved in Little [1961], Whitt [1991], Kim and Whitt [2013], has been studied under different situations for various purposes and in a wide range of fields. In general, Little’s Law studies a parallel-server system over finite time period. Suppose the queue length in the current parallel-server system at time t isl(t). The average queue length over ti...

  12. [12]

    For a more detailed description and overview of the theory and applications of Little’s Law, one may refer to Little [2011], Wolff [2011]

    This leads to the general version of Little’s Law, as stated in equation (29). For a more detailed description and overview of the theory and applications of Little’s Law, one may refer to Little [2011], Wolff [2011]. B Proof of Proposition 1 Proof of Proposition 1.In order to prove that the response-time-based and queue-length-based GTE are the same, it ...

  13. [13]

    Therefore, we haveC 0,q(λ) =E s∼π0 [cq(s)] andC 0,w(λ) =E s∼π0 [cw(s)], where π0 is the stationary distribution under the control policy

    and Luczak and McDiarmid [2006], the marginal distribution of the process will converge to the stationary distribution asT→ ∞, due to the ergodicity. Therefore, we haveC 0,q(λ) =E s∼π0 [cq(s)] andC 0,w(λ) =E s∼π0 [cw(s)], where π0 is the stationary distribution under the control policy. Now consider a parallel queue service system that has initial distrib...

  14. [14]

    λIndex GTE Naive qDQ wDQ mixDQ Est

    Table 9: Power-of-5/3 experiments in the homogeneous service time setting. λIndex GTE Naive qDQ wDQ mixDQ Est. 0.092 0.096 0.092 0.092 0.091 λ= 0.5 Std. Dev. 2.21×10 −4 0.022 0.005 0.003 Std. Err. 1.80×10 −4 0.023 0.006 0.003 Est. 0.132 0.137 0.133 0.132 0.132 λ= 0.6 Std. Dev. 2.69×10 −4 0.029 0.010 0.004 Std. Err. 2.16×10 −4 0.029 0.010 0.004 Est. 0.180 ...

  15. [15]

    Since the estimated queue length ˆcq only returns marginal observation of the system, the Markov property is no longer kept well

    = nT nC +n T .(50) Next, we introduce the regressive approximation of the Q-functions. Since the estimated queue length ˆcq only returns marginal observation of the system, the Markov property is no longer kept well. Hence, we adopt a linear regression model that takes the past five costs as independent variables, i.e, QREG j =β+ 4X u=0 βucq,j−u .(51) The...

  16. [16]

    0.095 0.049 0.101 0.098 0.095 λ= 0.5 Std

    Table 12: MJSQ-r/Power-of-dExperiments withr= 0.95,d= 1 λIndex GTE Naive qDQ wDQ mixDQ Est. 0.095 0.049 0.101 0.098 0.095 λ= 0.5 Std. Dev. 7.54×10 −4 0.081 0.040 0.004 Std. Err. 8.60×10 −4 0.077 0.038 0.003 Est. 0.174 0.072 0.159 0.165 0.173 λ= 0.6 Std. Dev. 1.10×10 −3 0.112 0.065 0.006 Std. Err. 1.06×10 −3 0.106 0.063 0.006 Est. 0.349 0.110 0.293 0.310 0...