{"id":"741da4b5-9404-44bb-99f5-2d45a1ab91d2","arxiv_id":"2606.13604","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"An offline multi-agent RL policy learned from delayed marketplace data selects discrete multipliers for a production dispatch optimizer in food delivery, increasing batching efficiency in a switchback experiment without harming delivery quality.","lead":"The paper describes an offline-trained reinforcement learning policy deployed at DoorDash that selects multipliers to adapt dispatch optimizer weights using delayed marketplace signals such as delivery speed and courier utilization. A smart generalist might read it to see how real-world logistics systems can safely incorporate world feedback into decision policies without replacing existing combinatorial solvers.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Logged data + conservative regularizer may not ensure reliable value estimates under live distribution shift","rationale":"The reader's weakest assumption directly identifies the same load-bearing point. The production experiment is positive evidence but does not substitute for explicit checks on shift robustness; therefore the UNVERDICTED verdict with low confidence is appropriate and no adjustment is warranted.","tokens_in":1686,"tokens_out":303,"duration_ms":11668,"concrete_test":"Partition the logged data into subsets with measurable distribution shift (e.g., by time-of-day or merchant density quantiles); recompute the learned Q-values and policy performance on these subsets versus the original training distribution. If the value estimates or realized batching/quality metrics diverge by more than the reported experiment effect size, the reliability assumption is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the offline-trained policy (Double Q-learning + conservative regularizer on centralized logged data) producing reliable values at decentralized store-level execution. The production switchback is presented as validation, but the argument requires that the regularizer and logged distribution suffice to bound overestimation and prevent degradation when the live state-action distribution diverges (due to policy-induced changes in batching, timing, or merchant/courier behavior). No explicit quantification of distribution shift or additional off-policy evaluation on shifted subsets is described in the abstract; if this gap exists in the full text, the experiment alone does not close it.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a deployed multi-agent RL system at DoorDash for adapting dispatch objective weights in a three-sided food-delivery marketplace. A store-level policy is trained offline on centralized logged data using Double Q-learning with a conservative regularizer; the policy selects discrete multipliers for an existing combinatorial dispatch optimizer. The approach preserves production feasibility constraints. Validation occurs via a production switchback experiment in which the learned policy increases batching, reduces courier-side time costs, and does not degrade customer-facing delivery quality.","tokens_in":1811,"tokens_out":393,"duration_ms":16129,"significance":"If the central experimental result holds after addressing the noted gaps, the work supplies a concrete, production-validated template for offline RL under delayed, noisy, and coupled marketplace feedback. The centralized-training/decentralized-execution interface together with the conservative regularizer offers a practical route to objective-weight adaptation while respecting operational safeguards; this is a rare documented case of live economic-system feedback being used to adapt a deployed decision policy.","major_comments":[{"comment":"The production switchback experiment is presented as the primary validation that value estimates remain reliable under live deployment. However, the manuscript provides no explicit quantification of distribution shift (e.g., divergence in state-action occupancy between logged and policy-induced trajectories) nor additional off-policy evaluation on shifted subsets. Without such analysis, it is unclear whether the conservative regularizer alone suffices to bound overestimation when batching patterns, timing, or merchant/courier behavior change after deployment.","section":"Training and Experiment sections (abstract paragraph on training and experiment)"}],"minor_comments":[{"comment":"The abstract states results qualitatively; reporting the magnitude of batching increase, courier time-cost reduction, and any statistical significance or confidence intervals from the switchback would strengthen the claim.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the distribution-shift analysis. We address the single major comment below and will revise the manuscript accordingly.","responses":[{"response":"We agree that the current manuscript lacks explicit quantification of distribution shift between the offline logged data and the policy-induced trajectories observed after deployment. The conservative regularizer was introduced specifically to mitigate overestimation under such shifts, and the production switchback experiment provides empirical evidence that the learned policy improved batching without degrading customer metrics. To strengthen the presentation, the revised manuscript will add (i) a quantitative comparison of state-action occupancy measures (e.g., via KL divergence or total variation on discretized state features) between the training logs and the post-deployment trajectories collected during the switchback, and (ii) off-policy value estimates on temporally or geographically shifted subsets of the logged data. These additions will clarify the degree of shift encountered and the extent to which the regularizer bounds overestimation in practice.","revision_made":"yes","referee_comment":"[Training and Experiment sections (abstract paragraph on training and experiment)] The production switchback experiment is presented as the primary validation that value estimates remain reliable under live deployment. However, the manuscript provides no explicit quantification of distribution shift (e.g., divergence in state-action occupancy between logged and policy-induced trajectories) nor additional off-policy evaluation on shifted subsets. Without such analysis, it is unclear whether the conservative regularizer alone suffices to bound overestimation when batching patterns, timing, or merchant/courier behavior change after deployment."}],"tokens_in":1310,"tokens_out":331,"duration_ms":10171,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a working system that learns a store-level policy to pick discrete multipliers for the dispatch objective, trained offline on logged marketplace data with Double Q-learning and conservative regularization, then run decentralized. The production switchback shows gains in batching and courier time costs without hurting customer delivery quality.\n\nThe paper does a few things cleanly. It keeps the combinatorial assignment solver untouched, which respects operational constraints and makes deployment feasible. Centralized training on shared data with decentralized execution fits the multi-store setting. The switchback experiment provides direct evidence from live traffic rather than simulation.\n\nThe soft spots are around robustness. The abstract gives no numbers on distribution shift between the logged data and the policy-induced live distribution, no ablations on the regularizer, and no off-policy evaluation on shifted subsets. The stress-test concern about value estimates under live divergence therefore stands on the information given; the experiment alone does not close that gap. Without those details it is difficult to judge how much the conservative regularizer actually buys in practice.\n\nThis is for practitioners who run large-scale dispatch or marketplace systems and want concrete examples of closing the loop with operational feedback. It is not a theoretical advance in RL methods.\n\nThe work deserves peer review because it reports a real deployment with measurable production outcomes. An editor should send it out, with the expectation that referees will press on the distribution-shift handling and ask for more training diagnostics.","headline":"The paper describes a deployed offline RL setup at DoorDash for tuning dispatch objective weights from delayed feedback while keeping the existing optimizer, backed by a production switchback experiment.","tokens_in":2321,"tokens_out":363,"would_cite":false,"duration_ms":15505,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An offline-trained policy adapts dispatch weights in a live three-sided marketplace to increase batching while holding delivery quality steady.","keywords":["multi-agent reinforcement learning","delayed feedback","offline RL","dispatch optimization","marketplace adaptation","objective weights","conservative regularization"],"falsifier":"A new switchback period in which the policy either fails to increase batching, increases courier time costs, or degrades measured delivery quality relative to the baseline would falsify the claim that the learned policy improves the intended metrics.","tokens_in":2575,"feed_emoji":"🚚","tokens_out":471,"duration_ms":12489,"temperature":0.7,"pith_summary":"The paper shows how reinforcement learning can tune the objective weights inside an existing combinatorial dispatch optimizer rather than replace it. A policy learned from historical marketplace logs chooses discrete multipliers that shift the balance between customer delivery times and courier batching efficiency. Training uses a centralized value function updated with Double Q-learning targets plus a conservative regularizer to limit overestimation on unseen states. The resulting policy is executed store-by-store and was tested in a production switchback experiment. If correct, the approach demonstrates that delayed, noisy operational feedback can safely drive online adaptation of high-stakes decisions without violating feasibility constraints.","feed_headline":"Offline RL policy increases batching in live dispatch without harming delivery quality","feed_subtitle":"Trained on marketplace logs with conservative regularization, the policy cuts courier time costs in a production switchback test.","key_machinery":"A store-level policy that outputs discrete multipliers for the existing combinatorial assignment optimizer, trained via centralized offline value learning with Double Q-learning and a conservative regularizer to bound out-of-distribution overestimation.","core_discovery":"A store-level policy trained offline on logged marketplace data selects discrete multipliers for the dispatch optimizer's tradeoff between delivery quality and batching efficiency. Using centralized training of a shared value function with Double Q-learning targets and a conservative regularizer, the policy increases batching rates and reduces courier-side time costs in a production switchback experiment while leaving customer-facing delivery quality unchanged.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Offline RL increases dispatch batching","RL policy from logs increases batching","Delayed feedback trains RL for dispatch","Offline RL adapts dispatch weights","RL reduces courier time in dispatch"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Logged marketplace data plus the conservative regularizer produces value estimates that stay reliable under live deployment without large distribution shift or violation of production constraints.","fun_headline_variants_meta":{"raw":{"variants":["Offline RL increases dispatch batching","RL policy from logs increases batching","Delayed feedback trains RL for dispatch","Offline RL adapts dispatch weights","RL reduces courier time in dispatch"]},"model":"grok-4.3","cost_usd":0.010512,"raw_usage":{"total_tokens":4628,"prompt_tokens":631,"num_sources_used":0,"completion_tokens":46,"cost_in_usd_ticks":105124500,"prompt_tokens_details":{"text_tokens":631,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3951,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":631,"tokens_out":46,"duration_ms":33280,"temperature":1.0,"reasoning_tokens":3951,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T06:45:37.875384+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new switchback period in which the policy either fails to increase batching, increases courier time costs, or degrades measured delivery quality relative to the baseline would falsify the claim that the learned policy improves the intended metrics.","supporting_citations":[],"review_version":1}