{"id":"94c0e80c-078d-4c6f-a00a-276bb1737302","arxiv_id":"2607.02196","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Additive regret is governed by the active weighted-mass exponent p near cutoffs: every policy incurs at least T^{1/2-1/(2p)} when p>1, and a sample-path marginal policy matches this (up to logs) while attaining O((log T)^2) when p=1, without non-degeneracy assumptions.","lead":"Online accept/reject resource allocation with continuous random rewards and sizes has regret controlled by a single exponent p measuring size-weighted mass near acceptance cutoffs. A sample-path marginal policy matches the resulting rates (polylog when p=1, polynomial T^{1/2-1/(2p)} when p>1) even under fluid degeneracy.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader's weakest_assumption correctly identifies Ass. 1 as the modeling condition on which both the product stability estimates and the matching lower-bound construction rest. That condition is stated explicitly (Defs. 2.2–2.4, Ass. 1) and is necessary for the claimed rates; it is not a soft spot in the argument itself. The continuous-consumption corner mechanism (independent V,β giving p=2; independent β,R giving p=1) is cleanly separated from dual multiplicity (App. C.1), and the capacity-local refinement (App. B) shows the polynomial price is a critical-capacity phenomenon, which strengthens rather than weakens the main claim. No formal verification or code is present, but the proofs are detailed and the matching lower bound is constructive. Therefore the ACCEPT / high-confidence verdict stands; no adjustment is warranted.","tokens_in":56250,"tokens_out":682,"duration_ms":6705,"concrete_test":"Independently re-derive the endpoint Hardy integral (Lem. A.7 / A.18) from the branch densities f(ω)≈ω^{α-1}, e(ω)≈ω^τ, λ^{br}(x)≈(x-e)^{γ-1} and the product cap d_ω μ(I_ω)≤r, without using the global active-mass exponent p; confirm that the integrated conditional curvature is still O(r log(e/r)). If the log factor fails or an extra power of r appears, the p>1 upper bound would need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (matching rates under Ass. 1 via SPM, Theorems 2.5 and 4.2) is internally consistent. The reader correctly flags Ass. 1 as the modeling regularity that makes the rates possible; that is not a hidden gap. The upper-bound chain (Bellman reduction Prop. 3.1 → marginal-to-cutoff Lem. 3.3 → projected-hull Jensen Lem. 3.4–3.5 → pre-Young stability Prop. 3.6 + Young absorption Lem. 3.7 → Prop. 3.8 → concentration + endpoint Hardy Lem. A.7 / Lem. 3.9) is self-contained and uses only the stated single-interval support, active-mass bounds, dominated cover, and contact-branch structure. The lower bound (Thm. 4.2) is a clean two-point dilemma on the same endpoint-contact family that realizes every p>1, so the polynomial exponent is tight inside the class. Dual non-uniqueness is handled by averaging cutoffs along the capacity sweep rather than selecting a fluid dual, which is the intended contribution. No load-bearing internal inconsistency or unstated assumption that would collapse the matching rates was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies online accept/reject resource allocation with continuous rewards and continuous scalar sizes that scale fixed type-specific consumption vectors, allowing fluid degeneracy. It introduces an active weighted-mass exponent p that measures size-weighted value-to-size ratio mass near active cutoffs (Assumption 1). Under this regularity, the sample-path marginal policy (SPM) attains O((log T)^2) regret when p=1 and O(T^{1/2-1/(2p)} polylog T) when p>1 (Theorem 2.5); a matching lower bound of order T^{1/2-1/(2p)} holds for every p>1 on one-resource endpoint-contact families (Theorem 4.2). The analysis reduces regret to Jensen gaps of pathwise marginals, averages cutoffs along capacity sweeps, controls the product of cutoff width and borderline mass via Hoffman-type stability plus Young absorption, and handles non-dominated corners by an endpoint Hardy estimate. Corollaries recover polylog rates for independent size/ratio with bounded density and T^{1/4} for independent uniform reward and size.","tokens_in":56639,"tokens_out":722,"duration_ms":7729,"significance":"If correct, the work cleanly separates dual non-uniqueness from the true price of continuous random consumption: the latter is the thinness of size-weighted ratio mass near cutoffs, quantified by a single exponent p. Matching upper and lower bounds without fluid non-degeneracy fill a genuine gap left by Jiang et al. (2025a), Besbes et al. (2025), Lueker (1998), and related CE/resolving analyses that either assume non-degeneracy or rule out continuous sizes. The technical toolkit (sweep averaging, active-mass product stability, endpoint Hardy) is reusable, and the capacity-local refinement (Appendix B) correctly isolates polynomial regret as a critical-capacity phenomenon. Full proofs are supplied end-to-end; the contribution is therefore a sharp, self-contained theory result of clear interest to online allocation and revenue management.","major_comments":[],"minor_comments":[{"comment":"The single-type sharpening remark after Examples 1–2 (that O((log T)^2) can be improved to O(log T) when there is only one scalar cutoff) is left as an informal note; a short formal statement or pointer would help readers who specialize to the classical knapsack.","section":null},{"comment":"Notation for the capacity-sweep parameter θ (Section 3) collides with the local contact exponent θ of Definitions 2.2–2.4; a one-line reminder at first use in each appendix would reduce cognitive load.","section":null},{"comment":"Appendix C.1’s dual calculations for the two running examples are useful; a brief cross-reference from the introduction examples to C.1 would make the dual-interval claim easier to verify on a first reading.","section":null},{"comment":"A few long sentences in Section 1.3 (roadmap) and the proof of Lemma 3.9 could be broken for readability; the logic is sound but dense.","section":null}],"recommendation":"accept","confidential_remarks":"The manuscript is long and technical but the central matching-rate claim is cleanly proved. Self-citations (Zhang 2026a,b) are complementary rather than load-bearing. Suitable for a top theory venue in OR/MS or learning theory; I see no reason to delay acceptance for presentation polish alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline: once both reward and consumption size are continuous, fluid degeneracy can force polynomial regret, and the paper identifies the exact order through a single active weighted-mass exponent p near the acceptance cutoff. For p=1 you get O((log T)^2); for p>1 the rate is T^{1/2-1/(2p)} up to logs, with matching lower bounds for every p>1. The sample-path marginal policy hits these rates without unique duals or non-degenerate fluid bases.\n\nWhat is actually new is the continuous-consumption corner mechanism. Prior logarithmic or constant-regret work either fixes size (Jiang et al., Besbes et al.), equalizes rewards (Arlotto–Xie), or assumes dual stability/non-degeneracy. Here both V and β can be continuous, duals may be non-unique, and the rates still go through. The two one-resource examples are the right illustration: same dual-optimal interval [0,1/2], but p=1 vs p=2 depending on whether you randomize the ratio or the reward independently of size. Appendix B also shows the polynomial hit is a critical-capacity effect—same primitives stay polylog at fixed interior capacities. That is useful.\n\nThe argument is careful and self-contained. Bellman reduction to Jensen gaps, averaging cutoffs along the capacity sweep instead of picking a dual, Hoffman-style stability for the active-mass product, Young absorption of the leftover length term, and a Hardy estimate when size-conditioned curvature fails pointwise domination. The lower bound is a clean two-point dilemma on the same endpoint-contact family. I do not see a load-bearing hole; Assumption 1 is the modeling price of admission and it is stated up front.\n\nSoft spots are real but proportionate. The policy needs Monte Carlo of the expected hindsight value, so it is not a cheap re-solving heuristic. Single-interval support is a modeling convention (split types if needed). The endpoint-contact machinery is technical, but the corollaries recover the common cases cleanly.\n\nThis is for people working on online allocation, NRM, and regret under degeneracy. It deserves a serious referee. I would bring it to reading group and cite it if I am writing in this area.","headline":"Matching rates under continuous size+reward and fluid degeneracy, pinned to one mass exponent p—clean theory that drops the usual non-degeneracy assumptions.","tokens_in":57189,"tokens_out":582,"would_cite":true,"duration_ms":15236,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C15","90B05","68W27","60G40"],"pacs":[],"model":"grok-4.5","headline":"When rewards and consumption sizes are both continuous, online allocation regret is set by how much size-weighted ratio mass sits near active cutoffs, not by fluid non-degeneracy.","keywords":["online resource allocation","additive regret","continuous random consumption","fluid degeneracy","value-to-size ratio","active weighted-mass exponent","sample-path marginal policy","network revenue management"],"falsifier":"Build a one-resource instance whose size-weighted ratio mass near the critical cutoff grows like width^p for a chosen p > 1 (e.g., independent uniform value and Beta size), run any online policy at capacity equal to mean total demand, and check whether regret stays o(T^{1/2-1/(2p)}) as T grows; the paper claims it cannot.","tokens_in":57165,"feed_emoji":"📦","tokens_out":800,"duration_ms":13296,"temperature":0.7,"pith_summary":"This paper studies irrevocable accept/reject decisions for scarce resources when each request has a continuous random reward and a continuous random size that scales a fixed type-specific consumption vector. It shows that additive regret against the fractional hindsight optimum is controlled by a single distributional quantity: the size-weighted mass of value-to-size ratios near the active acceptance cutoffs, measured by an active weighted-mass exponent p. When that mass grows linearly (p = 1), a sample-path marginal policy that prices capacity by the expected hindsight drop along a capacity sweep attains O((log T)^2) regret; when the mass is thinner (p > 1), every online policy must suffer order T^{1/2-1/(2p)} regret, and the same policy matches that polynomial rate up to logs. The result holds even when the fluid LP is primal-degenerate or dual-nonunique, so continuous random consumption can create a genuine polynomial price of degeneracy at critical capacities while non-critical capacities remain polylogarithmic for the same primitives.","feed_headline":"Cutoff mass, not dual uniqueness, sets online allocation regret","feed_subtitle":"Continuous sizes can force T^{1/2-1/(2p)} regret; a marginal policy matches it without non-degeneracy.","key_machinery":"The active weighted-mass exponent p (how fast size-weighted ratio mass accumulates near cutoffs) together with the sample-path marginal (SPM) policy, which prices each request by the average of pathwise offline bid-price cutoffs along the capacity segment it would consume; regret reduces to a stable product of cutoff width and borderline mass, closed by p.","core_discovery":"Additive regret in continuous-reward, continuous-size online resource allocation is governed by the active weighted-mass exponent p of the size-weighted ratio measures near acceptance cutoffs: the sample-path marginal policy achieves O((log T)^2) when p = 1 and O(T^{1/2-1/(2p)} polylog T) when p > 1, and for every p > 1 there exist instances on which every online policy incurs Ω(T^{1/2-1/(2p)}) regret, all without any fluid non-degeneracy assumption.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Cutoff mass exponent p sets continuous online allocation regret","Size-weighted mass near cutoffs governs regret without non-degeneracy","Active p-mass forces T^{1/2-1/(2p)} regret lower bounds","Sample-path marginal policy matches cutoff-mass regret rates","Regret in continuous-size allocation follows weighted ratio mass at cutoffs"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The size-weighted ratio mass near cutoffs must grow at least like a power of the band width, and conditional-on-size curvature can concentrate only in a limited power-law corner form that integrates after averaging over size.","fun_headline_variants_meta":{"raw":{"variants":["Cutoff mass exponent p sets continuous online allocation regret","Size-weighted mass near cutoffs governs regret without non-degeneracy","Active p-mass forces T^{1/2-1/(2p)} regret lower bounds","Sample-path marginal policy matches cutoff-mass regret rates","Regret in continuous-size allocation follows weighted ratio mass at cutoffs"]},"model":"grok-4.5","effort":"low","cost_usd":0.005044,"raw_usage":{"total_tokens":1498,"prompt_tokens":890,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":50440000,"prompt_tokens_details":{"text_tokens":890,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":512,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":890,"tokens_out":96,"duration_ms":4810,"temperature":1.0,"reasoning_tokens":512,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T08:20:31.543415+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Build a one-resource instance whose size-weighted ratio mass near the critical cutoff grows like width^p for a chosen p > 1 (e.g., independent uniform value and Beta size), run any online policy at capacity equal to mean total demand, and check whether regret stays o(T^{1/2-1/(2p)}) as T grows; the paper claims it cannot.","supporting_citations":[],"review_version":2}