{"id":"04ae6713-b3c1-4787-8289-e07b10833206","arxiv_id":"2405.03888","paper_version":5,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Authors introduce measurized MDPs as a measure-valued lifting of stochastic MDPs that generalizes the original setting and supports new constraints plus approximations via algebraic lifting and semicontinuous-semicompact analysis for average reward.","lead":"The paper lifts Markov Decision Processes to probability measure spaces, making states distributions over original states and actions stochastic kernels, creating deterministic measurized MDPs. This generalizes standard stochastic MDPs and enables new constraints and value approximations under average reward using semicontinuous frameworks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Whether lifted MDPs satisfy semicontinuous-semicompact assumptions after algebraic lifting is asserted but requires explicit verification for general (non-compact) state spaces.","rationale":"The reader's weakest_assumption directly identifies the same point. Because the full text is now available, the appropriate next step is the concrete verification above rather than a change in verdict category; the concern is technical rather than a contradiction with existing results.","tokens_in":1819,"tokens_out":339,"duration_ms":15832,"concrete_test":"In the section defining the algebraic lifting and the subsequent verification of H-L assumptions, extract the precise statement of the transition kernel on the measure space and check whether weak continuity (or the relevant semicontinuity) is proved without assuming compactness of the original state space; if the proof invokes an unstated compactness or tightness condition, the claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that measurized MDPs yield optimal Borel-measurable policies under milder conditions than Bertsekas-Shreve rests on the lifted processes satisfying Hernández-Lerma and Lasserre's semicontinuous-semicompact assumptions for both discounted and average-reward criteria. The state space becomes the (weakly) metrizable space of probability measures, whose compactness and continuity properties do not automatically inherit from the original MDP. The algebraic lifting procedure for MDPs with external shocks is claimed to produce non-deterministic measure-valued dynamics while preserving the required upper/lower semicontinuity of the reward and weak continuity of the transition kernel; this step is load-bearing because violation would invalidate the Borel-measurability and accessibility claims.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces measurized MDPs obtained by lifting standard MDPs to the space of probability measures, where states are probability measures on the original state space and actions are stochastic kernels. It claims that measurized MDPs generalize stochastic MDPs without loss of fidelity, enable incorporation of constraints and value-function approximations unavailable in the standard setting, and that an algebraic lifting procedure applied to MDPs subject to external random shocks produces non-deterministic measure-valued dynamics. By embedding the lifted processes in the semicontinuous-semicompact framework of Hernández-Lerma and Lasserre (rather than the universally measurable setting of Bertsekas-Shreve), the paper asserts that optimal Borel-measurable value functions and policies are obtained under milder, more verifiable assumptions for both discounted and average-reward criteria.","tokens_in":1987,"tokens_out":538,"duration_ms":16985,"significance":"If the lifting procedure is shown to preserve the required semicontinuity and semicompactness properties, the framework would supply a technically accessible route to measure-valued MDPs that yields Borel-measurable optima while supporting new modeling features such as explicit constraints on the measure-valued state.","major_comments":[{"comment":"The section introducing the algebraic lifting procedure: the claim that the lifted processes satisfy the semicontinuous-semicompact assumptions of Hernández-Lerma and Lasserre (upper/lower semicontinuity of the reward and weak continuity of the transition kernel) after the algebraic lifting is asserted but not accompanied by an explicit verification for general (non-compact) state spaces; the weak topology on the space of probability measures does not automatically inherit these properties from the original MDP, and this verification is load-bearing for the Borel-measurability and accessibility claims.","section":"Algebraic lifting procedure"},{"comment":"Paragraphs on framework choice and benefits: the central assertion that the measurized framework yields optimal Borel-measurable policies and value functions under milder conditions than Bertsekas-Shreve rests entirely on the lifted MDPs satisfying the semicontinuous-semicompact assumptions for both discounted and average-reward criteria; without a detailed argument or counterexample-free demonstration that the assumptions transfer, the comparison to the universally measurable framework cannot be substantiated.","section":"Framework choice and benefits"}],"minor_comments":[{"comment":"The abstract and introduction could more explicitly separate the deterministic measurized MDPs from the non-deterministic measure-valued processes that arise only after the algebraic lifting with external shocks.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and constructive comments on our manuscript. We address each major comment below and will incorporate revisions to strengthen the presentation of the algebraic lifting and framework comparison.","responses":[{"response":"We agree that an explicit verification is required for general (non-compact) state spaces, as the weak topology on probability measures does not automatically preserve the properties. The manuscript asserts the preservation via the algebraic nature of the lift but does not supply the detailed argument. In revision we will add a dedicated lemma and proof showing that if the original MDP satisfies the Hernández-Lerma–Lasserre semicontinuity and weak-continuity conditions, then the measurized process does as well, including the requisite arguments for the weak topology.","revision_made":"yes","referee_comment":"[Algebraic lifting procedure] The section introducing the algebraic lifting procedure: the claim that the lifted processes satisfy the semicontinuous-semicompact assumptions of Hernández-Lerma and Lasserre (upper/lower semicontinuity of the reward and weak continuity of the transition kernel) after the algebraic lifting is asserted but not accompanied by an explicit verification for general (non-compact) state spaces; the weak topology on the space of probability measures does not automatically inherit these properties from the original MDP, and this verification is load-bearing for the Borel-measurability and accessibility claims."},{"response":"The comparison to the Bertsekas–Shreve universally measurable setting indeed depends on the lifted processes satisfying the semicontinuous-semicompact assumptions. As indicated in our response to the first comment, the revised manuscript will contain the explicit verification for both criteria. This will substantiate that the assumptions are milder and more readily verifiable, thereby supporting the claimed advantages in accessibility and Borel measurability.","revision_made":"yes","referee_comment":"[Framework choice and benefits] Paragraphs on framework choice and benefits: the central assertion that the measurized framework yields optimal Borel-measurable policies and value functions under milder conditions than Bertsekas-Shreve rests entirely on the lifted MDPs satisfying the semicontinuous-semicompact assumptions for both discounted and average-reward criteria; without a detailed argument or counterexample-free demonstration that the assumptions transfer, the comparison to the universally measurable framework cannot be substantiated."}],"tokens_in":1507,"tokens_out":492,"duration_ms":20555,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper introduces an algebraic procedure that lifts any MDP to a deterministic process whose states are probability measures on the original space. This is new for the average-reward criterion and for Borel policies inside the Hernández-Lerma-Lasserre framework. It also shows how the lifted version can encode constraints and value-function approximations that do not fit the usual stochastic MDP setup. That part is useful and cleanly motivated from the abstract and the cited prior work on Bertsekas-Shreve and Hernández-Lerma-Lasserre.","headline":"The algebraic lifting to measurized MDPs lets you add constraints that standard MDPs block, but the claim that semicontinuous-semicompact properties survive the lift for general spaces is asserted rather than shown.","tokens_in":2505,"tokens_out":188,"would_cite":false,"duration_ms":10498,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"we cast lifted MDPs within the semicontinuous-semicompact framework of Hernández-Lerma and Lasserre... optimal Borel-measurable value functions and policies"},{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"novel algebraic lifting procedure... non-deterministic measure-valued MDPs can emerge from lifting MDPs impacted by external random shocks"}],"headline":"MDP measure-lifting and semicontinuous-semicompact optimality equations share no structural machinery with RS distinction-to-physics forcing.","alignment":"orthogonal","rationale":"The paper's central constructions (algebraic lifting of MDPs to MP(S) states with deterministic kernels F(ν,φ), inheritance of Hernández-Lerma–Lasserre Assumptions 2.1/2.2, Borel policies via measurable selection, CVaR/moment approximations on measures) operate entirely within classical dynamic programming on Borel spaces. RS theorems (reality_from_one_distinction, Jcost uniqueness via Aczél, phi_ladder constants, 8-tick/D=3 forcing in AlexanderDuality) derive spacetime and constants from a single non-trivial distinction with zero parameters; none of the paper's objects (stochastic kernels, sup-compact rewards, equilibrium problems for AR) appear in or parallel any RS module.","tokens_in":63218,"confidence":"high","tokens_out":368,"duration_ms":8132,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Lifting MDPs to probability measure states creates a deterministic generalization that supports constraints and approximations with Borel-measurable optimal policies.","keywords":["Markov decision processes","measurized MDPs","algebraic lifting","measure-valued states","semicontinuous assumptions","Borel measurable policies","average reward criterion"],"falsifier":"A concrete counterexample of an MDP with external shocks whose algebraic lifting satisfies the semicontinuous-semicompact assumptions but lacks a Borel-measurable optimal policy would falsify the accessibility and milder conditions claims.","tokens_in":2699,"feed_emoji":"","tokens_out":662,"duration_ms":22295,"temperature":0.7,"pith_summary":"The paper introduces measurized MDPs as deterministic processes on the space of probability measures, showing they generalize stochastic MDPs without losing fidelity. This allows embedding standard MDPs via an algebraic lifting that incorporates external shocks, leading to non-deterministic measure-valued processes. By using the semicontinuous-semicompact framework instead of universal measurability, the approach yields optimal Borel-measurable value functions and policies under milder, easier-to-verify assumptions for both discounted and average reward criteria. The framework further enables direct incorporation of constraints and value function approximations not feasible in the original MDP setting.","feed_headline":"Lifting MDPs to measures yields deterministic generalizations","feed_subtitle":"The framework supports constraints and approximations with Borel-measurable optima under milder assumptions than universal measurability.","key_machinery":"The algebraic lifting procedure that maps any MDP impacted by external random shocks to a non-deterministic measure-valued MDP analyzed under semicontinuous-semicompact assumptions.","core_discovery":"Measurized MDPs are deterministic MDPs whose states are probability measures on the original state space and whose actions are stochastic kernels; they generalize stochastic MDPs and, when the lifted processes satisfy semicontinuous-semicompact assumptions, admit optimal Borel-measurable value functions and policies under milder conditions than the universally measurable framework, for both discounted infinite-horizon and long-run average reward criteria. Any MDP can be algebraically lifted to such a process, and the setting permits constraints and approximations unavailable in standard MDPs.","pith_inferences":["This lifting could facilitate solving MDPs by embedding them into a space where deterministic optimization methods apply more readily.","It may enable new connections to problems like distributionally robust control by treating measures explicitly as states.","Applying the procedure to a simple MDP with additive shocks would test whether optimality is preserved exactly."],"forward_implications":["Optimal policies remain Borel-measurable rather than requiring universal measurability.","Constraints can be imposed directly on the measure-valued states.","Value function approximations become available in the lifted space.","The long-run average reward case is handled within the same framework with similar guarantees.","Non-deterministic measure-valued MDPs arise naturally from standard MDPs with shocks."],"fun_headline_variants":["MDPs lifted to measures become deterministic","Measurized MDPs use measure states and kernel actions","Deterministic MDPs from probability measure lifting","Borel measurable policies via measurized MDP framework","Lifted MDPs support constraints with milder assumptions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The lifted MDPs satisfy the semicontinuous-semicompact assumptions of Hernández-Lerma and Lasserre.","fun_headline_variants_meta":{"raw":{"variants":["MDPs lifted to measures become deterministic","Measurized MDPs use measure states and kernel actions","Deterministic MDPs from probability measure lifting","Borel measurable policies via measurized MDP framework","Lifted MDPs support constraints with milder assumptions"]},"model":"grok-4.3","cost_usd":0.004604,"raw_usage":{"total_tokens":2304,"prompt_tokens":710,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":46037000,"prompt_tokens_details":{"text_tokens":710,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1536,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":710,"tokens_out":58,"duration_ms":12220,"temperature":1.0,"reasoning_tokens":1536,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-24T01:22:28.469448+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A concrete counterexample of an MDP with external shocks whose algebraic lifting satisfies the semicontinuous-semicompact assumptions but lacks a Borel-measurable optimal policy would falsify the accessibility and milder conditions claims.","supporting_citations":[],"review_version":1}