{"id":"a546f6a5-3a3a-47c2-b8ee-d167c4953ef5","arxiv_id":"2606.27073","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces retention profiles r(x) and localization ratio L(P) for row-wise KL contraction analysis in finite Markov chains, with convexity-gap identity equating divergence gap to mutual information and constructions decoupling L(P) from spectral gap and mixing time.","lead":"The paper defines a state-indexed retention profile based on KL divergence from stationary distribution in finite Markov chains and introduces a localization ratio to separate localized versus global contraction issues. A smart generalist might read it for new analytical tools on mixing behavior that go beyond spectral methods in probability and sampling applications.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader verdict was UNVERDICTED solely from abstract-only access. Full text confirms the identity is standard and exact, the structural claims rest on symmetry plus explicit construction, and the finite-state setup is the minimal condition needed for the quantities to be defined. No derivation gaps or unsupported steps appear.","tokens_in":1914,"tokens_out":276,"duration_ms":47776,"concrete_test":"Re-derive the gap identity from the joint KL D(μ(x)P(x,y) || μ(x)π(y)) by splitting the log term into conditional and marginal parts; confirm equality holds for arbitrary μ, P, π (finite support).","verdict_should_be":"ACCEPT","load_bearing_attack":"The convexity-gap identity is an exact equality: E_μ[D(P(x,·)||π)] − D(μP||π) = I_μ(X;Y) follows from the joint KL expansion D(joint || μ ⊗ π) without further assumptions. Vertex-transitive symmetry forces constant r(x), hence L(P)=1 independently of mixing speed. The counterexample that L(P_n)→0 need not imply η_KL(P_n)/M_n→1 is supplied by explicit construction. All quantities are well-defined on finite irreducible chains with positive π.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript studies KL contraction in finite Markov chains from a row-wise perspective. It defines the retention profile r(x)=D_KL(P(x,·)||π)/log(1/π(x)) and the localization ratio L(P)=E_π[r]/M (with M=max r(x)). The central claims are (i) the exact convexity-gap identity E_μ[D(P(x,·)||π)]−D(μP||π)=I_μ(X;Y) together with a decomposition of the contraction ratio into entropy inflation and mutual-information penalty; (ii) a Cheeger-type lower bound on M; (iii) an explicit construction showing L(P_n)→0 need not imply η_KL(P_n)/M_n→1; and (iv) structural results including optimal tail bounds, Bhatia-Davis variance bound, two-sided spectral bounds with cubic correction, KL/Pinsker mixing-time bound, tensorization, and the fact that L(P) is decoupled from the spectral gap, Cheeger constant and mixing time (every vertex-transitive chain has L(P)=1 independently of mixing speed).","tokens_in":2045,"tokens_out":530,"duration_ms":38040,"significance":"If the identities and explicit construction hold, the work supplies a new row-wise lens on KL contraction that cleanly separates localized versus global obstructions via the retention profile and localization ratio. The exact convexity-gap identity (derived from the joint KL expansion) and the counter-example construction are concrete strengths; the decoupling from classical mixing invariants on vertex-transitive chains is also noteworthy. These tools could usefully complement existing contraction-coefficient analyses in information theory and Markov-chain theory.","major_comments":[],"minor_comments":[{"comment":"Abstract states that the numerical experiments are exploratory and not used as evidence for any universal classification; repeat this disclaimer explicitly in the main text (e.g., near the description of the test suite) to prevent misreading.","section":null},{"comment":"The definition of the SDPI ratio η_KL(P) is used throughout but is not restated in the introduction; add a one-sentence reminder of its standard definition for readers who may not recall the precise normalization.","section":"Introduction"},{"comment":"Notation for the averaged retention \bar r_π and the maximum M is introduced in the abstract; ensure both symbols are defined at first use in the body and that the interval [0,1] for L(P) is justified immediately after the definition.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment, accurate summary of the contributions, and recommendation of minor revision. The referee's description of the retention profile, localization ratio, convexity-gap identity, and decoupling results matches our manuscript closely. No specific major comments were listed in the report.","responses":[],"tokens_in":1540,"tokens_out":74,"duration_ms":21478,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main new elements are the retention profile r(x) = D_KL(P(x,·)||π) / log(1/π(x)), the localization ratio L(P) = average r over max r, the identity that E_μ[D(P(x,·)||π)] - D(μP||π) equals I_μ(X;Y), and the construction showing L(P_n)→0 need not force η_KL(P_n)/M_n →1. These follow from the joint KL expansion and standard definitions on finite irreducible chains.\n\nThe identity and decomposition are exact and organize the contraction ratio into entropy and mutual-information terms without extra assumptions. The Cheeger-type bound on M links retention directly to bottlenecks, and the vertex-transitive case correctly forces L(P)=1 regardless of mixing speed. The explicit construction on cardinality of high-retention states is verifiable and separates the quantities cleanly.\n\nThe listed tail bounds, Bhatia-Davis variance bound, spectral bounds with cubic term, and KL/Pinsker mixing bound are recorded as consequences. L(P) is shown decoupled from spectral gap, Cheeger constant, and mixing time on the vertex-transitive family and on a limited test suite where rank correlations are near zero.\n\nThe experiments are labeled exploratory and not used as evidence, so the zero correlations remain suggestive rather than conclusive. The abstract setup assumes finite chains with unique π and finite KLs, which is standard and sufficient for the claims.\n\nThis is for researchers working on SDPI ratios and information bounds for finite chains. The math rests on direct probability identities with no evident circularity or free parameters. It deserves a serious referee to verify the derivations and bounds in full.","headline":"The paper gives a row-wise view of KL contraction via retention profiles, an exact mutual-information identity for the convexity gap, and an explicit construction separating localization from normalized contraction.","tokens_in":2529,"tokens_out":425,"would_cite":false,"duration_ms":32322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The convexity gap between row-averaged KL divergence and D_KL(μP||π) equals the mutual information I_μ(X;Y) for finite Markov chains.","keywords":["finite Markov chains","KL contraction","retention profile","mutual information","localization ratio","SDPI","mixing time","Cheeger bound"],"falsifier":"Direct numerical verification on the paper's explicit sequence of chains P_n: compute L(P_n) and η_KL(P_n)/M_n for increasing n and check whether the ratio approaches 1 whenever L approaches 0.","tokens_in":2817,"feed_emoji":"📊","tokens_out":806,"duration_ms":50620,"temperature":0.7,"pith_summary":"Finite Markov chains admit a state-indexed retention profile r(x) obtained by normalizing each row's KL divergence to the stationary distribution π. A convexity-gap identity equates the difference between the π-averaged row divergences and the KL divergence of the pushed-forward measure to the mutual information between consecutive states under the initial distribution. This identity produces a decomposition of the KL contraction ratio into an entropy-inflation component and a mutual-information penalty. The work supplies an explicit sequence of chains demonstrating that the localization ratio L(P) can tend to zero without forcing the normalized contraction coefficient to approach one. It further records that L(P) equals one for every vertex-transitive chain and is uncorrelated with the spectral gap, Cheeger constant, and mixing time on the tested families.","feed_headline":"KL contraction gap equals mutual information in Markov chains","feed_subtitle":"Retention profiles decompose the contraction ratio and remain independent of spectral gap and mixing time.","key_machinery":"The retention profile r(x) together with the convexity-gap identity that sets the gap between averaged row KL divergences and D_KL(μP||π) equal to the mutual information I_μ(X;Y).","core_discovery":"For a finite Markov chain with unique stationary distribution π, the retention profile is defined by r(x) = D_KL(P(x,·)||π) / log(1/π(x)), the maximum retention is M = max r(x), and the localization ratio is L(P) = E_π[r]/M. The convexity-gap identity states that the difference between the row-averaged divergence and D_KL(μP||π) equals I_μ(X;Y). The resulting decomposition expresses the contraction ratio as an entropy term minus a mutual-information penalty. An explicit construction of chains P_n shows that L(P_n) → 0 does not force η_KL(P_n)/M_n → 1, with the number of high-retention states being the decisive factor rather than their total π-mass.","pith_inferences":["The decoupling of L(P) from classical invariants indicates that retention profiles track contraction features orthogonal to standard mixing metrics.","The counterexample construction implies that the cardinality of high-retention states, rather than their aggregate mass, governs whether localization controls the normalized contraction coefficient.","The same gap identity may yield sharper tail bounds or reverse-Markov inequalities once the finite-state restriction is relaxed."],"forward_implications":["The contraction ratio decomposes into entropy inflation minus a mutual-information penalty.","A Cheeger-type inequality supplies a lower bound on the maximum retention M in terms of the chain's bottleneck geometry.","Every vertex-transitive chain satisfies L(P) = 1 independently of its mixing speed.","L(P) is structurally independent of the spectral gap, Cheeger constant, and mixing time.","The retention profile tensorizes over product chains."],"fun_headline_variants":["Mutual information decomposes KL contraction in Markov chains","Localization ratio L decouples from mixing time in chains","High-retention state count controls KL contraction ratio","Convexity gap equals mutual information for row divergences"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The Markov chain is finite, possesses a unique stationary distribution π, and every row KL divergence to π is finite.","fun_headline_variants_meta":{"raw":{"variants":["Mutual information decomposes KL contraction in Markov chains","Localization ratio L decouples from mixing time in chains","High-retention state count controls KL contraction ratio","Convexity gap equals mutual information for row divergences"]},"model":"grok-4.3","cost_usd":0.005489,"raw_usage":{"total_tokens":2752,"prompt_tokens":898,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":54887000,"prompt_tokens_details":{"text_tokens":898,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1802,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":898,"tokens_out":52,"duration_ms":24836,"temperature":1.0,"reasoning_tokens":1802,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T03:37:44.540545+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Direct numerical verification on the paper's explicit sequence of chains P_n: compute L(P_n) and η_KL(P_n)/M_n for increasing n and check whether the ratio approaches 1 whenever L approaches 0.","supporting_citations":[],"review_version":1}