{"id":"6a2dcf1f-256a-4028-bd24-49b8d2a8fc97","arxiv_id":"2508.10060","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a four-arm trial with 7,711 Fitbit users, an RL-selected nudge algorithm increased average daily steps by 218 to 296 over other arms at one month and by 210 over control at two months.","lead":"This trial tested whether a reinforcement learning algorithm could choose personalized push notifications for physical activity using a Fitbit app, with 13,463 users randomized into four groups. It reports that the RL-personalized group took significantly more daily steps than control, random, and fixed nudge groups at one month, suggesting adaptive algorithms can improve mobile health interventions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Attrition of 43% of randomized participants could bias the RL-vs-control step difference; per-arm missingness and sensitivity analyses are needed.","rationale":"The paper's central claim is the RL group's significantly higher step counts versus all comparators at 1 month and versus control at 2 months. The abstract presents per-protocol/completer-style analyses with 7,711 analyzed out of 13,463 randomized, making attrition the most obvious threat. The reader identified this as the weakest assumption and gave a CONDITIONAL verdict. I agree: without transparency about who was excluded and why, the effect could be an artifact of differential missingness. I see no other concern from the abstract that is more load-bearing; multiple comparisons are partially mitigated by the fact that the 1-month p-values are small enough to withstand a simple Bonferroni correction, and the 2-month control comparison is secondary. The GEE result is similarly dependent on missingness assumptions. Therefore, I do not recommend changing the reader's verdict; the condition is already set correctly: the authors must provide missing-data analyses and retention data to confirm the effect is real.","tokens_in":973,"tokens_out":2154,"duration_ms":24943,"concrete_test":"Request per-arm retention rates and baseline characteristics (age, sex, baseline steps) for completers vs dropouts. Then re-estimate the 1-month RL-vs-control step difference using inverse probability weighting (IPW) with weights from a logistic model of observed primary outcome on arm, baseline steps, age, sex, and early wear adherence; also perform a tipping-point pattern-mixture analysis where missing step outcomes are imputed under differing assumptions about non-responders. If the +296-step effect (p=0.0002) shifts materially or loses significance under plausible missingness scenarios, the claim of RL superiority is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Abstract-only review. The primary analysis set includes 7,711 of 13,463 randomized participants (57%). The abstract reports no per-arm retention rates, no baseline table for the analysis set, and no missing-data methodology. If missingness is differential by arm and correlated with the outcome (e.g., less-active participants stop wearing the Fitbit differentially across arms), the observed +296-step RL-vs-control difference can be a selection artifact rather than a treatment effect. For instance, if the control arm retains more consistently inactive participants while the RL arm retains more responders, the comparison is biased in favor of RL. Conversely, if control dropouts are less active, bias could attenuate the effect. The GEE models cited do not by themselves address missingness unless they use robust weighting; typical GEE under missingness assumes missing completely at random. Given that nearly half the sample is excluded, this is the most load-bearing threat to the central claim. The reader's weakest_assumption correctly identifies this.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports results from the PEARL trial, a four-arm randomized controlled trial (control, random, fixed, and RL) testing whether a reinforcement-learning-based mHealth intervention can increase daily step counts. The abstract states that 13,463 Fitbit users were randomized, 7,711 were included in primary analyses, and the RL arm showed significantly higher average daily steps at 1 month relative to all other arms (e.g., +296 vs control, p=0.0002) and at 2 months relative to control (+210, p=0.0122), with a GEE analysis also favoring RL (+208, p=0.002). The central claim is that the RL algorithm provides a scalable, behaviorally-informed approach to personalizing physical activity nudges.","tokens_in":1126,"tokens_out":1430,"duration_ms":16465,"significance":"If the reported results are internally valid, this would be one of the first large-scale trials demonstrating that an RL-based JITAI can outperform both random and fixed nudge selection in a real-world mHealth setting. The randomized design, large sample, and theory-informed nudge bank are clear strengths, and the comparison against an active random arm and a fixed-logic arm is more informative than a simple waitlist control. However, the abstract alone cannot establish validity: nearly 43% of randomized participants are excluded from the primary analysis, and no missing-data or sensitivity information is provided. The significance of the finding therefore hinges on whether the attrition is non-differential and whether the analysis methods account for it. The paper's contribution would be substantial if these concerns are addressed in the full manuscript.","major_comments":[{"comment":"The abstract reports that 7,711 of 13,463 randomized participants were included in primary analyses, but it does not report per-arm retention, reasons for missingness, or any missing-data methodology. Since GEE under standard assumptions requires missing completely at random, and the outcome (daily steps from a Fitbit) is plausibly correlated with dropout, the +296-step RL-vs-control difference could be a selection artifact. The full manuscript must provide per-arm attrition tables, baseline characteristics for the analysis set, and sensitivity analyses (e.g., multiple imputation, pattern-mixture models) that bound the treatment effect under informative missingness. Without these, the central estimate is not credible.","section":"Abstract, primary analysis population"},{"comment":"All effect sizes are reported as point estimates with p-values, but no confidence intervals are given. For a trial with multiple comparisons across arms and time points (four arms at two time points plus GEE), the abstract also does not indicate whether any correction was applied as a primary confirmatory analysis. The reader cannot assess precision or whether the reported p-values would survive a conservative multiplicity control. The full manuscript should report 95% CIs for each contrast and specify a pre-specified confirmatory testing procedure.","section":"Abstract, statistical reporting"},{"comment":"The abstract does not describe how the RL policy was updated during the trial or whether treatment assignment groups were balanced at baseline for the analysis subset. Without a baseline table for the 7,711 analyzed participants, it is impossible to verify that randomization achieved balance after attrition. This is load-bearing because differential attrition across arms can induce confounding even in an RCT. The full manuscript should include a CONSORT flow diagram and a baseline table for both the randomized and analyzed populations.","section":"Abstract, randomization and intervention description"},{"comment":"The GEE result (+208 steps, p=0.002) is reported without the model specification, working correlation structure, standard error type (robust vs model-based), or covariates. Given that the 2-month contrast is only significant vs control and not vs random or fixed, the GEE analysis needs to be shown to be a pre-specified secondary analysis and not an ad hoc aggregated test. The full manuscript should clarify the analysis plan and whether the GEE model used all available observations or only completers.","section":"Abstract, GEE claim"}],"minor_comments":[{"comment":"The abstract reports the analysis-set demographics (mean age 42.1, 86.3% female, baseline steps 5,618.2) but not the corresponding demographics of all randomized participants. If the analyzed subset is not representative, the external validity of the findings is unclear; please report both populations.","section":"Abstract, demographic reporting"},{"comment":"The effect sizes (+296, +218, +238 steps at 1 month) are modest relative to baseline step count (~5,600). The manuscript should add a measure of clinical relevance (e.g., proportion meeting physical activity guidelines) to contextualize whether these differences are meaningful beyond statistical significance.","section":"Abstract, effect size interpretation"},{"comment":"The term 'first large-scale, four-arm randomized controlled trial' is a strong claim. Please specify what 'first' refers to (first RL JITAI? first four-arm mHealth RCT?) and provide a citation for prior work to substantiate novelty.","section":"Abstract, terminology"}],"recommendation":"major_revision","confidential_remarks":"This is an abstract-only review, so my confidence is necessarily low. However, the attrition issue is sufficiently load-bearing that the manuscript cannot be accepted in its current form; the full text may resolve my concern if it includes rigorous missing-data analysis. I would advise the editor to request the full manuscript and statistical appendix before making a final decision. The lack of confidence intervals and multiple-comparison handling in the abstract is troubling but could be addressed by better reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, This is a genuinely important trial to know about: 13,463 randomized, four arms, and a head-to-head of an RL-selected nudge policy against random and fixed nudge schemes plus a no-nudge control. That comparison is the real novelty. Most prior mHealth RL work has been pilot-sized or lacked a concurrent non-adaptive nudge arm. If the numbers hold up, this is the first large-scale RCT showing that RL personalization beats both no nudges and non-adaptive nudges on step counts. The abstract itself reports sensible-looking effects: +296 steps vs control at 1 month, with p=0.0002, and smaller but still significant differences vs random and fixed. The 2-month result is weaker (only vs control), which is worth knowing but not disqualifying. The soft spot is the one everyone will see first: 7,711 analyzed out of 13,463 randomized. The abstract gives no per-arm retention, no baseline table for the analysis set, and no missing-data methodology. GEE does not automatically handle missingness unless the model uses inverse probability weights or the missingness is genuinely MCAR. If dropouts were differential across arms, say, less-active participants in control more likely to stop wearing the Fitbit, the +296-step RL advantage could be a selection artifact. That is the load-bearing threat, and the abstract neither rules it out nor even names it. Also absent: confidence intervals, multiplicity correction across four arms and two timepoints, and any sensitivity analysis. The effect sizes are modest (3-5% of baseline steps) but within range of what mHealth trials typically report. I want to stress that I am not saying the result is wrong. Given the scale and design, this deserves a careful look. But as it stands, the preprint's abstract makes a strong causal claim without the information needed to evaluate the most obvious alternative explanation. The first question for the authors should be: what do the per-arm missingness patterns look like, and do the conclusions survive methods that handle informative dropout? Bottom line: this is worth a serious referee. The JITAI/RL-health community will want to see this in the literature, but the review should focus squarely on attrition, missing-data assumptions, and robustness. Bring it to reading group if you want to argue about whether the 'first large-scale' claim holds and how much weight to give self-reported Fitbit steps. Recommendation: send to peer review, but expect heavy revision before the claims are solid.","headline":"Large four-arm RCT with a useful RL-vs-random/fixed comparison, but the abstract alone cannot support the causal claim until attrition is addressed.","tokens_in":736,"tokens_out":1663,"would_cite":false,"duration_ms":30140,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A large randomized trial reports that a reinforcement-learning algorithm personalizing activity nudges increased daily steps by roughly 300 at one month versus control, beating random and fixed nudge arms.","keywords":["reinforcement learning","mobile health","just-in-time adaptive intervention","physical activity","randomized controlled trial","nudge personalization","Fitbit","step count"],"falsifier":"Re-analyze the trial using intention-to-treat principles with multiple imputation for missing step/wear data, and compare baseline characteristics and missingness across arms; the central claim would fail if the RL advantage over control and over the random and fixed arms disappears under this analysis.","tokens_in":1070,"feed_emoji":"🏃","tokens_out":1150,"duration_ms":32249,"temperature":0.7,"pith_summary":"The paper claims that a reinforcement-learning (RL) algorithm can personalize the content and timing of physical-activity nudges well enough to increase daily steps in a large real-world trial. In a four-arm study of 13,463 randomized Fitbit users, the RL arm gained significantly more average daily steps at one month than the control, random, and fixed-nudge arms, and sustained a significant gain over control at two months. The authors argue this is the first large-scale demonstration that an adaptive, behaviorally informed RL approach can outperform both no nudges and non-adaptive nudge selection. If the claim holds, it would support using adaptive algorithms to make digital health interventions more effective and scalable.","feed_headline":"Adaptive RL nudges beat random and fixed: +296 steps/day","feed_subtitle":"A four-arm trial of 13,463 Fitbit users shows learning-based nudge timing and content lift daily steps.","key_machinery":"A reinforcement-learning policy that, for each participant and time point, chooses which message from a 155-item theory-based nudge bank to send and when to send it. Unlike the random and fixed arms, the RL arm adapts to each user's step responses, which is the mechanism the paper credits for the improved engagement and step outcomes.","core_discovery":"The paper reports that a reinforcement-learning algorithm selecting nudges from a bank of 155 behavior-change-informed messages produced a significantly larger increase in average daily steps than three comparison arms: +296 steps versus the no-nudge control (p=0.0002), +218 steps versus random nudge selection (p=0.005), and +238 steps versus fixed, survey-based nudge selection (p=0.002) at one month. At two months, the RL arm remained significantly higher than control (+210 steps, p=0.0122), and generalized estimating equation models showed a sustained increase of +208 steps versus control (p=0.002). The authors interpret this as evidence that adaptive personalization of nudge content and t","pith_inferences":["The reported effect sizes are modest, so the practical health benefit depends on whether such gains persist beyond two months; a longer follow-up would test the durability implied by the RL mechanism.","Because the trial is abstract-only here, the key unresolved question is attrition: if the 7,711 analyzed participants differ systematically from the 13,463 randomized, the step differences may partly reflect who stayed engaged.","A natural next test is whether RL-selected nudges also improve retention or user satisfaction, which would explain the mechanism behind the step increases.","The same RL framework could be extended to other behavior-change targets such as sleep, medication adherence, or diet, where timing and content personalization matter similarly."],"forward_implications":["If the effect is real, adaptive RL nudge selection can be deployed on consumer wearable platforms without costly human coaching.","Personalizing the timing and content of nudges appears to add value beyond simply sending more messages, since the RL arm outperformed both random and fixed nudge arms.","A roughly 200–300 step-per-day gain, if sustained, could translate into meaningful weekly physical activity accumulation for sedentary users.","The trial's use of a large, four-arm design with a theory-informed nudge bank provides a template for evaluating JITAIs at scale.","The results suggest future mHealth interventions can treat message selection as a learning problem rather than a one-size-fits-all logic."],"supporting_citations":[],"fun_headline_variants":["RL personalization adds +296 steps/day over control in 13k-user trial","Adaptive RL nudges outperform random and fixed in step-count RCT","Reinforcement learning nudge timing boosts daily steps by 296","In 13,463 Fitbit users, RL-based nudges beat control by 296 steps","PEARL trial: RL-selected nudges increase steps vs all comparators"],"cache_read_input_tokens":3584,"weakest_assumption_plain":"The 7,711 participants included in the primary analyses represent all 13,463 randomized participants, with no differential dropout or missing step data that could bias the treatment effect.","fun_headline_variants_meta":{"raw":{"variants":["RL personalization adds +296 steps/day over control in 13k-user trial","Adaptive RL nudges outperform random and fixed in step-count RCT","Reinforcement learning nudge timing boosts daily steps by 296","In 13,463 Fitbit users, RL-based nudges beat control by 296 steps","PEARL trial: RL-selected nudges increase steps vs all comparators"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000662,"raw_usage":{"total_tokens":2951,"prompt_tokens":925,"completion_tokens":2026,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":1923}},"tokens_in":669,"tokens_out":2026,"duration_ms":13735,"temperature":1.0,"reasoning_tokens":1923,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:04:23.612813+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the trial using intention-to-treat principles with multiple imputation for missing step/wear data, and compare baseline characteristics and missingness across arms; the central claim would fail if the RL advantage over control and over the random and fixed arms disappears under this analysis.","supporting_citations":[],"review_version":1}