{"id":"1e52a596-f577-49dd-9a58-e142e5bc844a","arxiv_id":"2501.11814","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Contextual factors, including home location, sleepiness, valence, and multitasking, measurably affect how users respond to break interventions during infinite scrolling on social media.","lead":"A 7-day field study of 72 Android users found that contextual factors such as being at home, feeling sleepy, and emotional valence influence how quickly people stop infinite scrolling after a break reminder, and how much they resent the reminder. The results suggest that context-aware break reminders, timed for when people are tired or in distracting environments, could make digital well-being nudges more effective.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Responsiveness models condition on stopping: sessions where the intervention was ignored are censored, so the reported context–responsiveness interactions may be selection artifacts; survival or two-part analysis is required.","rationale":"The reader correctly identified the post-stop measurement timing as a key weakness, and the paper's own limitation section explicitly acknowledges missing data from participants who continued scrolling. My concern is distinct but closely related: even if context were measured at the exact moment of the intervention, the responsiveness outcome is only observed for sessions in which the user eventually stopped. This creates a censoring/selection problem that cannot be repaired by log-transforming the observed durations or by adding random intercepts. The interaction effects on responsiveness—the core of the abstract's claim—are therefore conditional on an outcome-dependent event, making them potentially biased. This is more load-bearing than the timing issue alone because it threatens the internal validity of the responsiveness analysis even under perfect self-report accuracy. The reader's verdict of CONDITIONAL remains appropriate; the authors should either provide a survival/two-part analysis of existing logs, collect data from ignored-intervention sessions, or substantially qualify the responsiveness findings. I set verdict_should_be to UNCHANGED because the recommended action is already embedded in the CONDITIONAL verdict, not because the concern is minor.","tokens_in":26687,"tokens_out":3182,"duration_ms":40010,"concrete_test":"Reanalyze with a Cox proportional-hazards or discrete-time survival model for time-to-stop, including all intervention events and coding sessions without a completed questionnaire as right-censored at the end of the session/study; if the intervention log does not record non-questionnaire events, collect these in a brief follow-up diary and then fit the survival model. Compare the At Home × Valence and Valence × Multitasking coefficients with Table 1 Model 4; if they shift materially or lose significance, the reported interactions are artifacts of conditioning on stopping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that contextual factors influence responsiveness rests on LMMs (Models 3–4, §4.5.3–4.5.5) where the dependent variable is time from intervention to stopping. By design, this variable is observed only when the participant actually stops scrolling and completes the post-stop questionnaire; intervention sessions that end in continued scrolling are missing, as acknowledged in §5.5. This is right-censoring/selection on the outcome: the missing sessions are exactly those with the longest (or undefined) responsiveness. If context—e.g., being at home with low valence—makes participants less likely to stop at all, then the observed interaction effects on responsiveness are not identified from a regression conditional on stopping. All five responsiveness interactions in §4.5.5 and Figure 5 are vulnerable. A log-transform LMM cannot correct this; the analysis needs a survival/censoring model or a two-part model for (a) whether the participant stops and (b) time-to-stop. Thus the abstract's statement that low valence at home 'slows down responsiveness' is not supported by the analysis as reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a 7-day longitudinal field study with N=72 participants who installed an Android app (InfiniteScape) that detects continuous infinite scrolling in six social media apps and, after 15 minutes, displays an intervention overlay. Once a participant stops scrolling, the app administers a questionnaire about reactance and six self-reported contextual factors (current activity, social situation, at home, multitasking, valence, sleepiness), and records responsiveness as the time from intervention to stopping. Linear mixed models are used to test main and interaction effects of these contextual factors on reactance and responsiveness. The main reported findings are that sleepiness reduces reactance and that several interactions involving valence, being at home, social situation, and multitasking affect responsiveness, e.g., low valence combined with being at home slows responsiveness. The authors conclude that context-aware interventions should account for these factors.","tokens_in":26851,"tokens_out":4414,"duration_ms":45404,"significance":"The paper addresses a timely and practically important question: whether intervention effectiveness during infinite scrolling depends on the user's context. The study design is ambitious for the field, combining real-world tracking of a specific interaction (infinite scrolling) with event-based experience sampling over seven days, and the authors make their code and anonymized data openly available. If the central claims were supported, the findings would usefully inform the design of context-aware digital well-being interventions. However, the current analysis has two load-bearing methodological limitations—outcome-dependent selection and post-hoc self-report of context—that substantially weaken the causal interpretation of the results. The paper is transparent about these limitations in Section 5.5, but the abstract and conclusions nonetheless assert causal contextual influences that the reported models cannot identify.","major_comments":[{"comment":"The responsiveness dependent variable is only observed for sessions in which the participant eventually stops scrolling and completes the questionnaire; sessions where the intervention is ignored and scrolling continues are missing. This is explicitly acknowledged in §5.5 ('we are missing data from those who continued scrolling'), but the linear mixed models in Models 3–4 condition on this selection. If context—for example, being at home with low valence—makes a user less likely to stop at all, the observed interaction effects on responsiveness in Table 1 and Figure 5 are not identified from a regression conditional on stopping. The log-transform of responsiveness does not address this issue. The manuscript needs a survival analysis treating continued scrolling as censored, or a two-part model separating the probability of stopping from the time-to-stop, to support the abstract's claim that context 'slows down responsiveness.'","section":"§4.1, §4.5.3–§4.5.5, §5.5"},{"comment":"The contextual factors (valence, sleepiness, location, multitasking, social situation, current activity) are self-reported after the participant has already stopped scrolling, not at the moment the intervention appears. The stopping event itself may change the user's affective state, sleepiness perception, or even location-related attention, so the post-stop reports may not reflect the state during the intervention. The paper acknowledges the reliance on self-report but does not discuss temporal validity of the context measures. This weakens all reported context–responsiveness associations as causal evidence; the findings should be framed as associations with post-stop context unless the questionnaire is moved to the intervention moment or validated against objective sensing.","section":"§4.1, §5.5"},{"comment":"The likelihood-ratio test comparing the interaction model (Model 4) with the main-effects model (Model 3) for responsiveness is not significant at the conventional level (χ²(60) = 78.265, p = .057), and the marginal R² of Model 4 is only 0.07. Despite this, the paper interprets five interaction coefficients from this model as robust evidence of contextual interplay. The model-comparison result and the small explained variance should temper the conclusion that context influences responsiveness; at minimum, the authors should report effect sizes and confidence intervals for the interaction terms and discuss the model-selection uncertainty explicitly.","section":"§4.5.6, Table 1"},{"comment":"The Valence × Social Situation [Strangers] interaction is estimated from very sparse data: only 0.43% of all observations are with strangers, and the authors themselves note that no data points were obtained for low and high valence in the strangers condition. A significant interaction estimated on such sparse cells is unstable and likely to be an artifact of a few influential observations. This interaction should be either removed from the set of reported findings or clearly flagged as exploratory and unreliable.","section":"§4.5.5, Figure 5a, Table 3"}],"minor_comments":[{"comment":"The model formula in the text uses 'Side Activity' while the table, questionnaire, and figures use 'Multitasking'; please use a single consistent term throughout.","section":"§4.5.3"},{"comment":"The sentence 'This observation aligns with the interaction effect between valence and multitasking (see Figure 5e)' appears to misrefer the figure; the valence × multitasking interaction is shown in Figure 5c, while Figure 5e shows the At Home × Valence interaction.","section":"§5.1.2"},{"comment":"The phrase 'high valance' should read 'high valence'.","section":"§5.1.2"},{"comment":"For the sleepiness main effect on reactance, the text reports t(917) = -3.40 and p < .001; please also report the exact p-value and a standardized effect size so readers can judge the practical importance of this effect, especially given the large sample size.","section":"§4.5.4"},{"comment":"The x-axis of panel (b) is labeled 'log(1+coef.)' but the response variable is log-transformed; please clarify how this back-transformation relates to the model coefficients, or plot the coefficients on a common scale for easier comparison.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of a CHI empirical studies paper and the authors are transparent about some limitations, which is commendable. However, the central causal claim is not supported by the analysis as presented because of outcome-dependent selection and the timing of the context questionnaire. These issues are potentially addressable if the raw logs allow a survival or two-part analysis and if the framing is revised to acknowledge the correlational nature of the post-stop context measures. I would not recommend rejection, as the underlying dataset and research question are valuable, but the revision must either supply the missing analysis or substantially soften the causal language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is the first field study I know of that directly tests whether context modulates how people respond to interventions during infinite scrolling. It deserves a serious referee, but the headline claims about responsiveness are weaker than the abstract suggests, because the outcome is only observed for sessions where the user eventually stopped.\n\nThe paper's strengths: a 7-day longitudinal deployment (N=72), a bespoke Android app that detects infinite scrolling via Accessibility Service, and open data and code. That is real work. The cleanest result is the main effect that sleepiness lowers reactance to the intervention, which is consistent with bedtime procrastination and does not depend on the censoring issue. The paper is also honest: it acknowledges the selection problem in §5.5 and the non-significant model comparison in §4.5.6.\n\nThe soft spots are in the responsiveness analyses. Responsiveness is time from intervention to stopping, measured only when the participant actually stops and completes the questionnaire. Sessions where the intervention is ignored are missing, so the LMMs in Models 3–4 are conditioned on stopping. That makes the context–responsiveness interactions in Figure 5 vulnerable to selection: if context affects the probability of stopping at all, the coefficients on time-to-stop are not identified from a regression on the completers. The authors acknowledge this, but the abstract still states that \"low valence coupled with being at home slows down responsiveness,\" which overreaches. A two-part model or survival analysis with censoring would fix this, or the claims need to be explicitly conditional on eventual stopping. Also, the interaction model for responsiveness does not significantly improve over the main-effects model (p=.057), yet the individual interactions are interpreted fairly strongly. And the \"strangers\" category has only four observations out of 927, so that interaction is on thin ice.\n\nThe post-stop self-report of context is a lesser issue, but it compounds the selection problem: the questionnaire reflects the state after stopping, not necessarily during the intervention.\n\nBottom line: the paper is a legitimate empirical contribution, best for HCI and digital wellbeing audiences. The sleepiness–reactance finding is likely solid; the responsiveness interactions need reanalysis or substantial qualification. Referee it, and push the authors on the censoring.","headline":"First real field data on context-aware scrolling interventions, but the responsiveness claims are selection-conditioned and need a reanalysis before they can be taken at face value.","tokens_in":27403,"tokens_out":2747,"would_cite":true,"duration_ms":28739,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Contextual factors such as being at home, sleepiness, and negative mood shape how effectively interventions interrupt infinite scrolling on social media.","keywords":["infinite scrolling","digital interventions","context-aware","field study","longitudinal study","reactance","responsiveness","digital well-being"],"falsifier":"A replication that measures context at intervention time — for instance with a second prompt right at the 15-minute mark or with passive sensors — and finds that post-stop self-reports differ systematically from the state during scrolling would invalidate the causal reading of the interaction effects. So would a controlled manipulation of location and sleepiness that shows no difference in responsiveness or reactance.","tokens_in":26489,"feed_emoji":"😴","tokens_out":6303,"duration_ms":56674,"temperature":0.7,"pith_summary":"This paper tries to show that the effectiveness of interventions against infinite scrolling on social media depends on the user's context, not just on the intervention itself. Through a 7-day field study with 72 participants who received a break prompt after 15 minutes of continuous scrolling, the authors find that being tired lowers reactance to the prompt, while being at home combined with low mood slows the user's response to it. The claim matters because current interventions are one-size-fits-all; if context systematically changes responsiveness and acceptance, interventions could be tailored to location, mood, and sleepiness to reduce regretful scrolling.","feed_headline":"Break prompts work worse at home, better when you're tired","feed_subtitle":"A 7-day field study of 72 users finds mood, location, and sleepiness change how people respond to anti-scrolling nudges.","key_machinery":"The central objects are the two dependent variables: reactance (measured with the Threat subscale of the Reactance Scale for Human-Computer Interaction, five Likert items averaged per observation) and responsiveness (the elapsed time from intervention display to the user stopping infinite scrolling, log-transformed for analysis). The argument is carried by linear mixed models with a random intercept per participant fitting main and interaction effects of six contextual factors: location (at home or not), current activity, valence, sleepiness, multitasking, and social situation. The key mechanism is the interaction term, since the paper's main claim is that context factors act together rather than independently.","core_discovery":"The paper's central discovery is that intervention effectiveness during infinite scrolling is contextually modulated: users who feel sleepier experience lower psychological reactance to the break prompt, and the time it takes to stop scrolling after the prompt is shaped by interactions among several contextual factors. Specifically, low valence together with being at home slows responsiveness, while multitasking during low valence hastens it, and multitasking's speeding effect is pronounced only when away from home. These results come from a 7-day in-the-wild study with 72 participants and 927 analyzed intervention events, using linear mixed models that treat the six contextual factors both as main effects and as interactions.","pith_inferences":["Editorial inference: if these interaction effects are stable, intervention systems could escalate prompt strength when the user is at home and in a negative mood, using GPS and sensors as proxies for the self-reported context.","Editorial inference: the non-significant improvement of the responsiveness interaction model (p = .057) relative to main effects suggests the specific interaction estimates for responsiveness may not replicate; a pre-registered study powered from these coefficients would settle that.","Editorial inference: a direct within-subjects comparison of a fixed 15-minute prompt versus a context-adaptive prompt (e.g., shorter delay when home and low mood) would translate the correlational findings into a causal design test."],"forward_implications":["Bedtime interventions should be easier to get accepted because sleepiness lowers reactance, but a single prompt will not stop scrolling, so escalating friction is needed.","Users at home with low mood will tend to ignore a break prompt, so interventions in that context need stronger design friction or a different trigger.","Because multitasking shortens response time in low-valence and out-of-home situations, nudging users toward a secondary activity may help disengagement.","The presence of multiple interaction effects supports designing interventions contextually rather than treating any single factor alone."],"supporting_citations":[{"why":"Supplies the infinite-scrolling session concept, the regret-based recruitment rationale, and the 15-minute latency for interventions, and motivates valence and activity as contextual factors.","marker":"[79]"},{"why":"Defines reactance for HCI and provides the Threat subscale used to measure the reactance dependent variable.","marker":"[23]"},{"why":"Justifies the event-based Experience Sampling Method over periodic sampling, which determines when the context questionnaire is delivered.","marker":"[99]"},{"why":"Self-determination theory is the basis for inviting only participants who regret scrolling, shaping the sample.","marker":"[83]"},{"why":"Supplies the five-area context model (location, social setting, internal state, situation, behavior patterns) that the six studied factors operationalize.","marker":"[74]"},{"why":"Contrasting evidence that users rarely personalise interventions by context, which the discussion weighs against its own findings.","marker":"[61]"},{"why":"Provides the Self-Assessment Manikin scale used to measure valence in the questionnaire.","marker":"[10]"},{"why":"Provides the Karolinska Sleepiness Scale used to measure the sleepiness factor.","marker":"[88]"}],"fun_headline_variants":["Sleepiness boosts acceptance of break prompts","At home, bad mood slows scrolling break response","Mood and location alter how well anti-scroll nudges work","Tired users are less resistant to scrolling interventions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the context questionnaire, which participants complete only after they have stopped scrolling, accurately reports their location, mood, sleepiness, and multitasking state at the moment the intervention appeared, even though stopping may change that state.","fun_headline_variants_meta":{"raw":{"variants":["Sleepiness boosts acceptance of break prompts","At home, bad mood slows scrolling break response","Mood and location alter how well anti-scroll nudges work","Tired users are less resistant to scrolling interventions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000316,"raw_usage":{"total_tokens":1738,"prompt_tokens":843,"completion_tokens":895,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":834}},"tokens_in":459,"tokens_out":895,"duration_ms":9238,"temperature":1.0,"reasoning_tokens":834,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:49:29.533285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that measures context at intervention time — for instance with a second prompt right at the 15-minute mark or with passive sensors — and finds that post-stop self-reports differ systematically from the state during scrolling would invalidate the causal reading of the interaction effects. So would a controlled manipulation of location and sleepiness that shows no difference in responsiveness or reactance.","supporting_citations":[],"review_version":1}