{"id":"644843f6-88ea-4aff-9d45-42a7dd55a7ac","arxiv_id":"2507.03520","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A public benchmark dataset of ten weeks of Fitbit sleep data from 139 hospital workers, with sleep analyses and machine learning baselines for sleep quality, demographics, and sleep stages.","lead":"Researchers released a public dataset of sleep tracking from 139 hospital workers over ten weeks, with more than 6,000 sleep sessions and surveys about sleep quality. It is meant as a benchmark for building and testing machine learning models that predict sleep stages, sleep quality, and demographics from wearable data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Between-group sleep comparisons assume Fitbit bias is shift-invariant; the cited validation study does not establish this, and if false, the REM findings in Table V and Fig. 5 could be device artifacts.","rationale":"The reader's weakest assumption correctly identifies Fitbit stage-label accuracy as the core vulnerability. My stress-test focuses on the specific form of that assumption that the paper actually relies on: not that Fitbit is accurate, but that its errors are shift-invariant. This is the single most load-bearing issue because the paper's headline behavioral analyses (night vs. day shift differences in REM and sleep-stage transitions) and the benchmark ground-truth labels are both downstream of it. The paper's own limitation section concedes the device overestimates REM, yet asserts without empirical support that the error cancels between groups. Other issues—the 5,000/6,000 recording count discrepancy, the PSQI being collected at baseline rather than concurrently with sleep episodes, and the lack of confidence intervals—are real but secondary; they affect presentation and the interpretation of one benchmark task, whereas the shift-invariance assumption affects the validity of the main sleep-behavior conclusions and the meaningfulness of the stage-label benchmark. The reader's CONDITIONAL verdict already accommodates this concern, so no change in verdict is needed; the dataset may still be a valuable public resource, but the shift-related conclusions should be tempered until the assumption is tested. The proposed check is feasible because a naturalistic validation study of the same device in shift workers already exists, making a stratified re-analysis a concrete and decisive test.","tokens_in":15030,"tokens_out":4955,"duration_ms":58318,"concrete_test":"Re-analyze the polysomnography validation data from Stucky et al. [30] (or obtain the underlying raw data) to compute Fitbit Charge 2 minus PSG bias for REM duration, REM percentage, and REM-to-light transition probabilities, stratified separately for day-shift and night-shift workers. Fit a model with device type-by-shift interaction; if the interaction is significant (e.g., REM overestimation differs between day and night sleep episodes), then the consistency assumption in Section VIII-C fails and the between-shift REM findings in Table V and Fig. 5 are not interpretable as behavioral differences. If the interaction is not significant, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central behavioral claims—night-shift workers have lower REM minutes and higher REM-to-light transition probabilities (Table V, Fig. 5)—rest on the assumption stated in Section VIII-C that Fitbit Charge 2's systematic errors are 'likely to be consistent across participants.' This assumption is load-bearing because the device is known to overestimate REM (Stucky et al. [30]), and the paper provides no evidence that the overestimation magnitude is independent of shift type. Night-shift workers sleep at different circadian phases, have more variable sleep-wake schedules, and sleep during daytime; these factors can plausibly change how the device's proprietary algorithm misclassifies sleep stages. If REM overestimation is larger (or smaller) for daytime sleep, then the observed day-night differences in REM duration and in REM-to-light transition probabilities could be measurement artifacts rather than genuine sleep-behavior differences. The same issue propagates into the benchmark tasks that use device stage labels as ground truth: differential label noise across shifts makes the task labels inconsistent. The paper acknowledges the device limitation but does not test the consistency assumption, and the cited naturalistic validation study is not shown to report the required shift-stratified bias. This is the weakest link in the argument that the dataset and its analyses support valid sleep-behavior conclusions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes TILES-2018 Sleep Benchmark, a longitudinal wearable sleep dataset collected from 139 hospital employees over 10 weeks using Fitbit Charge 2 devices. The dataset includes continuous heart-rate recordings during sleep, device-provided sleep stages (wake, light, deep, REM), sleep metadata, participant demographics, and baseline PSQI self-reports. The authors combine this new dataset with the earlier TILES-2018 dataset to form a 349-participant \"Combined TILES Sleep\" dataset, and use it for behavioral analyses comparing day- and night-shift workers (sleep duration, sleep stages, sleep onset/wake-up variability, and sleep-stage transition probabilities). They also present machine learning benchmarks for sleep-stage classification from heart rate, self-reported PSQI prediction, and demographic classification, with models trained on the earlier TILES-2018 data and evaluated on the new benchmark data. The abstract claims over 6,000 unique sleep recordings, while the main analysis section reports 6,012 unique main sleep records and the conclusion states over 5,000.","tokens_in":15298,"tokens_out":6310,"duration_ms":69157,"significance":"If the data release is implemented as described, this is a useful public resource for wearable sleep research in naturalistic settings. Its strengths include the longitudinal 10-week design, a shift-working hospital population that is underrepresented in open sleep datasets, concurrent heart-rate and device sleep-stage data, and the availability of PSQI and demographic metadata. The explicit train/evaluation split between the earlier TILES-2018 data and the new benchmark set is a sound design that avoids circular evaluation, and the paper is transparent about several limitations, including the lack of daily sleep-quality assessments and the absence of noisy-data mitigation experiments. However, the behavioral findings and benchmark labels rely heavily on Fitbit-derived sleep stages, and the paper's assumption that the device's known REM overestimation is consistent across day- and night-shift workers is not demonstrated. The count inconsistency and missing uncertainty quantification further reduce the current reliability of the reported results. With revision, the dataset could become a valuable community benchmark.","major_comments":[{"comment":"The central behavioral finding that night-shift workers have lower REM minutes and higher REM-to-light transition probabilities rests on the assumption stated in Section VIII-C that the Fitbit Charge 2's systematic errors are 'likely to be consistent across participants.' The cited validation study [30] is not described as reporting shift-stratified bias, and the manuscript's own §VI-B shows night-shift participants have much more variable sleep timing (within-subject SD of sleep onset 5.6 h vs 2.0 h), so daytime and circadian-phase-shifted sleep may plausibly be misclassified differently by the device. Without a shift-stratified validation or a sensitivity analysis under alternative label-noise assumptions, the Table V and Fig. 5 REM results could be device artifacts rather than genuine shift differences. Please either provide such evidence or explicitly relabel these results as exploratory and elevate this issue from 'likely consistent' to a primary limitation.","section":"Section VIII-C (with §VI-C, Table V, Fig. 5)"},{"comment":"The number of sleep recordings is reported inconsistently: the abstract says 'over 6,000 unique sleep recordings,' §V-C reports '6,012 unique main sleep records' after selecting participants with more than ten main sleeps, and §IX says 'over 5,000 unique sleep samples.' The released dataset size must be stated unambiguously, including which subset is public (e.g., all main sleeps vs only high-quality entries) and how the counts in the abstract and conclusion relate to the 6,012 figure. This is a basic reproducibility requirement for a dataset paper.","section":"Abstract; §V-C; §IX"},{"comment":"All benchmark results are reported as point estimates of macro-F1, ROC-AUC, and accuracy, without confidence intervals or significance tests, even though several subgroup comparisons involve very small sample sizes (e.g., n=28 night-shift participants in Fig. 6(b)). The discussion's statements such as 'TimesNet outperforms SleepNet' for PSQI and demographics, and the age-group difference in sleep-stage F1, are therefore not established. Please add bootstrap confidence intervals or repeated-run variability estimates, and clearly state the participant counts and sleep-recording counts underlying each estimate.","section":"Tables VI, VII and Fig. 6(b)"},{"comment":"The analysis pipeline uses several hand-chosen thresholds whose influence is not examined: the >10 main-sleep inclusion criterion, the >90% heart-rate coverage definition of 'high quality,' the PSQI binarization at 7, and the age binarization at 40. The PSQI cutoff of 7 is particularly consequential because the standard threshold in the cited reference [19] is greater than 5, not 7; the paper should justify the 7 cutoff or show that results are robust to it. For a benchmark intended for reuse, threshold sensitivity should be reported.","section":"§V-C, §VII-B"}],"minor_comments":[{"comment":"The caption states that approximately 70% of participants have 'more than 30 sleep hours,' but the table reports counts of main sleep entries, not sleep hours; please correct the caption or the table.","section":"Fig. 3(a) caption"},{"comment":"In §V-C, 'Table 3a' is actually a panel of Fig. 3, and in §VII-A, 'Table 6a' is a panel of Fig. 6; the cross-references should be corrected to the figure panels.","section":"§V-C and §VII-A"},{"comment":"The table lists the device as 'Fitbit Charge2' while the text uses 'Fitbit Charge 2'; please standardize the naming.","section":"Table I"},{"comment":"The transition analysis is described in the text as comparing 'nurses' (Fig. 5), whereas the rest of the sleep analyses concern all hospital employees; clarify whether the analysis is restricted to nurses and, if so, state the sample sizes.","section":"§VI-C"},{"comment":"The p-values reported in the Fig. 6(b) table are not linked to any described statistical test or adjustment for multiple comparisons; please state the test, the number of recordings, and whether the p-values are corrected.","section":"Fig. 6(b)"}],"recommendation":"major_revision","confidential_remarks":"This is a legitimate dataset contribution, and I would not reject it merely because the sleep-stage labels come from a consumer wearable. However, the manuscript's behavioral claims currently overreach what the device-accuracy evidence supports, and the count inconsistency should have been caught before submission. After the requested revisions, the paper could be suitable for publication in a venue that accepts dataset papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the TILES-2018 Sleep Benchmark paper. The main thing you should know: the dataset release is real and valuable. A new 139-person, 10-week Fitbit Charge 2 sleep collection from hospital workers, with heart rate, device sleep stages, demographics, and PSQI, is now public, and it can be combined with the earlier TILES-2018 cohort to give 349 participants and roughly 15,000 sleep recordings. That is a solid resource for sleep informatics and wearable ML. The benchmark tasks (sleep stage classification, PSQI prediction, demographics classification) use the new data as a held-out test split, trained on the earlier cohort. Clean setup.\n\nThe paper does well documenting protocol, data format, and demographics. The self-citations to TILES-2018 are appropriate since this is the holdout extension. The benchmark results are honestly reported: a simple LSTM beats TimesNet on sleep staging, the ensemble does best, and the LLM zero-shot numbers are poor, which matches expectations.\n\nNow the soft spots. The most load-bearing is the device-bias assumption. The paper's own limitation section cites Stucky et al. showing the Fitbit Charge 2 overestimates REM. The behavioral findings that night-shift workers have less REM and higher REM-to-light transition probability (Table V, Fig. 5) depend on the assumption that this overestimation is consistent across shifts. That is plausible but not demonstrated. The cited validation study is naturalistic, but I have not seen shift-stratified bias reported there, so the cancellation assumption is unsupported. If REM misclassification differs for daytime sleep or irregular schedules, those findings could be device artifacts. Note that in the new dataset alone, REM minutes difference is p=0.09, so the combined dataset carries the weight. The transition-probability results are subtler and therefore more vulnerable. This does not kill the dataset contribution, but the behavioral conclusions should be softened until validated against PSG or another reference.\n\nOther issues are minor: the count inconsistency (abstract says over 6,000, conclusion says over 5,000), benchmark point estimates without confidence intervals, and hand-chosen thresholds (PSQI cutoff 7, age cutoff 40, inclusion at 10 recordings). The paper does acknowledge missing per-night labels and noisy-data handling as limitations.\n\nBottom line: this is a dataset paper, and the dataset is the contribution. The analyses are illustrative, and the REM concern, while real, is secondary given the framing. I would bring it to a reading group for the dataset and cite it if I worked in wearable sleep. It deserves peer review with a request for revision: soften the shift-work claims, add intervals, and address the shift-invariance assumption directly.","headline":"A genuinely useful longitudinal wearable sleep dataset release, with benchmarks that work as demonstrations; the shift-work REM findings rest on an unverified device-bias assumption.","tokens_in":15793,"tokens_out":1698,"would_cite":true,"duration_ms":20867,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the TILES-2018 Sleep Benchmark dataset, 6,012 wearable sleep recordings from 139 hospital employees over ten weeks, and uses it to show that night-shift hospital workers sleep less and have more fragmented REM sleep…","keywords":["wearable sleep dataset","sleep stage classification","Fitbit Charge 2","hospital shift workers","PSQI","longitudinal study","heart rate monitoring","machine learning benchmark"],"falsifier":"A study that runs the same wrist-worn device against gold-standard clinical sleep recordings in both day-shift and night-shift hospital workers and checks whether the device's REM overestimation and sleep-stage misclassification differ statistically between the two shift groups; if they do, the paper's between-group comparisons and stage-label benchmarks would not hold.","tokens_in":14863,"feed_emoji":"😴","tokens_out":7918,"duration_ms":87851,"temperature":0.7,"pith_summary":"The paper's central claim is that the TILES-2018 Sleep Benchmark dataset—6,012 sleep recordings with continuous heart rate and device-labeled sleep stages from 139 hospital employees over ten weeks, plus demographics and PSQI self-reports—is a valid public resource for studying real-world sleep and benchmarking machine-learning models. The paper extends its earlier TILES-2018 dataset, and argues the combined cohort of 349 people and roughly 15,000 sleep recordings supports analyses that laboratory polysomnography datasets cannot: naturalistic, multi-week, shift-schedule comparisons. Its analyses find night-shift hospital workers report worse sleep quality, sleep fewer minutes, and exhibit more fragmented REM sleep, with REM-to-light transitions more likely than day-shift workers. It also reports benchmarks where deep time-series models outperform random forests and zero-shot LLMs at predicting self-reported PSQI scores, while random forests best predict demographics from sleep features. This matters because a public, longitudinal, wearable benchmark with both objective and subjective sleep measures is exactly what shift-work sleep research and model development need.","feed_headline":"Wearable sleep data from 139 hospital workers opens public benchmark","feed_subtitle":"Night shifts show shorter, more fragmented sleep; heart-rate models predict sleep quality and demographics.","key_machinery":"The central object is the dataset format itself: per-subject files with sleep metadata (start, end, total time, nap/main flag), minute-level heart-rate time series from the PPG sensor, and sleep-stage sequences with classic or stage labels, plus baseline PSQI scores and demographics. The argument runs on the pairing of continuous heart rate with device-labeled sleep stages across many nights, and on the use of the earlier TILES-2018 recordings as training data with this release as a fixed holdout evaluation set. For the behavioral analyses, the key mechanism is the per-participant sleep-stage transition probability graph, averaged within shift groups and compared with three-way ANOVA controlling for age and sex. For the benchmarks, the machinery is a set of models—a three-layer LSTM, a single-layer TimesNet block, a random forest on hand-crafted sleep features, and zero-shot large-language-model prompts—scored on the same held-out recordings.","core_discovery":"The paper's central discovery is that a ten-week, naturalistic wearable sleep collection from 139 hospital employees—6,012 sleep sessions with continuous heart rate and Fitbit-provided sleep-stage labels, matched with demographics and baseline PSQI scores—can serve as a public benchmark for real-world sleep research and machine learning. When combined with the earlier TILES-2018 dataset, the 349-participant cohort shows statistically significant differences between day- and night-shift workers: night-shift workers report worse PSQI scores, sleep fewer total minutes, have more variable sleep-onset and wake times, and are more likely to transition out of REM into light sleep, indicating more fragmented restorative sleep. Machine learning benchmarks trained on the earlier dataset and evaluated on this holdout show that deep time-series models reach macro-F1 around 0.57 for REM classification and beat both random forests and zero-shot LLMs on PSQI prediction, while a simple random forest best predicts demographics. These results are presented as evidence that the dataset supports both descriptive sleep-behavior science and reproducible model evaluation.","pith_inferences":["Editorial inference: the paper's assumption that Fitbit's systematic sleep-stage errors are shift-invariant could be tested directly by collecting a PSG-validated subsample; if REM bias differs between day and night shifts, the between-group comparisons would need correction.","Editorial inference: because nightly self-reported sleep quality was not collected, the PSQI benchmark likely predicts a stable trait-like sleep-quality score rather than night-by-night variation; adding daily EMA labels would create a stronger test of sleep-quality prediction.","Editorial inference: the strong shift-prediction performance of the random forest suggests sleep-schedule variability features could serve as a passive digital marker for shift-work disorder risk, though this would need clinical validation."],"forward_implications":["Night-shift hospital workers will show shorter total sleep, higher day-to-day sleep-schedule variability, and lower REM continuity than day-shift workers in this combined cohort.","Deep time-series models trained on one cohort of wearable sleep data can transfer to a held-out cohort with macro-F1 around 0.57 for three-class REM classification, making the dataset a usable testbed for sleep-stage modeling.","Self-reported PSQI scores can be predicted from sleep features better by deep time-series models than by zero-shot LLMs, suggesting physiological sleep data carries more signal than language-model reasoning about sleep.","Demographic attributes, especially age and shift type, are predictable from simple sleep features, meaning sleep physiology is structured enough to carry demographic information."],"supporting_citations":[{"why":"Supplies the earlier cohort used as training and validation data, and the basis for the combined 349-participant analyses.","marker":"[10]"},{"why":"Provides the Pittsburgh Sleep Quality Index scoring used as the self-reported sleep-quality labels.","marker":"[19]"},{"why":"Validation of the Fitbit Charge 2 against polysomnography that the paper cites to acknowledge device REM overestimation and to argue errors are systematic across groups.","marker":"[30]"},{"why":"LSTM architecture that underlies the SleepNet time-series sleep-stage and PSQI benchmarks.","marker":"[20]"},{"why":"TimesNet architecture used as the second time-series benchmark model.","marker":"[21]"},{"why":"Large-language-model family used for zero-shot self-reported sleep-quality prediction.","marker":"[22]"},{"why":"Prompt-construction approach for using LLMs to predict health variables from wearable data.","marker":"[23]"},{"why":"Tree explainer used to identify the top sleep features in demographic prediction.","marker":"[24]"},{"why":"Prior finding that deep sleep declines with age, used to interpret the age-prediction feature importance.","marker":"[25]"}],"fun_headline_variants":["Night shift workers show worse sleep in new public dataset","Public wearable sleep data from hospital workers for ML benchmarks","Sleep benchmark: 6,000+ sessions reveal shift work effects","New sleep dataset from 139 hospital employees aids ML research","Day vs night shift: fragmented sleep uncovered in public dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the wrist-worn device's automatic sleep-stage labels are accurate enough for these analyses, and specifically that its known tendency to overestimate REM sleep is the same for day-shift and night-shift workers; if the device's errors differ between shifts, the comparisons and model labels are not trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Night shift workers show worse sleep in new public dataset","Public wearable sleep data from hospital workers for ML benchmarks","Sleep benchmark: 6,000+ sessions reveal shift work effects","New sleep dataset from 139 hospital employees aids ML research","Day vs night shift: fragmented sleep uncovered in public dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1322,"prompt_tokens":986,"completion_tokens":336,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":602,"tokens_out":336,"duration_ms":4654,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:08:11.166825+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A study that runs the same wrist-worn device against gold-standard clinical sleep recordings in both day-shift and night-shift hospital workers and checks whether the device's REM overestimation and sleep-stage misclassification differ statistically between the two shift groups; if they do, the paper's between-group comparisons and stage-label benchmarks would not hold.","supporting_citations":[{"cited_title":"Tiles-2018, a longitudinal physiologic and behavioral data set of hospital workers,","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier cohort used as training and validation data, and the basis for the combined 349-participant analyses."},{"cited_title":"The pittsburgh sleep quality index: a new instrument for psychiatric practice and research,","cited_arxiv_id":null,"evidence_quote":"Provides the Pittsburgh Sleep Quality Index scoring used as the self-reported sleep-quality labels."},{"cited_title":"Validation of fitbit charge 2 sleep and heart rate estimates against polysomnographic measures in shift workers: Naturalistic study,","cited_arxiv_id":null,"evidence_quote":"Validation of the Fitbit Charge 2 against polysomnography that the paper cites to acknowledge device REM overestimation and to argue errors are systematic across groups."},{"cited_title":"From local explanations to global understanding with explainable ai for trees,","cited_arxiv_id":null,"evidence_quote":"Tree explainer used to identify the top sleep features in demographic prediction."},{"cited_title":"Sleep in normal aging,","cited_arxiv_id":null,"evidence_quote":"Prior finding that deep sleep declines with age, used to interpret the age-prediction feature importance."}],"review_version":1}