REVIEW 2 major objections 1 minor 22 references
Pixel Watch: Robust Heart Rate Sensing from Multipath PPG and On-Device Deep Learning Trained on 10,000 hours of Free-Living and Fitness Data
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Pixel Watch 2 pairs multipath PPG with a 300K-parameter neural net trained on 10,000 hours to reach heart rate limits of agreement under 11 BPM during exercise.
desk verdict The paper delivers concrete on-device HR sensing gains on Pixel Watch 2 from multipath PPG plus a 300k-param CNN trained at 10k-hour scale, with tighter LoA than prior Google devices, but the independence of the two validation sets from the training corpus is not fully shown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A 15-layer temporally dilated convolutional neural network with approximately 300,000 parameters that ingests 10 optical PPG channels to produce 1 Hz heart rate estimates.
What would settle it
A new study that records simultaneous ECG reference measurements on an independent cohort wearing the Pixel Watch 2 during comparable exercise and free-living tasks would test whether the reported limits of agreement hold.
Extended reading notes
Core claim
The Pixel Watch 2 is the first Google smartwatch to combine multipath photoplethysmography with deep learning-based heart rate inference. It processes 10 optical channels using an on-device 15-layer temporally dilated convolutional neural network of approximately 300,000 parameters to produce a 1 Hz heart rate output. Training on 10,000 hours of data from 962 participants enables 95 percent limits of agreement from -10.34 to 8.66 BPM during exercise and -6.57 to 7.48 BPM during free-living activities on two independent validation sets, outperforming previous devices and showing that deep learning fully exploits multipath PPG hardware.
Load-bearing premise
The two validation datasets are fully independent of the training corpus and representative of target users without curation-induced selection bias.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript describes the Pixel Watch 2 heart-rate system, which fuses multipath PPG (10 optical channels) with an on-device 15-layer temporally dilated CNN (~300K parameters) trained on 10,000 hours of curated data from 962 participants. On two held-out validation sets—an in-house fitness set (229 participants, 250 h) and an external free-living set (27 participants, >1000 h)—the system reports 95% limits of agreement of −10.34 to 8.66 BPM during exercise and −6.57 to 7.48 BPM during free-living activities, stated to be substantially tighter than prior Google devices.
Significance. If the reported limits of agreement are obtained on truly disjoint and representative validation cohorts, the result supplies concrete evidence that large-scale deep learning can extract substantially more signal from multipath PPG hardware than conventional processing pipelines, particularly under motion. The scale of the training corpus and the on-device deployment constitute clear engineering strengths.
major comments (2)
- [Abstract] Abstract: The central performance claim rests on the two validation sets being fully independent of the 962-participant training corpus and free of curation-induced selection bias. The abstract states only that the sets are 'independent' and that training data were 'curated from a broader corpus,' without participant-level overlap checks, explicit curation criteria, or confirmation that identical inclusion/exclusion rules were applied uniformly. These details are required to interpret the reported LoA values as evidence of generalization.
- [Abstract] Abstract / Methods: No description is provided of the reference device used to generate ground-truth heart-rate labels, the precise data-exclusion rules applied to the 10k-hour corpus, or the statistical procedure used to compute the 95% limits of agreement. These omissions directly affect the verifiability of the quantitative margins that constitute the paper's primary result.
minor comments (1)
- [Abstract] The abstract mentions 'statistical tests' implicitly through the LoA figures but does not report them; adding the exact test names and p-values (or confidence intervals on the LoA bounds) would improve clarity without altering the central claim.
Simulated Author's Rebuttal
We thank the referee for their thorough review and valuable feedback on our manuscript. We address each of the major comments below and will make revisions to enhance the transparency of our methods and validation procedures.
read point-by-point responses
-
Referee: [Abstract] Abstract: The central performance claim rests on the two validation sets being fully independent of the 962-participant training corpus and free of curation-induced selection bias. The abstract states only that the sets are 'independent' and that training data were 'curated from a broader corpus,' without participant-level overlap checks, explicit curation criteria, or confirmation that identical inclusion/exclusion rules were applied uniformly. These details are required to interpret the reported LoA values as evidence of generalization.
Authors: We agree that more explicit details on the independence of the validation sets would aid interpretation. The validation sets were collected from distinct participant cohorts with no overlap in participants from the training corpus. We will revise the abstract to state that the validation sets are participant-disjoint from the training data and that consistent inclusion/exclusion rules were applied. Further details on curation criteria will be elaborated in the Methods section of the revised manuscript. revision: yes
-
Referee: [Abstract] Abstract / Methods: No description is provided of the reference device used to generate ground-truth heart-rate labels, the precise data-exclusion rules applied to the 10k-hour corpus, or the statistical procedure used to compute the 95% limits of agreement. These omissions directly affect the verifiability of the quantitative margins that constitute the paper's primary result.
Authors: We acknowledge that these methodological details are not sufficiently described in the current version. In the revised manuscript, we will add concise descriptions of the reference device, data-exclusion rules, and the statistical method for computing the 95% limits of agreement to the abstract and ensure comprehensive coverage in the Methods section. revision: yes
Circularity Check
No circularity: empirical validation metrics on independent sets
full rationale
The paper reports measured 95% limits of agreement from direct evaluation of a trained 15-layer CNN on two explicitly described independent validation datasets (in-house fitness with 229 participants and external free-living with 27 participants) after training on a separate 10,000-hour corpus from 962 participants. No equations, derivations, or first-principles results are presented that reduce to fitted parameters or self-defined quantities by construction. No self-citations are invoked to justify uniqueness theorems or load-bearing premises, and no ansatz or renaming of known results occurs. The central claims are statistical performance numbers obtained from held-out evaluation, which by definition cannot be circular within the paper's own derivation chain.
Assumptions & free parameters
assumptions (1)
- domain assumption The validation sets are independent of the training data and representative of real-world use without selection bias from curation.
Cite this review
Pith. "Pith review of Pixel Watch: Robust Heart Rate Sensing from Multipath PPG and On-Device Deep Learning Trained on 10,000 hours of Free-Living and Fitness Data." pith.science (2026). https://pith.science/paper/2UFDLLV6
@misc{pith2026260621436,
author = {Pith},
title = {Pith review of: Pixel Watch: Robust Heart Rate Sensing from Multipath PPG and On-Device Deep Learning Trained on 10,000 hours of Free-Living and Fitness Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/2UFDLLV6}},
note = {Machine review of arXiv:2606.21436}
}
read the original abstract
The Pixel Watch 2 (PW2) is the first Google smartwatch to combine multipath photoplethysmography (PPG) with deep learning-based heart rate inference, designed to significantly improve sensing accuracy during motion-heavy activities. The device processes 10 optical channels using an on-device, 15-layer temporally dilated convolutional neural network (~300K parameters) to yield a 1 Hz heart rate output. Crucial to this model's performance was its training on a massive dataset comprising 10,000 hours of data from 962 participants, curated from a broader corpus of controlled and free-living activities. We evaluated the PW2's sensing performance across two independent validation sets: an in-house fitness dataset (229 participants, 250 hours) and an external free-living dataset (27 participants, 1000+ hours). The system achieved 95% Limits of Agreement of -10.34 to 8.66 BPM during exercise and -6.57 to 7.48 BPM during free-living activities, demonstrating substantially tighter error margins than previous Google devices. Finally, we discuss key design lessons, emphasizing that large-scale deep learning was instrumental in fully leveraging multipath PPG hardware over traditional signal processing approaches.
Figures
Reference graph
Works this paper leans on
-
[1]
Large scale population assessment of physical activity using wrist worn accelerometers: The UK Biobank study,
A. Dohertyet al., “Large scale population assessment of physical activity using wrist worn accelerometers: The UK Biobank study,”PLoS One, 2017
2017
-
[2]
The “All of Us
The All of Us Research Program Investigators, “The “All of Us” research program,” N Engl J Med, no. 7, pp. 668–676, 2019
2019
-
[3]
Fitbit Charge HR wireless heart rate monitor: Validation study conducted under free-living conditions,
A. W. Gornyet al., “Fitbit Charge HR wireless heart rate monitor: Validation study conducted under free-living conditions,”JMIR Mhealth Uhealth., 2017
2017
-
[4]
Guidelines for wrist-worn consumer wearable assessment of heart rate in biobehavioral research,
B. Nelsonet al., “Guidelines for wrist-worn consumer wearable assessment of heart rate in biobehavioral research,”NPJ Digit Med., vol. 3, no. 90, 2020
2020
-
[5]
Comprehensive comparison of Apple Watch and Fitbit monitors in a free-living setting,
Y . Baiet al., “Comprehensive comparison of Apple Watch and Fitbit monitors in a free-living setting,”PLoS One, 2021
2021
-
[6]
Wrist-worn devices for the measurement of heart rate and energy expenditure: A validation study for the Apple Watch 6, Polar Vantage V and Fitbit Sense,
G. Hajj-Boutroset al., “Wrist-worn devices for the measurement of heart rate and energy expenditure: A validation study for the Apple Watch 6, Polar Vantage V and Fitbit Sense,”Eur J Sport Sci., 2023
2023
-
[7]
Measurement of heart rate using the Polar OH1 and Fitbit Charge 3 wearable devices in healthy adults during light, moderate, vigorous, and sprint-based exercise: Validation study,
D. J. Muggeridgeet al., “Measurement of heart rate using the Polar OH1 and Fitbit Charge 3 wearable devices in healthy adults during light, moderate, vigorous, and sprint-based exercise: Validation study,”JMIR Mhealth Uhealth, 2021
2021
-
[8]
Assessment of Samsung Galaxy Watch4 PPG-based heart rate during light-to-vigorous physical activities,
C. S. Limaet al., “Assessment of Samsung Galaxy Watch4 PPG-based heart rate during light-to-vigorous physical activities,”IEEE Sensors Letters, 2024
2024
Show all 22 references
-
[9]
Preliminary assessment of the Samsung Galaxy Watch 5 accuracy for the monitoring of heart rate and heart rate variability parameters,
G. Rhoet al., “Preliminary assessment of the Samsung Galaxy Watch 5 accuracy for the monitoring of heart rate and heart rate variability parameters,” inProc. MEDICON’23 and CMBEBIH’23. Springer, 2023, pp. 22—-30
2023
-
[10]
Heart rate measurement accuracy of Fitbit Charge 4 and Samsung Galaxy Watch Active2: Device evaluation study,
M. Nissenet al., “Heart rate measurement accuracy of Fitbit Charge 4 and Samsung Galaxy Watch Active2: Device evaluation study,”JMIR F orm Res., 2022
2022
-
[11]
Validity of the wrist-worn Polar Vantage V2 to measure heart rate and heart rate variability at rest,
O.-P. Nuuttila, E. Korhonen, J. Laukkanen, and H. Kyröläinen, “Validity of the wrist-worn Polar Vantage V2 to measure heart rate and heart rate variability at rest,”Sensors, 2022
2022
-
[12]
Commercial smart watches and heart rate monitors: A concurrent validity analysis,
S. Montalvoet al., “Commercial smart watches and heart rate monitors: A concurrent validity analysis,”Journal of Strength and Conditioning Research, 2023
2023
-
[13]
Criterion validity and accuracy of a heart rate monitor,
V . O. Damascenoet al., “Criterion validity and accuracy of a heart rate monitor,” Human Movement, 2022
2022
-
[14]
Wrist-worn wearables for monitoring heart rate and energy expenditure while sitting or performing light-to-vigorous physical activity: Validation study,
P. Dükinget al., “Wrist-worn wearables for monitoring heart rate and energy expenditure while sitting or performing light-to-vigorous physical activity: Validation study,”JMIR Mhealth Uhealth., 2020
2020
-
[15]
Deep PPG: Large-scale heart rate estimation with convolutional neural networks,
A. Reiss, I. Indlekofer, P. Schmidt, and K. Van Laerhoven, “Deep PPG: Large-scale heart rate estimation with convolutional neural networks,”Sensors, 2019
2019
-
[16]
Deep learning fused wearable pressure and PPG data for accurate heart rate monitoring,
P. Mehrgardtet al., “Deep learning fused wearable pressure and PPG data for accurate heart rate monitoring,”IEEE Sensors Journal, 2021
2021
-
[17]
DeepHeart: A deep learning approach for accurate heartrate estimation from PPG signals,
X. Changet al., “DeepHeart: A deep learning approach for accurate heartrate estimation from PPG signals,”ACM Trans. Sen. Netw., 2021
2021
-
[18]
A review of deep learning methods for photoplethysmography data,
G. Nie, J. Zhu, G. Tang, D. Zhang, S. Geng, Q. Zhao, and S. Hong, “A review of deep learning methods for photoplethysmography data,”arXiv:2401.12783, 2024
2024 arXiv
-
[19]
A review of wearable multi-wavelength photoplethysmography,
D. Rayet al., “A review of wearable multi-wavelength photoplethysmography,” IEEE Reviews in Biomedical Engineering, 2023
2023
-
[20]
How to wear Google Pixel Watch,
Google, “How to wear Google Pixel Watch,” Apr 2026. [Online]. Available: https://support.google.com/googlepixelwatch/answer/12724980
2026
-
[21]
Reliability and validity of the combined heart rate and movement sensor Actiheart,
S. Brageet al., “Reliability and validity of the combined heart rate and movement sensor Actiheart,”Eur J Clin Nutr, 2005
2005
-
[22]
Gaussian process robust regression for noisy heart rate data,
O. Stegleet al., “Gaussian process robust regression for noisy heart rate data,” IEEE Trans Biomed Eng., 2008
2008
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.