{"id":"20a95331-bb90-4917-9265-1cd0324fe138","arxiv_id":"2502.07526","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"CodePhys casts remote heart-rate measurement as a code query task: a video encoder produces features matched to a learned codebook of clean PPG waveforms, and a pre-trained decoder reconstructs the pulse.","lead":"CodePhys measures heart rate from video by replacing noisy video-derived features with the closest clean PPG waveforms stored in a codebook learned from ground-truth pulse signals. On four benchmark datasets it reports lower heart-rate errors than prior methods, especially under blur, noise, low resolution, occlusion, and brightness changes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The code-query correction in Eq. (2)/(11) lacks a demonstrated correctness guarantee: a degraded video query can land in the wrong Voronoi cell, and the paper reports no code-selection accuracy on held-out or degraded data.","rationale":"The reader's weakest assumption identifies the same mechanism: the Stage I codebook may fail to cover unseen PPG variability and the hard nearest-neighbor selection may discard or misroute heart-rate-relevant detail. My stress-test refines this into a precise, testable condition: on a degraded query, Eq. (2) must select the same codebook item that a clean query would select, and the paper supplies no direct evidence of this. I am not objecting to the method's empirical results: the ablations in Table V and Table VI, the cross-dataset results in Tables III-IV, and the robustness evaluation in Section IV-H are internally consistent and provide real support for the overall pipeline. The Stage I reconstruction results in Table VII also show that the codebook can represent PPG signals well on the data used. The concern is specifically about the central mechanistic claim that codebook querying corrects degraded features by matching them to the most similar noise-free items. That claim requires either a coverage argument or a direct measurement of code-selection behavior; neither is present. The paper does not contain an internal contradiction, and the available evidence is sufficient for a conditional acceptance rather than rejection. The lack of released code, single-fold robustness evaluation, and absence of significance testing are secondary concerns that further justify keeping the reader's CONDITIONAL verdict unchanged. Therefore, I recommend no change to the reader's verdict: UNCHANGED.","tokens_in":22891,"tokens_out":5577,"duration_ms":58740,"concrete_test":"On held-out test clips in the setting of Table X, compute Qgt = QUERY(Es(sgt), C) from the clean GT-PPG and Qrppg = QUERY(Zrppg, C) from the video branch, then report the per-token argmin agreement between Qgt and Qrppg, together with the mean distance from Zrppg to its selected code, separately for clean and degraded inputs. If agreement is high on clean clips but degrades sharply under blur, noise, occlusion, or brightness changes, the claim that the code query 'corrects' degraded features rather than merely regularizing them is unsupported; a companion check is to replace the hard argmin in Eq. (2) with a soft top-k combination and verify whether HR metrics are preserved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Eq. (11) removes visual interference by replacing degraded rPPG features with the nearest noise-free codebook item. For this to hold on unseen subjects, two conditions must be met: (i) the 64-item codebook C learned in Stage I covers the latent PPG manifold of test subjects, and (ii) for any video-derived query Zrppg, the argmin in Eq. (2) selects the same code item that Es would assign to the clean GT-PPG. Condition (i) is only indirectly supported: Table VII shows near-perfect Stage I reconstruction, but no per-subject, per-clip, or held-out breakdown is given, and C is trained using only training-set GT-PPG signals. Condition (ii) is the load-bearing step. Equations (9)-(10) combine AdaIN-aligned video features with APB features, and Eq. (14) pulls Zrppg toward the selected code via a stop-gradient L2 loss, but nothing prevents a heavily degraded query from being closest to an incorrect codebook item. The robustness study in Section IV-H applies synthetic degradations to the same VIPL-HR Fold-1 used for training and reports only final HR metrics, not whether the selected codes match the codes of the corresponding clean GT-PPG. The Table V ablation shows the overall query process helps, but it does not separate 'correct code selection' from 'a learned bottleneck that happens to discard noise.' Without code-selection accuracy or distance-to-selected-code diagnostics, the paper does not establish that the mechanism works by choosing the right noise-free code rather than by acting as a generic learned regularizer.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"CodePhys is a two-stage framework for camera-based remote photoplethysmography (rPPG). In Stage I, a 1D signal autoencoder is trained to reconstruct GT-PPG waveforms, and a 64-item codebook of latent features is obtained via vector-quantization-style losses. In Stage II, a video feature extractor with spatial attention and a spatio-temporal encoder with an auxiliary prior branch produce query features; a soft feature distillation loss aligns the video features with GT-PPG features, and the query features are replaced by the nearest codebook item before the frozen decoder produces the predicted rPPG signal. The paper reports intra-dataset and cross-dataset heart-rate results on VIPL-HR, UBFC-rPPG, PURE, and COHFACE, component ablations, an efficiency comparison, a plug-and-play experiment with existing backbones, and a robustness study with five synthetic video degradations.","tokens_in":23242,"tokens_out":5006,"duration_ms":46237,"significance":"Conditional on the empirical claims, CodePhys is a novel application of discrete codebook priors to rPPG, and the plug-and-play integration of existing backbones in Table IX is practically useful. The paper is unusually broad in comparisons and includes an ablation for each component (CQP, APB, SFD, SAM, Stage I). However, the central mechanism, that code querying removes visual interference by selecting the correct noise-free code, is not directly evidenced, and the performance claims rest on single-seed point estimates and a one-fold robustness experiment.","major_comments":[{"comment":"The load-bearing claim that Eq. (11) corrects degraded rPPG features by retrieving the correct noise-free codebook item is not directly verified. The CQP ablation in Table V only toggles the query operation as a whole; it cannot distinguish 'correct code selection' from a learned projection that happens to discard noise. Furthermore, Table VII reports near-perfect Stage I reconstruction but does not state whether those reconstructions are obtained on held-out subjects/clips; if they are training-set reconstructions, they do not establish that the 64-item codebook covers test-subject PPG variability. Please report code-selection accuracy (agreement between the argmin in Eq. (2) for video queries and Qgt computed from GT-PPG), per-subject code-usage statistics, and distance-to-selected-code diagnostics on held-out and degraded test clips.","section":"Section III-B3, Eq. (14), Tables V-VII"},{"comment":"The robustness study applies synthetic degradations only to VIPL-HR Fold-1 and reports only final MAE/RMSE values. It is not specified whether 'trained on the Fold-1' means the test fold is held out, and if the same fold is used for hyperparameter selection the comparison is weakened. At minimum, report robustness across all five VIPL-HR folds or an equivalent held-out protocol, include code-selection accuracy under each degradation type, and provide confidence intervals across degradation realizations so the reader can see that the robustness claim is not specific to one fold.","section":"Section IV-H, Table X"},{"comment":"All results are point estimates from a single training run. Given the small absolute differences on some datasets (e.g., Table II, UBFC-rPPG MAE 0.21 bpm vs. 0.40 bpm for PhysFormer), the state-of-the-art claim is not established without multiple seeds or paired significance tests. Please report mean and standard deviation over at least three seeds for the main intra-dataset and cross-dataset tables, or provide statistical significance tests against the strongest baselines.","section":"Tables I-IV"}],"minor_comments":[{"comment":"The codebook is initialized randomly and optimized through the 'GLO strategy [50]', but neither the GLO update rule nor its role relative to standard VQ training is described; please add a sentence explaining this choice.","section":"Section III-A2"},{"comment":"The distance function D(·,·) uses the same symbol as the decoder D_s, which is confusing; please rename one of them.","section":"Eq. (12)"},{"comment":"The table footnotes define △, ‡, and ⋆, but the meaning of '--' entries (methods not evaluated on a dataset) is not stated; please clarify.","section":"Tables I and II"},{"comment":"The degradation ranges are given, but the number of random realizations per degradation type and whether the same degraded videos are used for both CodePhys and PhysFormer are not stated; please specify this protocol.","section":"Section IV-H"},{"comment":"The hyperparameter ablation figure lacks axis labels; please add them so the sensitivity results are interpretable.","section":"Figure 9"}],"recommendation":"major_revision","confidential_remarks":"I see no scope or ethical concerns; the paper fits JBHI. The main revision request is to substantiate the code-selection mechanism and to add statistical rigor to the empirical claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: CodePhys is the first rPPG method I know that builds a VQ codebook from ground-truth PPG signals and treats estimation as a code-query problem. On the evidence in the paper it works: state-of-the-art HR errors on four benchmarks, both intra- and cross-dataset, plus a plug-and-play demonstration where DeepPhys, EfficientPhys, and PhysFormer all improve when their extractors are swapped into the framework. The ablations support each component, and the Stage I reconstruction table shows the autoencoder/codebook is doing something sensible. That is real, reproducible-in-principle work, though no code or data artifacts are released.\n\nThe soft spots are mostly about verification. All metrics are point estimates; no significance tests, no multiple-seed variance. The robustness study in Section IV-H is on VIPL-HR Fold-1 only, with synthetic degradations applied at test time, and reports only final HR errors. The stress-test concern lands, but only partially: the paper's own description says the code query corrects degraded features by replacing them with the most similar noise-free code item, yet nothing in the paper measures whether the selected code matches the code a clean GT-PPG would have produced. Table VII shows good reconstruction, but that is an autoencoder sanity check, not a diagnosis of code-selection accuracy on video queries. So the mechanism is underdetermined: the gains could come from the learned bottleneck acting as a generic regularizer rather than from 'right code' selection. That does not kill the paper, but it means the explanatory story is stronger than the evidence.\n\nI also want to note the authors are honest about scope: they state in the conclusion that lab datasets may not capture real-world variability and that efficiency needs work. The ablation table V is ambiguously formatted (the first four rows are missing component indicators); that should be cleaned up.\n\nVerdict: the empirical claims are credible and the idea is worth building on, but independent verification requires release of code or at least per-seed variance and code-selection diagnostics. This deserves serious peer review.","headline":"Solid, well-executed rPPG paper with a genuinely new codebook-query idea; the central mechanism is under-verified, but the empirical case is strong enough to warrant peer review.","tokens_in":23783,"tokens_out":1875,"would_cite":true,"duration_ms":16562,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CodePhys recovers heart rate from video by looking up clean PPG codes instead of denoising the signal directly.","keywords":["remote photoplethysmography","rPPG","codebook querying","discrete representation learning","vector quantization","heart rate estimation","visual interference robustness","knowledge distillation"],"falsifier":"Train CodePhys on subjects whose heart rates never exceed 80 bpm and test on subjects whose heart rates exceed 120 bpm, then compare MAE against a version with the codebook removed; if the error gap reverses, the codebook coverage assumption fails. A finer check is to perturb query features adversarially and see whether two physiologically distinct waveforms collapse to the same codebook entry, which would show loss of detail.","tokens_in":22655,"feed_emoji":"🫀","tokens_out":5236,"duration_ms":46867,"temperature":0.7,"pith_summary":"This paper argues that accurate remote heart-rate measurement (rPPG) from facial video is best framed not as signal extraction but as retrieval: the system learns a codebook of clean, ground-truth photoplethysmography (PPG) fragments, and any degraded feature pulled from a video is replaced by its nearest clean entry before decoding into a pulse waveform. The claim is that visual interference such as motion blur, camera noise, changing resolution, occlusion, and brightness shifts gets corrected in one uniform mechanism rather than by hand-crafted modules tuned to each nuisance. If this is right, video-based vitals monitoring becomes substantially more resilient: CodePhys reports the lowest heart-rate errors on VIPL-HR, UBFC-rPPG, PURE, and COHFACE in both intra-dataset and cross-dataset settings, and it can be grafted onto existing rPPG backbones with measurable gains. The paper's bets are that a 64-item codebook built from ground-truth PPG signals covers enough pulse variability, and that hard nearest-neighbour lookup in the learned latent space preserves the heart-rate-relevant details.","feed_headline":"Video heart-rate measurement becomes a clean-code search","feed_subtitle":"Degraded rPPG features are replaced by their nearest clean PPG entries, cutting heart-rate error across four benchmarks.","key_machinery":"The central object is the noise-free PPG codebook $\\mathcal{C} = (\\mathbf{c}_1,\\ldots,\\mathbf{c}_N) \\in \\mathbb{R}^{N\\times D}$ with $N=D=64$, learned by reconstructing GT-PPG signals through a signal encoder $\\mathbf{E}_s$ and decoder $\\mathbf{D}_s$ using a quantization loss. The code query process is a hard nearest-neighbour lookup: for each query feature $\\mathbf{z}_i$, a one-hot coordinate $\\mathbf{q}_i$ marks the closest codebook item under the Euclidean norm, and the feature is replaced by $\\mathbf{q}_i\\cdot\\mathcal{C}$. The pre-trained decoder then turns the selected clean entries into a pulse signal, which is what converts denoising into a lookup against a noise-free prior.","core_discovery":"CodePhys treats rPPG measurement as a code query task in a noise-free proxy space. Stage I trains a signal autoencoder on ground-truth PPG signals with a vector-quantization objective, producing a codebook $\\mathcal{C}$ of 64 items in $\\mathbb{R}^{64}$ whose nearest-neighbour lookup, followed by decoding, reconstructs PPG nearly exactly. Stage II freezes $\\mathcal{C}$ and the decoder; a spatial-aware video encoder maps facial video into query features $\\mathbf{Z}_{\\text{rppg}}$, an auxiliary prior branch aligns their distribution, and the model outputs the decoded signal from the nearest clean codebook items. The central discovery claim is that this single replacement removes general visual interference, not just the specific kinds previous methods target, and yields state-of-the-art heart-rate accuracy: MAE of 4.27 bpm on VIPL-HR, 0.21 bpm on UBFC-rPPG, 0.39 bpm on PURE, and 1.19 bpm on COHFACE, with particular stability when test videos are synthetically degraded by blur, noise, resolution changes, occlusion, or brightness shifts.","pith_inferences":["A natural extension the authors leave implicit is building the Stage I codebook from contact PPG recorded across diverse populations, age groups, and sensor types, which would test whether the clean-code prior covers pulse morphologies outside the benchmark training sets.","Because retrieval is a hard nearest-neighbour lookup, a testable variant would replace it with soft or top-$k$ mixtures of code entries; this could preserve waveform morphology for very low or irregular heart rates where a single code may be too coarse.","The paper assumes each video query feature is a degraded version of a single clean PPG code, so a face video containing physiological signal from more than one source, such as two faces, might force the lookup to merge distinct pulses; enforcing a sparse spatial mixture of codes per region is a concrete way to address that."],"forward_implications":["Because the frozen codebook and decoder handle degradation jointly, visual-interference robustness in rPPG no longer needs a bespoke module for each nuisance (motion, blur, compression, occlusion), which is the paper's main bet.","Existing end-to-end rPPG networks can be upgraded by replacing only their video feature extractor with the query pipeline, and the paper reports that DeepPhys, EfficientPhys, and PhysFormer all improve in MAE this way.","Cross-dataset transfer improves: training on PURE plus COHFACE and testing on UBFC-rPPG reaches MAE 0.58 bpm, and training on VIPL-HR then testing on PURE reaches 4.03 bpm, evidence that the clean-code prior generalizes across domains.","The added modules are lightweight at inference, with 5.73 million parameters, 75.79 G MACs, and 0.12 ms per frame on the reported GPU, so the querying step need not block real-time use."],"supporting_citations":[{"why":"Supplies the vector-quantization autoencoder mechanism whose codebook and quantization loss Stage I adapts from image representation to PPG signals.","marker":"[21]"},{"why":"Provides the straight-through gradient estimator and codebook training recipe that make the non-differentiable argmin query trainable.","marker":"[22]"},{"why":"Demonstrates codebook-based restoration of degraded face images by replacing low-quality features with clean codebook entries, the template CodePhys transfers to rPPG signals.","marker":"[23]"},{"why":"Motivates using a discrete learned codebook as a prior for generating a target modality, here speech-driven facial motion, which the paper extends to pulse generation.","marker":"[24]"},{"why":"Supplies the GLO strategy used to initialize and optimize the codebook items in the latent embedding space.","marker":"[50]"},{"why":"Serves as a recent Transformer-based end-to-end rPPG baseline that CodePhys compares against and integrates with to show plug-and-play gains.","marker":"[14]"},{"why":"Provides a leading traditional rPPG method that is used as a baseline in both intra-dataset and cross-dataset heart-rate comparisons.","marker":"[5]"}],"fun_headline_variants":["Heart rate from video: query a clean codebook","Clean-code lookup makes video heart rate robust","CodePhys: replace noisy rPPG with clean code queries","Video pulse from nearest clean PPG code"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 64-entry codebook learned from training-set PPG signals covers the pulse variability of unseen subjects well enough that hard nearest-neighbour lookup never throws away heart-rate-relevant detail.","fun_headline_variants_meta":{"raw":{"variants":["Heart rate from video: query a clean codebook","Clean-code lookup makes video heart rate robust","CodePhys: replace noisy rPPG with clean code queries","Video pulse from nearest clean PPG code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1682,"prompt_tokens":1033,"completion_tokens":649,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":589}},"tokens_in":649,"tokens_out":649,"duration_ms":6255,"temperature":1.0,"reasoning_tokens":589,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T12:26:26.398557+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train CodePhys on subjects whose heart rates never exceed 80 bpm and test on subjects whose heart rates exceed 120 bpm, then compare MAE against a version with the codebook removed; if the error gap reverses, the codebook coverage assumption fails. A finer check is to perturb query features adversarially and see whether two physiologically distinct waveforms collapse to the same codebook entry, which would show loss of detail.","supporting_citations":[{"cited_title":"Neural discrete representation learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the vector-quantization autoencoder mechanism whose codebook and quantization loss Stage I adapts from image representation to PPG signals."},{"cited_title":"Taming transformers for high- resolution image synthesis,","cited_arxiv_id":null,"evidence_quote":"Provides the straight-through gradient estimator and codebook training recipe that make the non-differentiable argmin query trainable."},{"cited_title":"Towards robust blind face restoration with codebook lookup transformer,","cited_arxiv_id":null,"evidence_quote":"Demonstrates codebook-based restoration of degraded face images by replacing low-quality features with clean codebook entries, the template CodePhys transfers to rPPG signals."},{"cited_title":"Codetalker: Speech-driven 3d facial animation with discrete motion prior,","cited_arxiv_id":null,"evidence_quote":"Motivates using a discrete learned codebook as a prior for generating a target modality, here speech-driven facial motion, which the paper extends to pulse generation."},{"cited_title":"Physformer: Facial video-based physiological measurement with temporal difference transformer,","cited_arxiv_id":null,"evidence_quote":"Serves as a recent Transformer-based end-to-end rPPG baseline that CodePhys compares against and integrates with to show plug-and-play gains."},{"cited_title":"Algorithmic principles of remote PPG,","cited_arxiv_id":null,"evidence_quote":"Provides a leading traditional rPPG method that is used as a baseline in both intra-dataset and cross-dataset heart-rate comparisons."}],"review_version":1}