REVIEW 3 major objections 7 minor 11 references
A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps
T0 review · 3 major / 7 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read Predicted cortical drive from a top brain-encoding model does not track which moments of YouTube videos people re-watch, up to a bound of r≈0.14.
desk verdict A carefully bounded null: TRIBE predicted cortical drive does not track YouTube most-replayed beyond position and low-level baselines, with real artifact controls and released code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Position-controlled partial correlation of TRIBE global field power (root-mean-square over cortical vertices) against YouTube most-replayed heatmaps, with matched/mismatched leave-one-video-out probes and an equivalence/Bayes bound on the null.
What would settle it
A subcortex-inclusive encoder that includes nucleus accumbens, or a cleaner creator-side audience-retention target on a larger balanced sample, yielding a position-controlled partial correlation clearly above the loudness baseline and outside the current r≈0.14 equivalence bound.
Extended reading notes
Core claim
On this target and these readouts, a predicted-fMRI drive signal from TRIBE carries approximately no content-specific re-watch signal, up to a bound of r≈0.14. The pooled position-controlled partial correlation is +0.058 (95% CI [−0.04, 0.15]), indistinguishable from zero and from low-level baselines; the null holds for network GFP, signed value ROIs, permutation tests, and predicted ISC, while an apparent supervised r=0.47 is a shared temporal-shape artifact.
Load-bearing premise
That most-replayed heatmaps on already-popular clips, over a 60-second window and compared mainly via cortical-surface summaries, are a fair enough behavioral target to detect neuroforecasting transfer if it existed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript tests whether predicted cortical responses from TRIBE (the 2025 Algonauts winner) forecast moment-level YouTube re-watch behavior. TRIBE is run unmodified on 48 videos; its predicted surface response is reduced to a per-TR global field power (GFP) engagement curve and compared to each video’s most-replayed heatmap. The primary metric is a position-controlled partial correlation (quadratic detrend). The pooled partial r is +0.058 (95% CI [−0.04, 0.15]; t(47)=1.21, p=0.23), not above loudness/motion baselines, and near zero in raw form outside a music-video onset artifact. The null is stable across network and signed value/salience ROI readouts, a circular-shift permutation, and a supervised leave-one-video-out cortical probe that only appears successful (r=0.47) under a coarse detrend and collapses under spline + matched/mismatched controls. A borderline content-specific signal is reported only in the visual input stream; audio, text, and predicted cortex show none. Because the released TRIBE checkpoint is subject-averaged, the authors fit their own per-subject encoders and find that predicted ISC also fails to track re-watch (r=−0.04). They bound rather than merely fail to reject the null (BF01=3.2; effects ≳0.14 excluded; target split-half reliability ≈0.82). Code, video-ID manifest, and a SABR-resilient acquisition method are released.
Significance. If the bounded null holds, the paper is a useful negative result for anyone tempted to treat modern multimodal brain encoders (or their foundation features) as off-the-shelf engagement predictors. Its main contributions are (i) a carefully scoped claim with Bayes factor, equivalence-style bound, and reliability ceiling rather than a bare non-significant p-value; (ii) a concrete methodological caution that most-replayed labels plus a coarse detrend can manufacture large cross-validated correlations that are pure shared temporal shape; and (iii) a mechanistic sketch that group-trained encoding collapses the weak video-specific structure present in visual features. Full release of code, video IDs, and an acquisition pipeline that works under SABR is a genuine reproducibility strength and should be credited. The result is narrow by design (one model, cortical surface, biased target, 60 s window), but within that scope it is informative.
major comments (3)
- [§6.4, Abstract, Conclusion] §6.4 vs abstract/conclusion: the pre-specified TOST against δ=0.10 does not establish equivalence (p=0.20). The claim that effects larger than r≈0.14 are excluded comes from the descriptive 90% CI, not from a successful TOST at a pre-registered SEI. Please state explicitly what was pre-specified versus post-hoc, and align the abstract wording (“an equivalence test excludes effects above r≈0.14”) with §6.4 so readers do not read a failed TOST as a passed equivalence test. BF01=3.2 and the reliability ceiling already support a bounded null; the equivalence language just needs to match the procedure.
- [§3.2, §5, §6.8, §8] §3.2 and §5 fix the analysis to the first 60 s (≈60 TRs) of every clip. Most-replayed markers span the full video, and intros/onsets are exactly the position artifact the paper works hard to remove. §6.8 shows that marker density and duration do not correlate with the effect, which is helpful, but does not test whether the null is specific to the opening minute. A sensitivity check—e.g., a mid-clip 60 s window, or full-length analysis where TRIBE segment onsets allow—would make the “moment-level” claim less dependent on a single, intro-heavy window. If that is infeasible for encoding cost, state the limitation more prominently in §8 and qualify the claim accordingly.
- [§6.7] §6.7 is presented as a direct test of the closest prior positive result (ISC). The custom per-subject ridge encoders are validated at only r≈0.15 in-domain and r≈0.10 cross-domain. That is within the published range for feature-to-cortex models, but it leaves open how large a true ISC–rewatch effect could have been and still been missed. Please add a brief power or attenuation note (e.g., expected attenuation of a behavioral correlation given encoder r≈0.10–0.15) so the ISC null is interpretable as evidence against transfer rather than only as a weak encoder. The paper already notes that the released TRIBE checkpoint cannot supply per-subject predictions; that disclosure is good and should stay.
minor comments (7)
- [Table 1] Table 1 leaves raw r blank for the loudness and motion baselines; either fill those cells or note in the caption that only partial r is reported for baselines by design.
- [Figure 1c] Figure 1c category labels include “?” for n=4; replace with the “misc” label used in §5 for consistency.
- [Figure 2a, §6.4] Figure 2a labels the equivalence region ±0.139 while the text uses r≈0.14; pick one rounding and use it throughout abstract, §6.4, and the figure.
- [§3.4, Eq. (2)] Eq. (1) GFP is clear; Eq. (2) would benefit from an explicit statement that the same B=[1,t,t²] is fit independently to e and g (which is implied but not written).
- [§6.1] §6.1 video-level ranking (views/likes) is a useful negative control but sits somewhat apart from the moment-level claim; a one-sentence pointer in the introduction or discussion would help readers see why it is included.
- [Metadata / venue fit] Primary arXiv category is listed as cs.SE; for discoverability the authors may want cs.AI / q-bio.NC / cs.CV as cross-lists if the venue allows, since the contribution is brain encoding and engagement prediction rather than software engineering per se.
- [§1–§2] Minor prose: “theneuroforecastingliterature” and similar missing spaces appear in the introduction PDF text; check for tokenization artifacts before camera-ready.
Circularity Check
No significant circularity: null result is tested against independent external YouTube heatmaps and standard reductions of an off-the-shelf encoder.
full rationale
The paper's load-bearing claim is an empirical null (position-controlled partial r ≈ +0.058, not above loudness/motion, BF01=3.2, equivalence bound r≈0.14) between TRIBE's predicted cortical drive (GFP and network/ROI variants) and public YouTube most-replayed heatmaps on 48 held-out videos. TRIBE is used frozen from released weights trained on Algonauts fMRI; the behavioral target is external public metadata never used in TRIBE training or in the authors' own ridge encoders. GFP, partial correlation after quadratic/spline position basis, matched/mismatched CV probes, permutation nulls, and reliability ceilings are standard reductions and controls, not definitions of success in terms of the target. The supervised LOVO probe that appears to reach r=0.47 is explicitly shown to collapse to a shared temporal-shape artifact under stronger position control and is not treated as a positive result. Per-subject encoders for the ISC test are fit and validated on Algonauts (in-domain r≈0.15, cross-domain r≈0.10) then applied to the YouTube clips; this is ordinary transfer, not fitting the quantity being predicted. No self-definitional equations, no fitted parameter renamed as prediction of a closely related quantity, no load-bearing self-citation uniqueness claim, and no ansatz smuggled via citation. The derivation chain is therefore self-contained against external benchmarks.
Assumptions & free parameters
free parameters (5)
- analysis_window_seconds
- position_basis_order
- probe_n_components_and_ridge
- equivalence_delta
- per_subject_ridge_encoders
assumptions (6)
- domain assumption Global field power (RMS over vertices) is a reasonable unsupervised scalar for cortical drive/engagement.
- domain assumption YouTube most-replayed heatmaps are a valid passive proxy for moment-level re-watch engagement.
- domain assumption Position-controlled partial correlation isolates content-level moment prediction from shared temporal shape.
- domain assumption TRIBE released weights without fine-tuning, averaged over subject embeddings, preserve the behaviorally relevant structure if any exists in predicted cortex.
- standard math Standard Pearson/partial correlation, t-tests, circular-shift permutation, JZS Bayes factor, and TOST are appropriate for these autocorrelated series.
- ad hoc to paper Cortical-surface (fsaverage5) readouts are informative enough to test transfer even though NAcc is absent.
Cite this review
Pith. "Pith review of A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps." pith.science (2026). https://pith.science/paper/N6BF2J6F
@misc{pith2026260701400,
author = {Pith},
title = {Pith review of: A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps},
year = {2026},
howpublished = {\url{https://pith.science/paper/N6BF2J6F}},
note = {Machine review of arXiv:2607.01400}
}
read the original abstract
Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy; whether their predicted neural signals also forecast behavioral engagement is unknown. We run TRIBE, the winning model of the 2025 Algonauts challenge (Llama-3.2 + V-JEPA 2 + Wav2Vec-BERT), on 48 YouTube videos and reduce its predicted cortical response to a per-second engagement curve, the global field power. Correlated against each video's "most replayed" heatmap, a proxy for re-watch, it shows no evidence of prediction: the pooled position-controlled partial correlation is +0.058 (95% CI [-0.04, 0.15]; t(47)=1.21, p=0.23), and not above simple loudness/motion baselines. The raw correlation is also near zero; the moderate values for music videos are an onset-replay artifact. The null holds across six cortical-network readouts, value/salience ROIs, and a permutation test; a supervised leave-one-video-out probe appears to reach r=0.47 but collapses to a temporal-shape artifact under a proper position control. Running the probe on TRIBE's input streams reveals at most a small, borderline visual-stream signal (matched vs. mismatched p=0.004-0.06) and none in audio, text, or the predicted cortex. The inter-subject-correlation readout, the closest prior positive result, is unavailable from the subject-averaged released model, so we fit our own per-subject encoders on the Algonauts fMRI (validated in-domain at r=0.15 and cross-domain, Friends-to-film, at r=0.10); the predicted ISC still does not track re-watch (r=-0.04, p=0.34). We bound rather than merely fail to reject the null: a Bayes factor gives moderate evidence for it (BF01=3.2), an equivalence test excludes effects above r=0.14, and the target's split-half reliability (0.82; ceiling r=0.9) rules out a noisy-label artifact. We release code, a video-ID manifest, and a heatmap-acquisition method robust to YouTube's SABR streaming.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
G. S. Berns and S. E. Moore. A neural predictor of cultural popularity. Journal of Consumer Psychology, 22(1):154--160, 2012
2012
-
[3]
S. d'Ascoli, J. Rapin, Y. Benchetrit, H. Banville, and J.-R. King. TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction. arXiv:2507.22229, 2025
arXiv 2025
-
[4]
J. P. Dmochowski, M. A. Bezdek, B. P. Abelson, J. S. Johnson, E. H. Schumacher, and L. C. Parra. Audience preferences are predicted by temporal reliability of neural processing. Nature Communications, 5:4567, 2014
2014
-
[5]
Genevsky, C
A. Genevsky, C. Yoon, and B. Knutson. When brain beats behavior: Neuroforecasting crowdfunding outcomes. Journal of Neuroscience, 37(36):8625--8634, 2017
2017
-
[6]
A. T. Gifford, D. Bhattacharjee, M. St-Laurent, B. Pinsard, et al. The Algonauts Project 2025 challenge: How the human brain makes sense of multimodal movies. arXiv:2501.00504, 2025
arXiv 2025
-
[7]
A. Grattafiori et al. The Llama 3 herd of models. arXiv:2407.21783, 2024
arXiv 2024
-
[8]
Hasson, Y
U. Hasson, Y. Nir, I. Levy, G. Fuhrmann, and R. Malach. Intersubject synchronization of cortical activity during natural vision. Science, 303(5664):1634--1640, 2004
2004
Show all 11 references
-
[9]
A. G. Huth, W. A. de Heer, T. L. Griffiths, F. E. Theunissen, and J. L. Gallant. Natural speech reveals the semantic maps that tile human cerebral cortex. Nature, 532(7600):453--458, 2016
2016
-
[10]
Naselaris, K
T. Naselaris, K. N. Kay, S. Nishimoto, and J. L. Gallant. Encoding and decoding in fMRI. NeuroImage, 56(2):400--410, 2011
2011
-
[11]
Scholz, E
C. Scholz, E. C. Baek, M. B. O'Donnell, H. S. Kim, J. N. Cappella, and E. B. Falk. A neural model of valuation and information virality. Proceedings of the National Academy of Sciences, 114(11):2881--2886, 2017
2017
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.