Pith. sign in

REVIEW 3 major objections 7 minor 11 references

A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps

T0 review · 3 major / 7 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Predicted cortical drive from a top brain-encoding model does not track which moments of YouTube videos people re-watch, up to a bound of r≈0.14.

desk verdict A carefully bounded null: TRIBE predicted cortical drive does not track YouTube most-replayed beyond position and low-level baselines, with real artifact controls and released code. read the letter →

arxiv 2607.01400 v2 pith:N6BF2J6F submitted 2026-07-01 cs.SE cs.LGq-bio.NC

classification cs.SEcs.LGq-bio.NC
keywords brainencodingneuroforecastingfMRIpredictionYouTubemost-replayedTRIBEglobalfieldpowerinter-subjectcorrelationengagement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Deep multimodal models can now predict fMRI responses to naturalistic video with high accuracy. This paper asks whether those predicted neural signals also forecast real-world engagement, using YouTube “most replayed” heatmaps as a passive, moment-level proxy for re-watch. The authors run TRIBE, a trimodal encoder, on 48 videos, collapse its predicted cortex into a per-second global-field-power curve, and find no content-specific association after position control: the pooled partial correlation is near zero, no better than loudness or motion baselines, and robust across network and ROI readouts. A supervised cortical probe that first looks successful collapses into a shared temporal-shape artifact once position is controlled properly. A small, borderline video-specific signal appears only in the visual input stream and is lost by the encoding step; predicted inter-subject correlation, the closest prior positive route, also fails. The authors bound the null with a Bayes factor, an equivalence test, and a high target reliability ceiling, and they release the pipeline so others can test subcortical or cleaner-target variants.

What carries the argument

Position-controlled partial correlation of TRIBE global field power (root-mean-square over cortical vertices) against YouTube most-replayed heatmaps, with matched/mismatched leave-one-video-out probes and an equivalence/Bayes bound on the null.

What would settle it

A subcortex-inclusive encoder that includes nucleus accumbens, or a cleaner creator-side audience-retention target on a larger balanced sample, yielding a position-controlled partial correlation clearly above the loudness baseline and outside the current r≈0.14 equivalence bound.

Watch

Extended reading notes

Core claim

On this target and these readouts, a predicted-fMRI drive signal from TRIBE carries approximately no content-specific re-watch signal, up to a bound of r≈0.14. The pooled position-controlled partial correlation is +0.058 (95% CI [−0.04, 0.15]), indistinguishable from zero and from low-level baselines; the null holds for network GFP, signed value ROIs, permutation tests, and predicted ISC, while an apparent supervised r=0.47 is a shared temporal-shape artifact.

Load-bearing premise

That most-replayed heatmaps on already-popular clips, over a 60-second window and compared mainly via cortical-surface summaries, are a fair enough behavioral target to detect neuroforecasting transfer if it existed.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The manuscript tests whether predicted cortical responses from TRIBE (the 2025 Algonauts winner) forecast moment-level YouTube re-watch behavior. TRIBE is run unmodified on 48 videos; its predicted surface response is reduced to a per-TR global field power (GFP) engagement curve and compared to each video’s most-replayed heatmap. The primary metric is a position-controlled partial correlation (quadratic detrend). The pooled partial r is +0.058 (95% CI [−0.04, 0.15]; t(47)=1.21, p=0.23), not above loudness/motion baselines, and near zero in raw form outside a music-video onset artifact. The null is stable across network and signed value/salience ROI readouts, a circular-shift permutation, and a supervised leave-one-video-out cortical probe that only appears successful (r=0.47) under a coarse detrend and collapses under spline + matched/mismatched controls. A borderline content-specific signal is reported only in the visual input stream; audio, text, and predicted cortex show none. Because the released TRIBE checkpoint is subject-averaged, the authors fit their own per-subject encoders and find that predicted ISC also fails to track re-watch (r=−0.04). They bound rather than merely fail to reject the null (BF01=3.2; effects ≳0.14 excluded; target split-half reliability ≈0.82). Code, video-ID manifest, and a SABR-resilient acquisition method are released.

Significance. If the bounded null holds, the paper is a useful negative result for anyone tempted to treat modern multimodal brain encoders (or their foundation features) as off-the-shelf engagement predictors. Its main contributions are (i) a carefully scoped claim with Bayes factor, equivalence-style bound, and reliability ceiling rather than a bare non-significant p-value; (ii) a concrete methodological caution that most-replayed labels plus a coarse detrend can manufacture large cross-validated correlations that are pure shared temporal shape; and (iii) a mechanistic sketch that group-trained encoding collapses the weak video-specific structure present in visual features. Full release of code, video IDs, and an acquisition pipeline that works under SABR is a genuine reproducibility strength and should be credited. The result is narrow by design (one model, cortical surface, biased target, 60 s window), but within that scope it is informative.

major comments (3)
  1. [§6.4, Abstract, Conclusion] §6.4 vs abstract/conclusion: the pre-specified TOST against δ=0.10 does not establish equivalence (p=0.20). The claim that effects larger than r≈0.14 are excluded comes from the descriptive 90% CI, not from a successful TOST at a pre-registered SEI. Please state explicitly what was pre-specified versus post-hoc, and align the abstract wording (“an equivalence test excludes effects above r≈0.14”) with §6.4 so readers do not read a failed TOST as a passed equivalence test. BF01=3.2 and the reliability ceiling already support a bounded null; the equivalence language just needs to match the procedure.
  2. [§3.2, §5, §6.8, §8] §3.2 and §5 fix the analysis to the first 60 s (≈60 TRs) of every clip. Most-replayed markers span the full video, and intros/onsets are exactly the position artifact the paper works hard to remove. §6.8 shows that marker density and duration do not correlate with the effect, which is helpful, but does not test whether the null is specific to the opening minute. A sensitivity check—e.g., a mid-clip 60 s window, or full-length analysis where TRIBE segment onsets allow—would make the “moment-level” claim less dependent on a single, intro-heavy window. If that is infeasible for encoding cost, state the limitation more prominently in §8 and qualify the claim accordingly.
  3. [§6.7] §6.7 is presented as a direct test of the closest prior positive result (ISC). The custom per-subject ridge encoders are validated at only r≈0.15 in-domain and r≈0.10 cross-domain. That is within the published range for feature-to-cortex models, but it leaves open how large a true ISC–rewatch effect could have been and still been missed. Please add a brief power or attenuation note (e.g., expected attenuation of a behavioral correlation given encoder r≈0.10–0.15) so the ISC null is interpretable as evidence against transfer rather than only as a weak encoder. The paper already notes that the released TRIBE checkpoint cannot supply per-subject predictions; that disclosure is good and should stay.
minor comments (7)
  1. [Table 1] Table 1 leaves raw r blank for the loudness and motion baselines; either fill those cells or note in the caption that only partial r is reported for baselines by design.
  2. [Figure 1c] Figure 1c category labels include “?” for n=4; replace with the “misc” label used in §5 for consistency.
  3. [Figure 2a, §6.4] Figure 2a labels the equivalence region ±0.139 while the text uses r≈0.14; pick one rounding and use it throughout abstract, §6.4, and the figure.
  4. [§3.4, Eq. (2)] Eq. (1) GFP is clear; Eq. (2) would benefit from an explicit statement that the same B=[1,t,t²] is fit independently to e and g (which is implied but not written).
  5. [§6.1] §6.1 video-level ranking (views/likes) is a useful negative control but sits somewhat apart from the moment-level claim; a one-sentence pointer in the introduction or discussion would help readers see why it is included.
  6. [Metadata / venue fit] Primary arXiv category is listed as cs.SE; for discoverability the authors may want cs.AI / q-bio.NC / cs.CV as cross-lists if the venue allows, since the contribution is brain encoding and engagement prediction rather than software engineering per se.
  7. [§1–§2] Minor prose: “theneuroforecastingliterature” and similar missing spaces appear in the introduction PDF text; check for tokenization artifacts before camera-ready.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: null result is tested against independent external YouTube heatmaps and standard reductions of an off-the-shelf encoder.

full rationale

The paper's load-bearing claim is an empirical null (position-controlled partial r ≈ +0.058, not above loudness/motion, BF01=3.2, equivalence bound r≈0.14) between TRIBE's predicted cortical drive (GFP and network/ROI variants) and public YouTube most-replayed heatmaps on 48 held-out videos. TRIBE is used frozen from released weights trained on Algonauts fMRI; the behavioral target is external public metadata never used in TRIBE training or in the authors' own ridge encoders. GFP, partial correlation after quadratic/spline position basis, matched/mismatched CV probes, permutation nulls, and reliability ceilings are standard reductions and controls, not definitions of success in terms of the target. The supervised LOVO probe that appears to reach r=0.47 is explicitly shown to collapse to a shared temporal-shape artifact under stronger position control and is not treated as a positive result. Per-subject encoders for the ISC test are fit and validated on Algonauts (in-domain r≈0.15, cross-domain r≈0.10) then applied to the YouTube clips; this is ordinary transfer, not fitting the quantity being predicted. No self-definitional equations, no fitted parameter renamed as prediction of a closely related quantity, no load-bearing self-citation uniqueness claim, and no ansatz smuggled via citation. The derivation chain is therefore self-contained against external benchmarks.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The claim is empirical, not axiomatic derivation. Load-bearing background is standard encoding/neuroforecasting practice plus operational choices (GFP, 60s window, position basis, popular-video sample). No new physical entities are postulated; TRIBE and foundation encoders are prior artifacts. Free parameters are analysis hyperparameters and probe settings that could shift secondary numbers but are stress-tested by multiple readouts and permutation/Bayes bounds.

free parameters (5)
  • analysis_window_seconds
    First 60s (~60 TRs) retained per video; longer clips and full-timeline dynamics are not tested and could change moment-level associations.
  • position_basis_order
    Primary partial correlation uses quadratic B=[1,t,t^2]; supervised probe later uses cubic spline with 7 df. Choice of detrend strength determines whether the r=0.47 artifact appears.
  • probe_n_components_and_ridge
    Supervised cortex/input probes reduce to 100 PCs and fit ridge under leave-one-video-out; these knobs affect apparent CV r though matched/mismatched controls still null the cortex.
  • equivalence_delta
    TOST uses pre-specified SEOI δ=0.10; descriptive exclusion of effects above ~0.14 depends on this choice and N=48 resolution.
  • per_subject_ridge_encoders
    ISC arm uses authors’ own ridge encoders (not TRIBE’s missing per-subject head), validated at r≈0.15 in-domain / 0.10 cross-domain; encoder capacity and feature set are free modeling choices.
assumptions (6)
  • domain assumption Global field power (RMS over vertices) is a reasonable unsupervised scalar for cortical drive/engagement.
    Introduced in §3.2 as the primary readout; later relaxed via network and signed ROI variants that remain null.
  • domain assumption YouTube most-replayed heatmaps are a valid passive proxy for moment-level re-watch engagement.
    §3.3 defines the target; §2 and §8 acknowledge bias from intros, chapters, and seek-back.
  • domain assumption Position-controlled partial correlation isolates content-level moment prediction from shared temporal shape.
    Primary metric in §3.4–3.5; strengthened to splines in §6.5–6.6 when quadratic control proved insufficient for the probe.
  • domain assumption TRIBE released weights without fine-tuning, averaged over subject embeddings, preserve the behaviorally relevant structure if any exists in predicted cortex.
    §3.1 operational setup; the paper’s mechanism claim is that group averaging removes that structure.
  • standard math Standard Pearson/partial correlation, t-tests, circular-shift permutation, JZS Bayes factor, and TOST are appropriate for these autocorrelated series.
    Used throughout §3.5 and §6.3–6.4; permutation preserves autocorrelation as a nonparametric check.
  • ad hoc to paper Cortical-surface (fsaverage5) readouts are informative enough to test transfer even though NAcc is absent.
    §6.2 and §8 explicitly note subcortical absence; the central claim is scoped to available cortical predictions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps." pith.science (2026). https://pith.science/paper/N6BF2J6F

@misc{pith2026260701400,
  author       = {Pith},
  title        = {Pith review of: A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N6BF2J6F}},
  note         = {Machine review of arXiv:2607.01400}
}
read the original abstract

Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy; whether their predicted neural signals also forecast behavioral engagement is unknown. We run TRIBE, the winning model of the 2025 Algonauts challenge (Llama-3.2 + V-JEPA 2 + Wav2Vec-BERT), on 48 YouTube videos and reduce its predicted cortical response to a per-second engagement curve, the global field power. Correlated against each video's "most replayed" heatmap, a proxy for re-watch, it shows no evidence of prediction: the pooled position-controlled partial correlation is +0.058 (95% CI [-0.04, 0.15]; t(47)=1.21, p=0.23), and not above simple loudness/motion baselines. The raw correlation is also near zero; the moderate values for music videos are an onset-replay artifact. The null holds across six cortical-network readouts, value/salience ROIs, and a permutation test; a supervised leave-one-video-out probe appears to reach r=0.47 but collapses to a temporal-shape artifact under a proper position control. Running the probe on TRIBE's input streams reveals at most a small, borderline visual-stream signal (matched vs. mismatched p=0.004-0.06) and none in audio, text, or the predicted cortex. The inter-subject-correlation readout, the closest prior positive result, is unavailable from the subject-averaged released model, so we fit our own per-subject encoders on the Algonauts fMRI (validated in-domain at r=0.15 and cross-domain, Friends-to-film, at r=0.10); the predicted ISC still does not track re-watch (r=-0.04, p=0.34). We bound rather than merely fail to reject the null: a Bayes factor gives moderate evidence for it (BF01=3.2), an equivalence test excludes effects above r=0.14, and the target's split-half reliability (0.82; ceiling r=0.9) rules out a noisy-label artifact. We release code, a video-ID manifest, and a heatmap-acquisition method robust to YouTube's SABR streaming.

Figures

Figures reproduced from arXiv: 2607.01400 by the authors.

Figure 1
Figure 1. No content-level prediction of re-watch behavior. (a) Per-video raw and position￾controlled correlations with most-replayed; the partial correlation (mean ± 95% CI) is centered on zero and the CI crosses it. (b) Pooled partial correlation: TRIBE is statistically indistinguishable from the loudness baseline and near zero. (c) Per-category partial correlations are small, sign￾inconsistent, and dominated by noise at sm… view at source ↗
Figure 2
Figure 2. Bounding the null, and locating where the signal is lost. (a) The pooled partial correlation (point ± 95% CI) lies inside an equivalence region of 𝛿=0.14; a Bayes factor gives BF01=3.2 (moderate evidence for the null). (b) Every readout is near zero: whole-cortex GFP, five functional networks, and signed value/salience ROIs (*). (c) Under the spline position control, matched vs. mismatched CV correlation by source: … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references · 4 linked inside Pith

  1. [1]

    Assran, A

    M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Muckley, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv:2506.09985, 2025

  2. [2]

    G. S. Berns and S. E. Moore. A neural predictor of cultural popularity. Journal of Consumer Psychology, 22(1):154--160, 2012

  3. [3]

    d'Ascoli, J

    S. d'Ascoli, J. Rapin, Y. Benchetrit, H. Banville, and J.-R. King. TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction. arXiv:2507.22229, 2025

  4. [4]

    J. P. Dmochowski, M. A. Bezdek, B. P. Abelson, J. S. Johnson, E. H. Schumacher, and L. C. Parra. Audience preferences are predicted by temporal reliability of neural processing. Nature Communications, 5:4567, 2014

  5. [5]

    Genevsky, C

    A. Genevsky, C. Yoon, and B. Knutson. When brain beats behavior: Neuroforecasting crowdfunding outcomes. Journal of Neuroscience, 37(36):8625--8634, 2017

  6. [6]

    A. T. Gifford, D. Bhattacharjee, M. St-Laurent, B. Pinsard, et al. The Algonauts Project 2025 challenge: How the human brain makes sense of multimodal movies. arXiv:2501.00504, 2025

  7. [7]

    Grattafiori et al

    A. Grattafiori et al. The Llama 3 herd of models. arXiv:2407.21783, 2024

  8. [8]

    Hasson, Y

    U. Hasson, Y. Nir, I. Levy, G. Fuhrmann, and R. Malach. Intersubject synchronization of cortical activity during natural vision. Science, 303(5664):1634--1640, 2004

Show all 11 references
  1. [9]

    A. G. Huth, W. A. de Heer, T. L. Griffiths, F. E. Theunissen, and J. L. Gallant. Natural speech reveals the semantic maps that tile human cerebral cortex. Nature, 532(7600):453--458, 2016

  2. [10]

    Naselaris, K

    T. Naselaris, K. N. Kay, S. Nishimoto, and J. L. Gallant. Encoding and decoding in fMRI. NeuroImage, 56(2):400--410, 2011

  3. [11]

    Scholz, E

    C. Scholz, E. C. Baek, M. B. O'Donnell, H. S. Kim, J. N. Cappella, and E. B. Falk. A neural model of valuation and information virality. Proceedings of the National Academy of Sciences, 114(11):2881--2886, 2017

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.