{"id":"9288df22-cafe-46cd-a9a5-2cc275eb744c","arxiv_id":"2608.07282","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"CLIP attribution scores, normalized and aggregated over objects, reproduce the anticipatory gaze effect of Altmann and Kamide (1999) without any training on eye-tracking data.","lead":"This paper shows that off-the-shelf CLIP image-text models, used with no fine-tuning, can reproduce a classic eye-tracking finding: listeners look at an object before it is named when the verb predicts it. This is an existence proof that a simple contrastive encoder can mimic anticipatory human gaze, with implications for how predictive language processing may arise from static correlational structure.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported predictive-gaze effect is not robust across normalization choices: ReLU was selected post hoc as the one transform that produces the pattern, and the population-level claim is fit only on that output.","rationale":"The reader's weakest assumption was the general linking hypothesis that CLIP attributions correspond to human fixations; I agree that this is a deep issue, but it is explicitly acknowledged in §3.2 and §6 as an approximation, and the paper's stated goal is qualitative replication rather than quantitative fixation prediction. The more concrete and testable vulnerability is the post hoc selection of ReLU normalization, which the reader mentions in the rationale but does not elevate to the top concern. The evidence in Appendix B makes this concern concrete: the effect that forms the central claim disappears or weakens under other reasonable normalizations, and the population-level analysis is only reported for ReLU. This is a researcher-degrees-of-freedom problem that a single additional analysis can settle. If the population interaction remains significant under Shift, the concern weakens and the current CONDITIONAL verdict might move toward ACCEPT; if not, the claim of robustness fails. I therefore keep CONDITIONAL, but the condition is now sharpened: the authors must demonstrate that the result is not an artifact of the chosen normalization and exclusion policy.","tokens_in":13375,"tokens_out":2746,"duration_ms":31159,"concrete_test":"Fit the population-level mixed-effects model of §5 to target preferences computed with the Shift normalization and with the raw (None) normalization, using the same random effects structure, and report the verb×region interaction term (β, p). Also re-fit the ReLU-based global model with the excluded OpenCLIP ViT-B/16 (dfn2b) checkpoint included. If the interaction is no longer significant under Shift or after including the excluded checkpoint, the headline result is conditional on a post hoc pipeline choice and the central claim is not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (constraining verb induces predictive target preference that attenuates at the noun) is evaluated exclusively on ReLU-normalized attributions. Section 3.5 presents ReLU as merely one of several plausible ways to make attributions non-negative, but Appendix B shows that the qualitative result is strongly transformation-dependent: with ReLU, all four models show a significant pre-object constraining-verb effect; with Shift, only three of four do; with Softmax, only one; and with no normalization, none do. The global mixed-effects model that supports the headline interaction (verb×region, β=−0.235, p=0.005) is fit only on ReLU outputs, so the 'across models' claim is really 'across models under the normalization that best matches the desired outcome.' This is a garden-of-forking-paths risk: ReLU won over three competitors after inspecting results, and the paper offers no a priori theoretical reason why clamping negative attributions to zero (rather than shifting, softmaxing, or not transforming) is the correct linking function. Additionally, a fifth checkpoint (OpenCLIP ViT-B/16, dfn2b) was excluded because its attributions were degenerate under this pipeline, further tightening the set of models on which the claim is demonstrated. The load-bearing premise is therefore that ReLU is a principled, not merely convenient, mapping from attribution scores to fixation saliencies; the paper does not establish this, and the reported robustness across models does not extend across the plausible space of preprocessing choices.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a method to model gaze in visual world experiments without any fine-tuning or human gaze training data. It encodes each image with each linguistic prefix of the sentence using off-the-shelf CLIP bi-encoders, computes Integrated Jacobians attributions between text tokens and image patches, applies a ReLU normalization to the per-patch attributions, and aggregates them over manually annotated bounding boxes to obtain a target-preference score. The authors test the approach on the 18 stimuli of Altmann and Kamide (1999) with four CLIP checkpoints and report that a constraining verb induces a significantly stronger target-object preference before the noun is encountered than a non-constraining verb, and that this advantage attenuates at the noun, mirroring human anticipatory gaze. They interpret this as evidence that predictive processing can emerge from correlational structure in contrastively trained encoders without a generative or autoregressive objective.","tokens_in":13625,"tokens_out":4274,"duration_ms":43961,"significance":"If the result holds, it is a valuable contribution to computational psycholinguistics: it shows that a simple, non-generative, off-the-shelf vision-language encoder can qualitatively reproduce a seminal anticipatory eye-movement effect without being trained for incremental language processing. The paper has several genuine strengths: the prefix-presentation design in Section 3.6 prevents the encoder from using future context, which is a clean solution to a real leakage problem; the evaluation uses multiple CLIP checkpoints and standard mixed-effects analyses; the failure-mode analysis in Section 5 is honest and informative; and the authors state that code and materials will be released. The central empirical claim, however, rests on two analytic choices that are not robustly justified: the post hoc selection of ReLU normalization and the exclusion of one checkpoint based on its outcome. These issues are fixable with additional sensitivity analyses or pre-specified decision rules, and the linking hypothesis needs either validation against human gaze data or a more careful scoping of the claim.","major_comments":[{"comment":"The choice of ReLU normalization is load-bearing but was made after inspecting the results: the text states that ReLU \"turned out to be the best choice,\" and Table 3 shows that the pre-object constraining-verb effect appears in all four models only under ReLU (three models under Shift, one under Softmax, and none under no normalization). The population-level interaction reported in Section 5 (β = −0.235, p = 0.005) is computed exclusively on ReLU-normalized attributions, so the \"across models\" claim is not robust to an analytic choice that lacks an a priori theoretical motivation. Please either justify ReLU as the linking function on independent grounds, or report the global mixed-effects model under all four normalization strategies and discuss the resulting sensitivity.","section":"Section 3.5 / Appendix B, Table 3"},{"comment":"The exclusion of the fifth checkpoint (OpenCLIP ViT-B/16, dfn2b) is based on the outcome of interest: its attributions are described as degenerate because the target object receives negative attribution at verb onset. This is a selection rule that can only make the reported pattern easier to obtain. Please report the results with this checkpoint included under the same pipeline, or motivate the exclusion with a pre-specified, outcome-independent quality criterion such as a failure of the attribution method to satisfy a basic sanity check.","section":"Footnote 6 / Section 4.2"},{"comment":"The paper claims that normalized Integrated Jacobians attributions \"are directly linearly related to human fixations\" (Section 1) and that the resulting probabilities \"can be interpreted as fixation saliencies\" (Section 3.1), but the linking hypothesis is not validated against any human gaze data. For the simple static materials this may be \"arguably sufficient,\" as Section 3.2 says, but the headline claim that gaze behavior \"can be modeled\" is stronger than what is demonstrated. I recommend either validating the mapping against empirical fixation proportions (e.g., from the original Altmann and Kamide data or a comparable visual world study) or explicitly scoping the claim to attribution-derived target preference and supporting the linking assumption with evidence from prior work.","section":"Section 3.2 / Section 6, Limitations"},{"comment":"The dependent variable depends directly on manually annotated bounding boxes plus a one-patch margin, but the paper reports no reliability assessment or sensitivity analysis for this step. A one-patch change in box boundaries can change the aggregated target preference and therefore the statistical results. Please provide annotation details, inter-annotator agreement if applicable, or a robustness check that varies the margin size and shows that the qualitative conclusions are unchanged.","section":"Section 3.5 / Section 4.3"}],"minor_comments":[{"comment":"The capitalization of ReLU is inconsistent (\"ReLU\" in the text and figures, \"ReLu\" in Section 3.5 and Appendix B); please standardize.","section":"Throughout"},{"comment":"The caption of Figure 1 describes the panels as attribution heatmaps, while the text in Section 1 says the colored overlay indicates fixation probabilities; please clarify that the figure shows model-derived attributions, not human data.","section":"Figure 1 and Section 1"},{"comment":"Section 6 cites Möller et al. (2024) for the finding that CLIP models identify object categories, but Section 3.2 attributes this finding to Möller et al. (2025); please check which reference is intended and make the citations consistent.","section":"Section 6"},{"comment":"There is a typographical error in the model size discussion: \"VIT-B/16\" should be \"ViT-B/16.\"","section":"Section 5"},{"comment":"The sentence \"This dataset consists of 18 images which is considered as dataset size with sufficient statistical power for its experimental design\" is awkward; consider rephrasing to state that the original study used 18 items and that the authors follow that design.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for a computational psycholinguistics venue, and the core idea is interesting and clearly presented. The main risk is analytic flexibility: the normalization choice and the checkpoint exclusion are both selected after seeing the outcome, and the linking hypothesis is asserted rather than validated. These issues are addressable with additional experiments or sensitivity analyses, so I would not reject the paper, but I would not accept it in its current form. I would also encourage the editor to ask the authors to make the code and annotation files available at revision time."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper is that the core idea is both simple and new: take an off-the-shelf CLIP bi-encoder, run Integrated Jacobians, and read the token-by-patch attributions as fixation saliencies. Nobody has done that for the visual-world paradigm, and it works well enough to reproduce the Altmann–Kamide anticipation pattern without fine-tuning. That is a real contribution, both methodologically and conceptually, because it suggests predictive gaze can fall out of static correlational structure rather than a generative objective.\n\nWhat the paper does well: it is transparent about its own decisions. The appendix gives results for four normalization schemes, the failure modes are discussed with heatmaps, and the limitations section openly states the linking assumption and the non-incrementality of the encoders. The mixed-effects analysis is appropriate for 18 items, and the population-level interaction (verb type by region) is the right way to frame the claim. The authors also correctly note that with only four models, the model-size correlation is not statistically testable.\n\nThe soft spots are real and proportionally serious. Appendix B shows the result is transformation-dependent: with ReLU all four models show the pre-object effect; with Shift only three; with Softmax one; with none, zero. Since ReLU was selected after looking at the results, this is a garden-of-forking-paths concern. The paper says ReLU is \"the best choice\" but gives no theoretical reason why clamping negatives to zero is the correct linking function. The exclusion of the fifth checkpoint is reported honestly, but it still tightens the club of models that work. The linking hypothesis is asserted as \"arguably sufficient\" rather than validated; there is no direct comparison to human fixation probabilities, only a qualitative match to one classic experiment. None of these flaws refute the central finding, but they do mean the word \"robustly\" in the abstract is doing more work than the evidence supports.\n\nWho is this for? Computational psycholinguists and anyone working on multimodal benchmarks, especially those interested in whether encoders can serve as process models. It deserves a serious referee: the method is cheap, easy to extend to other materials, and the paper is honest enough that the flaws are addressable. I would accept it for review and ask the authors to pre-register or justify the normalization choice, and to report all models and all normalizations without selective emphasis. If the ReLU-specificity is confirmed as a real constraint, that is still an interesting finding, just a narrower one.\n\nBring it to reading group—it will generate a good argument about what counts as a linking hypothesis.","headline":"Genuinely novel zero-shot method for visual-world gaze, with a plausible qualitative replication that is unfortunately dependent on a post hoc normalization choice; worth serious engagement, but the robustness claim needs revision.","tokens_in":14182,"tokens_out":1893,"would_cite":true,"duration_ms":21891,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No fine-tuning: CLIP attributions replicate predictive gaze in visual-world experiments.","keywords":["visual world paradigm","CLIP","Integrated Jacobians","predictive processing","gaze modeling","attribution methods","vision-language models","psycholinguistics"],"falsifier":"Record human eye movements on a second set of visual-world stimuli with the same Altmann-Kamide design, run the identical attribution pipeline on those stimuli, and check whether the item-level verb-type by region interaction in CLIP attributions predicts the item-level interaction in human target preference; if the model interaction does not track the human interaction across items, the linking hypothesis fails.","tokens_in":13135,"feed_emoji":"👀","tokens_out":4539,"duration_ms":45242,"temperature":0.7,"pith_summary":"This paper claims that off-the-shelf CLIP vision-language encoders, paired with an attribution method called Integrated Jacobians, can reproduce the anticipatory gaze pattern observed in human visual-world experiments. When the input contains a constraining verb such as \"eat,\" the model assigns higher attribution to the target object (the cake) before that object is named, and this advantage disappears once the object is mentioned. The effect holds across four CLIP models without any fine-tuning or generative decoding, suggesting that predictive gaze behavior can emerge from correlational multimodal similarity rather than dedicated prediction machinery.","feed_headline":"No fine-tuning: CLIP attributions replicate predictive gaze","feed_subtitle":"Integrated Jacobian scores on a vision-language encoder reproduce the Altmann-Kamide constraining-verb effect before the noun is spoken.","key_machinery":"Integrated Jacobians, a generalization of Integrated Gradients to dual encoders, decompose CLIP's similarity score between a caption and an image into attribution scores for each token and image patch. ReLU normalization clamps negative attributions to zero so the scores behave like fixation saliencies, and bounding-box aggregation sums patch scores into per-object probabilities with a one-patch margin. Crucially, the encoder sees only each prefix of the linguistic input in isolation, preventing the attribution from using the sentence-final object as context.","core_discovery":"The paper's central claim is that ReLU-normalized Integrated Jacobian attributions from CLIP models qualitatively replicate the time-course of human predictive processing in the Altmann-Kamide stimuli: a constraining verb reliably drives stronger target-object preference in the pre-object region, and this advantage attenuates once the object noun is encountered. Across the four models, the population-level interaction between verb type and region is significant and negative, confirming that the predictive advantage is localized to the window before the object is named.","pith_inferences":["A direct quantitative comparison against recorded fixation proportions on the same or new stimuli would be a stronger test than the qualitative interaction reported here; the paper's linking hypothesis does not yet license precise probability predictions.","The same pipeline could be pointed at other visual-world phenomena, such as semantic competitors or color-based preactivation, to test whether CLIP attributions capture featural anticipation and not just selectional restrictions of verbs.","Varying the one-patch bounding-box margin and comparing against human-defined object regions would show how much of the reported interaction depends on the annotation choice, which the normalization ablation does not cover."],"forward_implications":["The qualitative time course of human predictive gaze can be reproduced with no training, no fine-tuning, and no generative decoder.","Because the CLIP language encoder is non-incremental and sees each prefix separately, the predictive signal is attributable to the constraining verb's content rather than to incremental processing dynamics.","Larger, higher-accuracy CLIP models tended to show larger pre-object effects in the four models tested, though the paper treats this pattern as hypothesis-generating given the small number of checkpoints.","The prediction step itself needs no object labels or bounding boxes; bounding boxes are used only for evaluation aggregation, so the approach avoids manual object categorization during prediction."],"supporting_citations":[{"why":"Supplies the visual-world stimuli and the human predictive-processing pattern that the paper aims to replicate.","marker":"Altmann and Kamide (1999)"},{"why":"Introduces the CLIP bi-encoder architecture whose similarity scores are the substrate for attribution.","marker":"Radford et al. (2021)"},{"why":"Introduces Integrated Jacobians, the method that decomposes dual-encoder similarity into token-patch attribution scores.","marker":"Möller et al. (2024)"},{"why":"Shows CLIP models acquire object-category knowledge from matching captions, motivating why attributions mark relevant objects, and informs model-difference observations.","marker":"Möller et al. (2025)"},{"why":"Defines Integrated Gradients, the axiomatic attribution framework that Integrated Jacobians generalizes.","marker":"Sundararajan et al. (2017)"},{"why":"Provides the earlier computational model of visual-world gaze and the linking-assumption precedent that fixations reflect linguistic processing.","marker":"Allopenna et al. (1998)"},{"why":"Guides the maximal random-effects structure used in the mixed-effects statistical analysis.","marker":"Barr et al. (2013)"}],"fun_headline_variants":["CLIP without fine-tuning predicts human gaze","Off-the-shelf CLIP replicates predictive eye movements","Vision-language encoders model visual world gaze","No training: CLIP attributions match gaze"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim stands on the assumption that ReLU-normalized, bounding-box-aggregated CLIP attribution scores can stand in for human fixation probabilities; the paper calls this \"arguably sufficient\" for simple static displays and does not validate it against recorded eye movements.","fun_headline_variants_meta":{"raw":{"variants":["CLIP without fine-tuning predicts human gaze","Off-the-shelf CLIP replicates predictive eye movements","Vision-language encoders model visual world gaze","No training: CLIP attributions match gaze"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1473,"prompt_tokens":795,"completion_tokens":678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":411,"completion_tokens_details":{"reasoning_tokens":620}},"tokens_in":411,"tokens_out":678,"duration_ms":6732,"temperature":1.0,"reasoning_tokens":620,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:05:40.235543+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record human eye movements on a second set of visual-world stimuli with the same Altmann-Kamide design, run the identical attribution pipeline on those stimuli, and check whether the item-level verb-type by region interaction in CLIP attributions predicts the item-level interaction in human target preference; if the model interaction does not track the human interaction across items, the linking hypothesis fails.","supporting_citations":[{"cited_title":"and Magnuson, James S","cited_arxiv_id":null,"evidence_quote":"Provides the earlier computational model of visual-world gaze and the linking-assumption precedent that fixations reflect linguistic processing."}],"review_version":1}