{"id":"38b3e8c8-c9ee-4419-8755-c1059f419d83","arxiv_id":"2506.13409","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Neural and behavioral responses differ modestly across Exact, Substitute, Complement, and Irrelevant product recommendations, with the smallest neural separation between Exact and Substitute items.","lead":"This study measured brain activity and self-report ratings while participants evaluated product recommendations labeled Exact, Substitute, Complement, or Irrelevant. It reports that behavioral ratings differ across categories and that EEG classifiers slightly exceed chance, suggesting brain signals might one day guide personalized recommender systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The neural-signature claim is confounded by the oddball design: Exact targets are always preceded by Exact context products, so category is entangled with novelty/repetition.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the oddball design confounds recommendation category with novelty/repetition. This is the most serious threat to the paper's central neural claim because it attributes EEG differences to semantic category when the stimulus sequence guarantees that Exact targets differ from all other targets along a well-known cognitive dimension (stimulus probability/novelty). The proposed k-dependent analysis would test whether the observed effects scale with oddball strength, a concrete, data-internal check. The reader's REJECT verdict is appropriate: the neural evidence as presented cannot support the abstract's claim without either redesigning the experiment to orthogonalize category and novelty or substantially reinterpreting the results as novelty-driven. The behavioral findings remain useful, but they are not the paper's stated primary contribution. No change to the reader's verdict is warranted.","tokens_in":19622,"tokens_out":4113,"duration_ms":47745,"concrete_test":"Exploit the fact that the number of preceding Exact context products k varies from 1 to 3. Recompute the within-subject classification accuracy and the ERP amplitude differences for Exact vs Irrelevant (and Exact vs Substitute) separately for k=1, k=2, and k=3. If the oddball confound drives the Exact-vs-other differences, accuracy should increase monotonically with k, because a non-Exact target becomes a more salient deviant after more Exact primes; if the differences reflect stable category identity, accuracy should be roughly k-invariant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that distinct recommendation categories evoke differentiable neural signatures rests on the within-subject and between-subject EEG analyses in Sections 3.7.1 and 3.7.2. However, the experimental protocol in Sections 3.3.4 and 3.4 presents a search query followed by 1–3 Exact products before every target recommendation. Consequently, an Exact target is a repetition of the context class, while Substitute, Complement, and Irrelevant targets are oddball deviations. The paper explicitly invokes the oddball paradigm, and the P300/novelty response is a well-established EEG generator. Any above-chance classification involving Exact (e.g., Exact vs Irrelevant at 0.545, Table 2) could therefore reflect target-vs-context oddball status rather than the semantic recommendation category. The claim that Exact and Substitute are most similar is especially vulnerable: the low accuracy for that pair could arise from cancellation between category similarity and oddball-driven dissimilarity, making the reported similarity hard to interpret as a neural property of the categories rather than an artifact of the design. The behavioral ratings are not similarly confounded, but the abstract's neural claims, which are the paper's stated novelty, are not supported without controlling for this asymmetry. Post-hoc exclusion of non-significant subjects from accuracy averages (Tables 2 vs 3) further weakens the statistical evidence, but the design confound is the more fundamental threat to internal validity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a controlled EEG and behavioral study (N=21, 19 after exclusions) in which participants viewed a search query, a sequence of one to three Exact products, and then a target product labeled Exact, Substitute, Complement, or Irrelevant according to the ESCI benchmark. The authors report within-subject pairwise SVM classification accuracies (Tables 2 and 3), between-subject ERP-based meta-classifier accuracies (Table 4), behavioral ratings for relevance, likelihood of purchase, and diversity (Table 5), and engagement metrics including frontal alpha asymmetry (FAA), theta, and beta power (Section 4.3). Their central claims are that recommendation categories evoke differentiable neural signatures, that Exact and Substitute evoke the most similar neural responses, that behavioral ratings distinguish Exact and Irrelevant at the extremes with Substitute and Complement in between, and that FAA/theta/beta index category-dependent engagement. The behavioral analysis is largely convincing, but the neural analyses are compromised by a stimulus-design confound, a post-hoc subject-exclusion procedure, and an internal contradiction in the FAA interpretation.","tokens_in":19835,"tokens_out":3929,"duration_ms":43689,"significance":"If the neural claims were valid, the paper would be a useful first step toward using EEG to study how different recommendation types, beyond binary relevance, are processed, with potential implications for personalized recommender systems. The paper has genuine strengths: it uses real e-commerce data (ESCI labels), a within-subjects design, standard nonparametric tests for the behavioral ratings, and a rich feature exploration. However, the central neural-signature claim is not currently supported. The oddball design means that the target category is confounded with stimulus novelty, the reported classification accuracies are close to chance and partly reflect post-hoc exclusions, and the FAA result as written contradicts the equation defining the measure. The behavioral findings are solid and could stand on their own, but the abstract and contributions place the neural novelty first, so the main scientific contribution is not established.","major_comments":[{"comment":"The design confounds recommendation category with oddball novelty. Every trial presents a query followed by one to three Exact context products before the target recommendation, so an Exact target is a repetition of the context class while Substitute, Complement, and Irrelevant targets are oddball deviations. This is explicitly an oddball paradigm (Section 3.4, citing Sutton et al. 1965). Any EEG differentiation involving Exact, such as the Exact vs. Irrelevant accuracies in Tables 2 and 3 (0.545 and 0.599), could be driven by target-vs-context novelty rather than by semantic recommendation category. The paper's further claim that Exact and Substitute are the most similar is especially vulnerable, because that comparison would pit category similarity against oddball-driven dissimilarity and could yield low accuracy for either reason. The authors need a control condition in which all four categories appear as context products, or an analysis that isolates category identity from repetition/novelty; without this, the neural claims in RQ1 are not supported.","section":"§3.3.4 and §3.4"},{"comment":"The reported above-chance accuracies are inflated by post-hoc exclusion of non-significant subjects. Table 3 reports mean accuracies after removing individual models with near-chance performance, but the criterion for exclusion appears to be applied after seeing the results, and no pre-registered rule or multiple-comparison correction is described. The unexcluded accuracies in Table 2 are only 0.545–0.549 for the best features, i.e., 9% above chance at most. The statement that these are 'meaningful neurophysiological differences' is therefore too strong. The authors should report classification accuracy on all subjects, with a proper statistical test against chance (e.g., a permutation test at the group level), and treat the post-hoc exclusions only as a hypothesis-generating sensitivity analysis.","section":"§4.1.1, Tables 2 and 3"},{"comment":"The FAA interpretation is internally inconsistent. Equation (1) defines FAA = log(P_F4) − log(P_F3), and Section 3.7.3 correctly states that higher values indicate greater relative right-frontal activity (withdrawal-related) and lower/negative values indicate left-frontal dominance (approach-related). Section 4.3 then claims that the Exact condition's only positive mean value (Mean = 9.91e-14) indicates 'greater left frontal activity' associated with approach motivation. This directly contradicts the equation and the definition given in the Methods. Moreover, a mean of 9.91e-14 is effectively zero, so it cannot support any claim about hemispheric asymmetry. The authors must correct this misreading of the sign and, more importantly, report the actual effect sizes and confidence intervals for the FAA comparison.","section":"§3.7.3 vs. §4.3"},{"comment":"The between-subject classification results do not include a significance test against chance. Table 4 reports accuracies from five runs for each comparison, with run-to-run variation that is substantial (e.g., E vs. I: 0.568 ± 0.055; I vs. S: 0.495 ± 0.063), yet the text claims these results are 'comparable to or even higher than' individual models and that certain comparisons 'remain among the top performers.' With only five stochastic runs, a permutation test or a confidence interval against 0.5 is necessary before drawing conclusions. The data augmentation procedure, which multiplies epoch weights by uniform noise in [0.5, 1] to create synthetic ERPs, should also be validated more carefully; as described, it may artificially inflate agreement among the meta-model's training and test samples if the augmentation is not performed independently within each cross-validation fold.","section":"§3.7.2 and Table 4"}],"minor_comments":[{"comment":"The sentence 'The study used a within-subjects design with one independent variables' contains a grammatical error: 'variables' should be 'variable.'","section":"§3.1"},{"comment":"The figure presents a single-channel ERP plot (channel Oz) without error bars or the number of participants/trials used to compute the average. The caption 'average neural responses' is insufficiently precise for an EEG result figure.","section":"Figure 2"},{"comment":"The definition of 'Diff (%)' is not stated. The reader must infer that it is (accuracy − 0.5)/0.5; this should be stated in the table caption.","section":"Tables 2 and 3"},{"comment":"The text says 'even after excluding individual models (42%) with near-chance performance'; it is unclear whether 42% refers to subjects or models, and this percentage appears to refer only to the E vs. I comparison, not to all models. The wording should be clarified.","section":"§4.1.1"},{"comment":"Using an LLM (Llama 3) to shorten product descriptions may alter the semantic cues that drive category perception. The paper does not provide any validation that the shortened descriptions preserve the ESCI label distinctions; a small human validation study or example-based consistency check would strengthen the stimulus-construction section.","section":"§3.3.4"},{"comment":"The paper does not mention whether the data or analysis code will be made available. Given the small sample and the number of analysis choices (band-pass range, PCA variance threshold, MLP regularization, augmentation noise), a data/code availability statement would aid reproducibility.","section":"General"}],"recommendation":"reject","confidential_remarks":"For the editor: this is a borderline case between 'major revision' and 'reject.' The behavioral findings are credible and the topic is timely, but the paper's headline contribution is the neural differentiation of recommendation categories, and that claim is not supportable from the current design. The oddball confound is fundamental: every target category other than Exact is, by construction, an oddball relative to the context products. This cannot be fixed by reanalysis alone; it would require a new experiment with a balanced design. The post-hoc exclusion and FAA sign errors are fixable in revision, but they do not repair the confound. I therefore recommend rejection, although I would encourage the authors to resubmit a version focused on the behavioral results or with a redesigned neural experiment."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The behavioral results are real, but the neural story has a hole in it. The design always shows one to three Exact products before the target, so an Exact target is a repetition of the context and the other three categories are oddball deviations. The paper even describes the setup as an oddball paradigm. That means any EEG difference between Exact and the rest could be novelty detection, not the semantic category. The \"Exact vs Substitute\" classification, which the paper interprets as neural similarity, is exactly the pair most contaminated by this confound: the low accuracy could be category similarity cancelling against oddball-driven dissimilarity. I don't see how you can separate those without a control condition where category and repetition are unconfounded.\n\nWhat the paper does well: it is the first to bring EEG to ESCI's four recommendation categories, and the behavioral ratings are clean. The Friedman/Wilcoxon tests on relevance, purchase likelihood, and diversity are significant, and the medians tell a sensible story. The 71% self-label accuracy with confusion mostly between Exact and Substitute is a nice behavioral finding. The between-subject ERP analysis is a reasonable attempt, and the topomaps are useful.\n\nThe soft spots beyond the design: the post-hoc exclusion of non-significant subjects from Table 2 to Table 3 inflates accuracies without a proper statistical accounting, and the within-subject accuracies are barely above chance before exclusion (0.52-0.55). The FAA result is a simple error: a mean of 9.91e-14 is zero, not \"the highest and only positive value,\" and it cannot support a claim about left frontal approach motivation. The between-subject meta-model results are inconsistent, with several comparisons at chance.\n\nMy take: the behavioral contribution could stand alone if reframed, but the neural claims as presented are not supported. This needs a redesign or a substantial reinterpretation. I'd still send it to peer review because the research question is legitimate and the behavioral data appear solid, but I would expect the reviewers to demand major changes. If you read it, use it as a case study in how a well-run study can be undone by a stimulus confound.","headline":"Behavioral results are solid, but the neural-signature claim is undermined by the oddball design and a misread FAA value.","tokens_in":20437,"tokens_out":2598,"would_cite":false,"duration_ms":26050,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Different recommendation categories—Exact, Substitute, Complement, Irrelevant—are decodable from EEG signals and from user ratings, with Exact and Substitute the most similar.","keywords":["Recommender Systems","Electroencephalography","Recommendation categories","ESCI labels","User study","Behavioural analysis","E-commerce","Neural variability"],"falsifier":"Run the same task with context products drawn from all four categories in balanced proportions; if pairwise EEG classification of Exact versus Irrelevant (or Exact versus Substitute) collapses to chance when Exact is no longer the repeated baseline, then the reported neural signatures reflect stimulus novelty rather than recommendation semantics. A simpler check: compare classification accuracy between trials with k=1 and k=3 Exact context products, since an oddball effect would grow with the number of repetitions.","tokens_in":19385,"feed_emoji":"🧠","tokens_out":5757,"duration_ms":57807,"temperature":0.7,"pith_summary":"This paper tries to show that the category of a product recommendation—Exact, Substitute, Complement, or Irrelevant—leaves a measurable trace in brain activity and in users' own ratings, not just in the recommender's accuracy. Using EEG from nineteen participants who evaluated query-product pairs from an e-commerce corpus, it reports above-chance pairwise classification of the four categories from neural features, with Exact and Substitute the most confusable pair. Behaviourally, Exact and Irrelevant sit at opposite ends of relevance, purchase-likelihood, and diversity ratings, while Substitute and Complement occupy the middle and keep relevance without sacrificing diversity. If the claim holds, recommender systems could move beyond relevance scoring toward user-centred, category-aware, and eventually personalised optimisation.","feed_headline":"EEG distinguishes four recommendation categories","feed_subtitle":"Exact and substitute products look alike to the brain, while complements balance relevance and diversity.","key_machinery":"The argument runs on a within-subjects pairwise classification machinery: a support vector machine trained on one-second EEG windows after each recommendation, with features that include single-trial event-related potentials, power spectral density in delta, theta, alpha and low-beta bands, Kullback-Leibler divergence from mean responses, and signal complexity measures (approximate entropy, Higuchi and Katz fractal dimension, detrended fluctuation analysis), reduced by PCA and evaluated on repeated down-sampled train/test splits. A second layer averages epochs into synthetic ERPs and trains per-channel multilayer perceptrons whose probabilistic outputs feed a logistic-regression meta-classifier for between-subject decoding. The experimental scaffold is an oddball-like trial in which a query is followed by one to three Exact products and then the target recommendation, so Exact is a repetition of the context class and the other categories are deviations from it; this scaffold is what makes the category signals measurable, and also what leaves the novelty confound open.","core_discovery":"In the paper's own terms, the discovery is that recommendation semantics are decodable from the user's brain: pairwise SVM classifiers operating on single-trial EEG features (ERPs, band power, KL divergence, complexity measures) distinguish Exact from Irrelevant recommendations at around 0.545 accuracy before subject exclusions and up to 0.599 afterward, while Exact versus Substitute is consistently the hardest pair (0.524 before exclusions, about 0.567 afterward). Between-subject ERP meta-classifiers reproduce the ordering, with Complement versus Substitute unexpectedly strongest in that setting. Behavioural ratings show significant category effects on relevance, likelihood of purchase, and diversity, and engagement markers (frontal alpha asymmetry, theta and beta power) peak for Exact items. Together the results are interpreted as evidence that the brain treats exact matches and functional substitutes in a closely related way, separates relevant from irrelevant recommendations, and exhibits large inter-subject variability that speaks against one-size-fits-all recommendation strategies.","pith_inferences":["A control experiment that varies the context type (not always Exact) is needed to separate semantic category identity from oddball novelty; the current design cannot rule out that some of the EEG separation is a response to any non-Exact item.","If the neural decodability survives that control, EEG-derived labels could be used to improve or re-rank recommendations in categories like Complement, where behavioural ratings alone are ambiguous.","The behavioural overlap of Substitute and Complement clusters suggests a testable extension: users may accept a related-items shelf mixing substitutes and complements, provided per-user diversity thresholds are tuned.","The between-subject strength of Complement-versus-Substitute decoding hints that semantic association processing is a separate neural axis from relevance; this could be probed with queries that deliberately vary association strength."],"forward_implications":["Recommender systems could use neural signatures of the Exact–Irrelevant distinction as an implicit relevance signal, reducing reliance on clicks and ratings.","Because Exact and Substitute produce similar neural and behavioural responses, systems may safely rank substitutes close to exact matches without confusing users.","Complement recommendations occupy a middle ground on relevance and diversity, so blending them into result lists can broaden choice without sacrificing perceived quality.","Engagement markers (frontal alpha asymmetry, theta and beta power) favour Exact items, giving a physiological target for evaluating recommendation quality.","Large inter-subject variability in decoding accuracy implies that category-aware ranking should be personalised rather than globally optimised."],"supporting_citations":[{"why":"Supplies the within-subjects classification-and-significance procedure used to test whether pairwise EEG differences exceed chance.","marker":"[17]"},{"why":"Provides the MNE-Python library used for filtering, epoching, and cleaning the EEG data.","marker":"[20]"},{"why":"Defines the ESCI recommendation categories (Exact, Substitute, Complement, Irrelevant) that structure the experiment.","marker":"[43]"},{"why":"Supplies the Amazon product corpus from which the stimuli were drawn.","marker":"[45]"},{"why":"Provides the FASTER artifact-rejection pipeline used to denoise the EEG epochs.","marker":"[46]"},{"why":"Establishes the prior EEG-based relevance decoding approach that this study extends to multiple recommendation categories.","marker":"[54]"},{"why":"Supplies the Shopping Queries dataset with its ESCI relevance judgments, the ground truth for category labels.","marker":"[56]"},{"why":"Provides the oddball paradigm that shapes the trial structure and is also the source of the novelty confound.","marker":"[62]"}],"fun_headline_variants":["EEG decodes four product recommendation types","Neural signals distinguish exact from irrelevant picks","Brain waves reveal how users value recommendations","Recommendation categories leave distinct brain traces","EEG shows substitutes look like exact matches to brain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every trial shows the user one to three Exact products before the target, so an Exact target repeats the pattern while Substitute, Complement, and Irrelevant targets break it; the paper's central claim assumes that category identity, not that breaking-the-pattern effect, produces the neural differences.","fun_headline_variants_meta":{"raw":{"variants":["EEG decodes four product recommendation types","Neural signals distinguish exact from irrelevant picks","Brain waves reveal how users value recommendations","Recommendation categories leave distinct brain traces","EEG shows substitutes look like exact matches to brain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000728,"raw_usage":{"total_tokens":3228,"prompt_tokens":881,"completion_tokens":2347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":2281}},"tokens_in":497,"tokens_out":2347,"duration_ms":17224,"temperature":1.0,"reasoning_tokens":2281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:02:09.748114+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same task with context products drawn from all four categories in balanced proportions; if pairwise EEG classification of Exact versus Irrelevant (or Exact versus Substitute) collapses to chance when Exact is no longer the repeated baseline, then the reported neural signatures reflect stimulus novelty rather than recommendation semantics. A simpler check: compare classification accuracy between trials with k=1 and k=3 Exact context products, since an oddball effect would grow with the number of repetitions.","supporting_citations":[{"cited_title":"McGeown, and Yashar Moshfeghi","cited_arxiv_id":null,"evidence_quote":"Establishes the prior EEG-based relevance decoding approach that this study extends to multiple recommendation categories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the oddball paradigm that shapes the trial structure and is also the source of the novelty confound."}],"review_version":2}