{"id":"08ffb565-5365-4130-9a69-1c4752e571bb","arxiv_id":"2505.12279","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive survey that organizes side-information-driven session-based recommendation by data type, datasets, encoding, injection, and techniques.","lead":"This paper surveys roughly 60 research papers on session-based recommendation systems that use extra data such as price, text, images, time, and behavior type. It organizes the field by data type, lists available datasets, and proposes a data-centric taxonomy for future work.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Literature collection protocol is too underspecified to support the 'first comprehensive survey' claim; an independent keyword-expanded search is needed to verify completeness.","rationale":"The paper delivers a clearly organized, genuinely useful map of side-information-driven SBR: the task formulation (Section II), dataset inventory (Table I), data utility discussion (Section IV), and the taxonomy of methods by side-information type (Table II) are all coherent and will help practitioners. The central claim, however, is a novelty/comprehensiveness claim: 'the first effort' to survey this topic, covering 'over 60 papers'. That claim depends entirely on the completeness of the search protocol. The protocol described in Section I is too underspecified to be falsifiable: only three keyword phrases, a subjective venue whitelist, and 'selected' arXiv preprints. Because the field uses many synonyms and adjacent terms, and because the paper itself includes SR models under a non-operationalized criterion, the boundaries of the reviewed set are not reproducible. This is the same load-bearing assumption the reader identified. The two textual defects (the '[?]' citation in Section III and the 'GNN'/'CNN' slip in Section V-C2) are not load-bearing by themselves, but they strengthen the case that the final vetting pass was incomplete. The proposed concrete test—an independent re-run of the search with expanded keywords—would settle whether significant work was missed. If it finds few or no missing papers, the conditional concerns are resolved; if it finds many, the comprehensiveness claim needs to be softened. Since the reader already issued CONDITIONAL, my assessment leaves that verdict unchanged.","tokens_in":32608,"tokens_out":5874,"duration_ms":56169,"concrete_test":"Independently re-run the Section I search protocol on DBLP and Google Scholar for 2016–2025, then run a second search adding adjacent terms: 'feature-rich session-based recommendation', 'context-aware session-based recommendation', 'attribute-aware session-based recommendation', 'multi-modal session-based recommendation', 'side information session-based recommendation'. Compare the union of hits against the papers in Table II and the reference list. Count unique papers that satisfy the survey's inclusion criteria (use at least one side-information type in a session-based setting, or an SR model trainable under SBR splits) but are absent from the survey. If the count exceeds 10% of the roughly 60 covered papers, the comprehensiveness claim is materially weakened; if it is near zero, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that this is the first comprehensive data-centric survey of side-information-driven session-based recommendation (SIDSBR)—rests on the completeness of the literature collection protocol (Section I, 'Paper collection'). That protocol is described only qualitatively: DBLP and Google Scholar, three keyword phrases ('session-based recommendation', 'session recommendation', 'sequential recommendation'), post-2016 only, a curated venue list, and subjectively 'selected' arXiv preprints. It cannot establish comprehensiveness because relevant work in this area often appears under adjacent terms the protocol does not list: 'feature-rich session-based recommendation', 'context-aware/attribute-aware session-based recommendation', 'multi-modal session-based recommendation', and 'side-information-enhanced sequential recommendation'. Moreover, the stated criterion for including SR models—'effective when trained using the aforementioned SBR experimental implementations' (Section II-B)—is not operationalized, so the boundary of the reviewed set is not reproducible. If a nontrivial body of work using those terms, or appearing in venues outside the whitelist, is absent from Table II and the reference list, both the 'first' and 'comprehensive' claims are weakened. Two textual defects reinforce this concern: the unresolved citation placeholder '[?]' in the Instacart paragraph (Section III) and the CNN subsection mislabeling its formula as 'processing of GNN' (Section V-C2). These are symptomatic of an incomplete final vetting pass. The survey's primary contribution is organizational, so a reference map with gaps is a materially weaker contribution than claimed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys side-information-driven session-based recommendation (SIDSBR) from a data-centric perspective. It formulates the task, distinguishes SBR from sequential recommendation, reviews benchmark datasets and their available side information in Table I, analyzes the characteristics and utility of ten types of side information (time, category, brand, price, text, image, address, rating, review, and behavior), discusses data encoding and three injection manners (item-level, session-level, and prompt-level), summarizes neural techniques (RNN, CNN, attention, GNN, contrastive learning, and LLMs), and organizes research progress by side-information type in Table II. The authors claim that this is the first comprehensive survey of SIDSBR and identify under-explored data types, missing benchmark resources, and future directions such as joint multi-information incorporation, cold-start SBR, explainable SBR, and LLM-based SBR.","tokens_in":32950,"tokens_out":6864,"duration_ms":60988,"significance":"If the survey's coverage is indeed representative, it provides a valuable reference map for a fast-growing area. The data-centric organization, the explicit dataset-information matrix (Table I), and the side-information-type taxonomy (Table II) are useful resources. The paper also makes concrete, falsifiable observations, e.g., that brand, price, and review information are under-explored in SBR, and that no existing dataset covers all listed side-information types. The survey includes machine-checkable elements: the tables can be cross-verified against the reference list, and the procedural description of the literature search can be audited. However, the core contribution—the \"first comprehensive survey\" claim—depends on the completeness of the search protocol, which is currently underspecified.","major_comments":[{"comment":"The literature search protocol is too underspecified to support the \"first comprehensive survey\" claim. The protocol mentions only three keyword phrases (\"session-based recommendation\", \"session recommendation\", and \"sequential recommendation\"), a post-2016 cutoff, a whitelist of venues, and \"selected\" arXiv preprints. It does not include adjacent terminologies commonly used in this area, such as \"feature-rich session-based recommendation\", \"attribute-aware session-based recommendation\", \"multi-modal session-based recommendation\", or \"side-information-enhanced sequential recommendation\". As a result, the search may miss relevant work, and the completeness claim is not reproducible. Please operationalize the protocol: list exact query strings, databases, screening criteria, and the number of papers retrieved and excluded, or soften the \"first/comprehensive\" claim accordingly.","section":"Section I, Paper collection"},{"comment":"The criterion for including SR models in a survey titled \"session-based recommendation\" is stated as \"the selected SR models should be effective when trained using the aforementioned SBR experimental implementations\" (Section II-B), but this criterion is not operationalized, and it is not evident how Table II was populated from it. For instance, entries such as S3-Rec [68], SASRec [30], and BERT4Rec [31] are originally proposed for sequential recommendation with user profiles, and the text does not clarify under what conditions they qualify as SBR methods. This makes the boundary of the reviewed set ambiguous and further weakens the reproducibility of the survey's coverage. Please make the inclusion/exclusion decision for each category of method explicit, or provide a supplementary list of the included SR-derived methods and the justification for each.","section":"Section II-B and Table II"}],"minor_comments":[{"comment":"The sentence \"Generally, We can formulate the processing of GNN as follows\" should refer to CNN, not GNN; the equation is a convolution over item embeddings, not a graph neural network update.","section":"Section V-C2, Eq. (2)"},{"comment":"The unresolved citation placeholder \"[?]\" in the sentence about sampling subsets must be replaced with the intended reference(s); an unresolved placeholder in a survey undermines the traceability of the claim.","section":"Section III, Instacart paragraph"},{"comment":"The model name \"Reformer [160]\" should be \"Recformer [160]\" to match the reference title (\"Text is all you need: Learning language representations for sequential recommendation\") and the entry in Table II.","section":"Section VI-E"},{"comment":"The contrastive loss in Eq. (5) is missing the exponential function and the logarithm; as written, it is not the standard InfoNCE objective. Please correct the formula or add a note that it is a simplified variant.","section":"Section V-C5, Eq. (5)"},{"comment":"\"In addtion\" should be \"In addition\".","section":"Section III, Tmall paragraph"},{"comment":"The phrase \"the first proposal and wide acceptance of SBR in GRU4Rec [3]\" is imprecise; GRU4Rec is one of the first deep-learning SBR models, not the first proposal of session-based recommendation in general. Consider rewording to avoid a historically inaccurate statement.","section":"Section I"},{"comment":"The text states that \"over 60 papers\" are reviewed, but Table II lists approximately 53 approaches; please clarify how the 60 count is derived, e.g., whether it includes foundational SBR/SR papers cited outside Table II or papers discussed in the text but not tabled.","section":"Section I and Table II"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of TKDE and, after strengthening the literature-collection methodology and fixing the listed defects, would be a useful reference. I did not find evidence of a prior dedicated survey on SIDSBR, but the authors should re-verify this claim carefully; if a prior survey exists (e.g., in non-English venues or as an arXiv preprint), the novelty claim must be revised. The self-citations (MMSBR, CoHHN, BiPNet, DIMO, FineRec) are to legitimate, peer-reviewed papers and are not a concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you the short version: this is a solid, useful survey, not a breakthrough. The value is organizational — a data-centric taxonomy of side-information-driven session-based recommendation, a dataset availability table, and a methods summary table covering ~60 papers. If you work near this area, you'll want it on hand.\n\nWhat's genuinely new: the framing around data characteristics and utility, the grouping of side info into numerical vs descriptive, and the explicit comparison table of datasets with which info types they carry. They also do a decent job separating SBR from sequential recommendation in a way that's often muddled in the literature, and they include recent LLM-based methods. The open-directions section (cold-start, explainable SBR, missing benchmarks) is reasonable and not just boilerplate.\n\nWhere I'd push back: the literature collection protocol is described in four sentences and isn't reproducible. DBLP and Google Scholar with three keywords, post-2016, top venues plus selected preprints — that's not enough to support 'first comprehensive survey.' I don't think the paper misses a huge body of work, but the claim is stronger than the protocol. If I were refereeing, I'd ask for a broader keyword set (e.g., 'feature-rich', 'attribute-aware', 'multi-modal session-based recommendation'), a PRISMA-style flow or at least the search string, and a count of papers retrieved vs included. That's a moderate fix, not a fatal one.\n\nThe structural content holds up. Table II is a useful snapshot, and the discussion of injection levels (item, session, prompt) is a genuinely useful lens. Minor defects: the Instacart entry leaves an unresolved '[?]' citation, and Section V-C2 labels the CNN formula as 'processing of GNN.' Copyediting issues, but symptomatic of a slightly rushed final pass.\n\nSelf-citations are fine here — those are independent published methods being reviewed, and the authors don't privilege them.\n\nBottom line: this paper deserves peer review and will be useful to the community. The soft spots are addressable in a revision. I'd engage with it.","headline":"Useful data-centric survey of side-information-driven session-based recommendation; the taxonomy and dataset tables are the value, and the 'first comprehensive' claim needs a more rigorous literature protocol.","tokens_in":33395,"tokens_out":4155,"would_cite":true,"duration_ms":36107,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first data-centric review of side-information-driven session-based recommendation, sorting more than sixty methods by the type of data they exploit and identifying which data types remain unused.","keywords":["session-based recommendation","side information","data-centric survey","benchmark analysis","multi-modal recommendation","sequential recommendation","large language models","recommender systems"],"falsifier":"Checking whether any dataset already in routine use covers all ten side-information types in Table I, especially both behavior and image, would settle the survey's claim that no existing benchmark supports holistic comparison of methods across information types.","tokens_in":32387,"feed_emoji":"🗺️","tokens_out":6784,"duration_ms":68294,"temperature":0.7,"pith_summary":"Session-based recommendation predicts an anonymous user's next action from a short stream of clicks, and it suffers from data scarcity. This survey claims that the emerging response—side-information-driven session-based recommendation (SIDSBR)—has matured enough to deserve its own map, and that the map should be drawn from the data outward: what side information exists, what each type reveals about intent, and how models encode and inject it. It presents itself as the first survey to take that data-centric view, organizing over sixty methods into a taxonomy by side-information type. A reader comes away with concrete gaps: brand, price, and review information are underexplored, no benchmark covers all information types, and side information is proposed as the main route to cold-start, explainable, and LLM-based session recommendation.","feed_headline":"First data-centric survey maps side-information session recommenders","feed_subtitle":"Sixty-plus methods organized by the data they use, from price and image to behavior type.","key_machinery":"The machinery is the survey's double taxonomy: side information organized by type (time, category, brand, price, text, image, address, rating, review, behavior) and injection organized by level (item, session, prompt). The type taxonomy sorts research progress so a reader can find all methods that use, say, price, while the injection taxonomy sorts the technical how—look-up embeddings for numerical data, pretrained encoders for text and images, heterogeneous graphs or hypergraphs for fused item representations, and prompt templates for LLMs. These taxonomies carry the argument: they turn the general idea that side information helps into a structured claim about which data, injected where, reveals which user intent.","core_discovery":"The paper's central claim is that side-information-driven session-based recommendation is its own maturing topic, and that the most useful way to organize it is by the side information itself: time, category, brand, price, text, image, address, rating, review, and behavior type. It argues that conventional session-based recommendation using only item IDs can mine co-occurrence but misses user intent, while different data types reveal different intents—price signals budget sensitivity, images expose taste, and behavior type distinguishes casual browsing from purchase. On that basis the survey provides a task formulation that separates session-based recommendation from sequential recommendation, a catalogue of nineteen benchmarks and the side information each carries, an account of how each information type is encoded and injected at item, session, or prompt level, and a taxonomy of research progress by information type. It then concludes that brand, price, and reviews are under-exploited, that no existing dataset covers all the information types needed for holistic evaluation, and that side information is the main lever for cold-start, explainable, and LLM-based session-based recommendation.","pith_inferences":["A testable consequence of the data-centric claim is that a model's advantage from side information should grow as sessions get shorter; one could degrade sessions on the surveyed benchmarks to one or two clicks and measure whether side-information methods lose less accuracy than their ID-only counterparts.","The paper's examples of information conflict, such as image and text disagreement, suggest that dataset-level consistency statistics could become a benchmark quality metric, a step the survey itself does not propose.","The injection-level taxonomy could double as a design guide: item-level injection suits fused item representations, session-level injection suits per-modality preferences, and prompt-level injection suits text-only LLM pipelines, so a practitioner could choose a method family by which side information their production data actually contains."],"forward_implications":["A researcher who wants to use a specific side-information type, such as price or behavior type, can now locate the relevant methods and datasets in one place, which should make replication and comparison cheaper.","Underexplored data types—brand, price, and reviews—and the absence of a benchmark combining behavior with image information become concrete research targets rather than vague impressions.","Side information is framed as the main route to cold-start session recommendation, because new items can be linked to past behavior through shared features such as category or cast even without co-occurrence data.","LLM-based session recommenders are not yet competitive with ID-based models according to the surveyed evidence, and item text and images are identified as the leverage point for closing that gap.","The explicit distinction between session-based and sequential recommendation clarifies evaluation practice: SBR groups sessions for train and test splits and never uses user profiles, so sequential-recommendation methods need modification before being applied to SBR."],"supporting_citations":[{"why":"Introduces the recurrent-neural-network session-based recommendation setting that anchors the survey's post-2016 collection window and defines what a session is.","marker":"[3]"},{"why":"Supplies the standard attentive session-based recommendation baseline that side-information methods extend.","marker":"[2]"},{"why":"Existing survey of session-based recommenders that this paper distinguishes itself from by covering side information.","marker":"[4]"},{"why":"Sequential-recommendation review used to delineate the difference between session-based recommendation and sequential recommendation.","marker":"[16]"},{"why":"Second sequential-recommendation review used to position this survey against prior work on sequence modeling.","marker":"[17]"},{"why":"Multimodal recommender survey whose reliance on user profiles is contrasted with the anonymous-session setting of SBR.","marker":"[18]"},{"why":"Earlier side-information survey in general recommendation that the paper extends to the session-based task.","marker":"[19]"},{"why":"Empirical analysis showing that existing SBR surveys focus on item-ID methods, motivating the data-centric gap this survey fills.","marker":"[20]"},{"why":"Price-driven session recommendation used to motivate price-level encoding and price-sensitivity modeling.","marker":"[12]"},{"why":"Multi-modal session recommendation used to illustrate image-text fusion and the conflict problem between modalities.","marker":"[7]"}],"fun_headline_variants":["Data-centric map of side-info session recommenders","Side info: key to session-based recommendation","Survey: how side data shapes session recommenders","Side-information taxonomy for session-based recs","Data-first look at session recommendation side info"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the literature search was broad enough: if a meaningful body of side-information-driven session-based recommendation work falls outside the chosen keywords, venues, or post-2016 window, the survey's claim to be the first comprehensive map of the topic would not hold.","fun_headline_variants_meta":{"raw":{"variants":["Data-centric map of side-info session recommenders","Side info: key to session-based recommendation","Survey: how side data shapes session recommenders","Side-information taxonomy for session-based recs","Data-first look at session recommendation side info"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1387,"prompt_tokens":931,"completion_tokens":456,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":387}},"tokens_in":547,"tokens_out":456,"duration_ms":5624,"temperature":1.0,"reasoning_tokens":387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:35:55.640283+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Checking whether any dataset already in routine use covers all ten side-information types in Table I, especially both behavior and image, would settle the survey's claim that no existing benchmark supports holistic comparison of methods across information types.","supporting_citations":[],"review_version":1}