REVIEW 3 major objections 5 minor 36 references
Balancing Accuracy and Novelty with Sub-Item Popularity
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read By scoring items through the sub-IDs they share with a user's listening history, the paper shows that music recommenders can gain substantial personalized novelty without sacrificing accuracy.
desk verdict Useful incremental idea—sPPS at sub-ID level—but the headline PPS-vs-sPPS comparison is confounded by training on both signals and ablating at inference; the claim needs a cleaner experiment before the 20% novelty gain is taken seriously. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is RecJPQ's sub-item decomposition, which represents each item by $m$ sub-IDs drawn from learned codebooks; the new signal sPPS sums, over the $m$ code positions, $\log(\text{count of that sub-ID in the user's history} + \epsilon)$, then z-score normalizes. It is combined with the recommender's logits and item-level PPS in a convex weighted score (Eq.~7), whose weights $\alpha$ and $\beta$ let one sweep the accuracy-novelty frontier at inference time without retraining. The sub-IDs are load-bearing: they must carry semantic grouping for the novelty gain to be relevant.
What would settle it
Run the same experiments with sub-IDs randomly permuted before computing sPPS (keeping everything else fixed); if sPPS still shows the same or larger novelty-accuracy frontier, the semantic-grouping premise is wrong. Alternatively, measure the genre or artist purity of the items sharing each sub-ID: if purity is near chance, the proposed mechanism is not operating.
Extended reading notes
Core claim
Modelling personalised popularity at the sub-ID level, not the item-ID level, captures latent structural recurrence in a user's listening history and yields a better accuracy-novelty trade-off. The final score is $\text{logits}_{\text{final}} = \gamma\,\text{logits}_{\text{rec}} + \alpha\,\text{PPS}_{\text{std}} + \beta\,\text{sPPS}_{\text{std}}$ with $\gamma = 1 - \alpha - \beta$, where PPS is the standardised log-frequency of whole items and sPPS is the standardised sum of log-frequencies of the item's constituent sub-IDs. When $\beta$ alone is active, sPPS reaches roughly 20% higher personalised novelty at the same NDCG; when both signals are active, item-level accuracy can be added with
Load-bearing premise
The whole gain rests on RecJPQ's learned sub-IDs actually grouping items by shared stylistic or structural traits relevant to the user; if the SVD-derived sub-IDs are semantically arbitrary, then the 'novel' items sPPS promotes would be unrelated to the user's tastes and the accuracy-novelty gain would disappear.
Editorial extensions
If this is right
- With sPPS, a music recommender can push novelty at a fixed accuracy level: at NDCG@40 near 0.32 on Last.fm, novelty rises from about 10 (PPS) to about 12 (sPPS), a roughly 20% relative gain.
- At a personalised-novelty threshold of 12, NDCG@40 improves from 0.2749 (PPS) to 0.3016 (sPPS) on Last.fm and from 0.1150 to 0.1242 on Yandex, so the gain is not limited to one dataset.
- Because $\alpha$ and $\beta$ are varied only at inference time, the same trained model can serve different accuracy-novelty operating points without retraining.
- Adding item-level PPS on top of a strong sub-ID signal ($\beta=0.9$, varying $\alpha$) improves NDCG while losing less novelty than PPS alone, indicating the two signals are complementary.
- The integration operates directly on the output scoring function, so the authors expect the approach to transfer to other sub-ID-based sequential recommenders.
Reading between the lines
- If sub-ID semantic grouping transfers, sPPS could apply to any catalogue with compositional items—outfits, news topics, code snippets—where sub-item recurrence encodes shared attributes; that extension goes beyond the paper's music experiments.
- A direct test of the mechanism would be to permute the sub-ID assignments before computing sPPS; the novelty-accuracy advantage should vanish under permutation if semantic grouping is the cause.
- The novelty metric used here measures deviation from the user's own listening history, so it does not by itself prove long-term engagement; an online or user-study check would be needed to confirm the added novelty is desirable.
- Per-user tuning of $\alpha$ and $\beta$ (for instance, higher $\beta$ for users with long, repetitive histories) is a natural extension implied by the convex score, but the paper only sweeps global values.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes sPPS, a personalized popularity signal computed at the sub-ID level using RecJPQ's item decomposition, and combines it with item-level PPS and the recommender's logits via a weighted score function (Eq. (7)). Experiments on Yandex and Last.fm-1K with BERT4Rec+RecJPQ report that sPPS improves the accuracy–personalized-novelty trade-off compared with item-level PPS, citing a 20% relative novelty gain at matched NDCG and about 8–10% NDCG gains at novelty threshold 12. The code and data are public.
Significance. If the empirical claim were solid, this would be a modest but useful contribution: repurposing an efficiency-oriented sub-item representation for a beyond-accuracy objective, with a simple inference-time control. The paper is clearly written, the method is easy to implement, and the public code/data support reproducibility. The scientific value, however, depends almost entirely on the reliability of the two-dataset comparison, because the method is an inference-time re-ranking heuristic rather than a new learning framework.
major comments (3)
- [§3 Experimental Setup and §4 RQ2] The central PPS-only vs sPPS-only comparison is confounded by the training-time configuration. The paper states that α=0.4 and β=0.4 are fixed at training time and only varied at inference. Therefore the red and blue frontiers in Figure 2 are obtained by ablating one additive signal from a single checkpoint trained with both signals. The sub-embeddings and logits_rec in that checkpoint were optimized under Eq. (7) with both PPS and sPPS present; removing one signal at inference tests a scoring function the model never trained with. The 20% relative novelty gain attributed to sPPS could be an artifact of training with sPPS in the loss, not an intrinsic property of sub-ID granularity. A fair comparison requires separate checkpoints trained with only PPS, only sPPS, and both, and then reporting inference-time frontiers for each.
- [Table 2] The NDCG values at novelty thresholds (≥0, ≥10, ≥12, ≥14) are not matched-accuracy comparisons if, as the column headings suggest, the metric is computed only over users/lists whose personalized novelty reaches the threshold. PPS and sPPS will generally have different qualifying user subsets at each threshold, so the +9.7% and +8.0% gains at threshold 12 may reflect population differences rather than recommendation quality. The paper should report the number of qualifying users per method and threshold, or use a threshold-free frontier comparison on the same user set.
- [§4 and Figure 2] All claims of 'significantly higher novelty' rest on a single training run per dataset, with no error bars, confidence intervals, or significance tests. The margins are small (e.g., 0.1242 vs 0.1150 on Yandex), and the comparisons in Figure 2 are read off interpolated curves. At least three seeds with standard deviations, or paired tests across users, are needed to support the central accuracy–novelty claim and the word 'significantly' in the abstract.
minor comments (5)
- [§2.4, Eq. (7)] The text says γ=1−α−β 'ensures a convex combination.' This is only true under the constraints α,β≥0 and α+β≤1. The green-curve sweep described in §4 ('fix β=0.9 and gradually increase α') can violate the latter if α>0.1. State the feasible region explicitly or use unconstrained weights with a different normalization.
- [Figure 2] The legend and labels are hard to read, and the caption text ('β=0 (PPS)', 'β=0 (Sub-ID)', 'β=0.9 (varying)') is inconsistent with the body text describing the green curve. Clarify which parameter is on the x-axis of each sweep and what the marker values are.
- [§4 RQ2] The example 'at fixed NDCG@40 approximately 0.32, PPS yields novelty about 10 whereas sPPS achieves roughly 12' is read off a plot. Provide the underlying data points or a table of (α, β, NDCG, novelty) values to make the 20% claim auditable.
- [§2.3] The semantic-coherence assumption—that RecJPQ's SVD-based sub-IDs group items by genre, artist, or similar latent attributes—is central to the method's interpretation but is not validated. An analysis of sub-ID–attribute association would strengthen the causal story behind the novelty gains.
- [Abstract and §4] The word 'significantly' is used without statistical support. Please soften to 'higher' or provide significance tests.
Circularity Check
No circular derivation: sPPS is an independent empirical signal; the central claim rests on experiments, not on re-using its inputs.
full rationale
The paper's central claim—that sub-ID-level personalised popularity (sPPS) yields a better accuracy–novelty trade-off than item-level PPS—is an empirical comparison, not a derivation from an input. sPPS (Eq. 6) is a new function of user history counts and RecJPQ's sub-ID codes, and the final score (Eq. 7) is a convex combination that is swept at inference. No fitted parameter is renamed as a prediction, and no equation defines its output in terms of the target metric. The building blocks PPS [1] and RecJPQ [15] are prior works by overlapping authors, but the paper implements and evaluates them on standard datasets; the comparison is external to those papers' conclusions. The assumption that RecJPQ sub-IDs are semantically coherent is a correctness risk (and the training-time α=β=0.4 vs inference-time sweep is a potential confound), but these are experimental validity concerns, not circular reductions. The reviewer's guide requires a specific reduction (Eq. X = Eq. Y by construction) to flag circularity; none is present here.
Assumptions & free parameters
free parameters (5)
- alpha =
0.4 at training; swept 0 to 0.9 at inference
- beta =
0.4 at training; swept 0 to 0.9 at inference
- smoothing epsilon =
not reported
- embedding dimension and split count =
d=256, m=32
- novelty thresholds =
10, 12, 14
assumptions (3)
- domain assumption RecJPQ's SVD-quantized sub-IDs group items by latent stylistic or structural similarity.
- domain assumption Additive combination of the Transformer logits with popularity scores in Eq. (7) preserves ranking semantics; weights are convex only if alpha+beta <= 1.
- domain assumption Evaluation protocol from [1] (Global Temporal Split, NDCG with relevance mapping, negative labels as zero) is a valid measure of accuracy.
Cite this review
Pith. "Pith review of Balancing Accuracy and Novelty with Sub-Item Popularity." pith.science (2026). https://pith.science/paper/DFB3YYKI
@misc{pith2026250805198,
author = {Pith},
title = {Pith review of: Balancing Accuracy and Novelty with Sub-Item Popularity},
year = {2026},
howpublished = {\url{https://pith.science/paper/DFB3YYKI}},
note = {Machine review of arXiv:2508.05198}
}
read the original abstract
In the realm of music recommendation, sequential recommenders have shown promise in capturing the dynamic nature of music consumption. A key characteristic of this domain is repetitive listening, where users frequently replay familiar tracks. To capture these repetition patterns, recent research has introduced Personalised Popularity Scores (PPS), which quantify user-specific preferences based on historical frequency. While PPS enhances relevance in recommendation, it often reinforces already-known content, limiting the system's ability to surface novel or serendipitous items - key elements for fostering long-term user engagement and satisfaction. To address this limitation, we build upon RecJPQ, a Transformer-based framework initially developed to improve scalability in large-item catalogues through sub-item decomposition. We repurpose RecJPQ's sub-item architecture to model personalised popularity at a finer granularity. This allows us to capture shared repetition patterns across sub-embeddings - latent structures not accessible through item-level popularity alone. We propose a novel integration of sub-ID-level personalised popularity within the RecJPQ framework, enabling explicit control over the trade-off between accuracy and personalised novelty. Our sub-ID-level PPS method (sPPS) consistently outperforms item-level PPS by achieving significantly higher personalised novelty without compromising recommendation accuracy. Code and experiments are publicly available at https://github.com/sisinflab/Sub-id-Popularity.
Figures
Reference graph
Works this paper leans on
-
[1]
Davide Abbattista, Vito Walter Anelli, Tommaso Di Noia, Craig MacDonald, and Aleksandr Vladimirovich Petrov. 2024. Enhancing Sequential Music Recommen- dation with Personalized Popularity Awareness. In RecSys. ACM, 1168–1173
work page 2024
-
[2]
Vito Walter Anelli, Tommaso Di Noia, Eugenio Di Sciascio, Azzurra Ragone, and Joseph Trotta. 2019. Local Popularity and Time in top-N Recommendation. In ECIR (1) (Lecture Notes in Computer Science, Vol. 11437) . Springer, 861–868
work page 2019
-
[3]
Òscar Celma. 2010. Music Recommendation and Discovery - The Long Tail, Long Fail, and Long Play in the Digital Music Space . Springer
work page 2010
-
[4]
Junsu Cho, Dongmin Hyun, SeongKu Kang, and Hwanjo Yu. 2021. Learning Het- erogeneous Temporal Patterns of User Preference for Timely Recommendation. In WWW ’21: The Web Conference 2021, Virtual Event / Ljubljana, Slovenia, April 19-23, 2021, Jure Leskovec, Marko Grobelnik, Marc Najork, Jie Tang, and Leila Zia (Eds.). ACM / IW3C2, 1274–1283. doi:10.1145/34...
arXiv 2021
-
[5]
Szu-Yu Chou, Yi-Hsuan Yang, and Yu-Ching Lin. 2015. Evaluating music recom- mendation in a real-world setting: On data splitting and evaluation metrics. In ICME. IEEE Computer Society, 1–6
work page 2015
-
[6]
Mouzhi Ge, Carla Delgado-Battenfeld, and Dietmar Jannach. 2010. Beyond accuracy: evaluating recommender systems by coverage and serendipity. In Proceedings of the Fourth ACM Conference on Recommender Systems (Barcelona, Spain) (RecSys ’10). Association for Computing Machinery, New York, NY, USA, 257–260. doi:10.1145/1864708.1864761
arXiv 2010
-
[7]
Lukas Gienapp, Maik Fröbe, Matthias Hagen, and Martin Potthast. 2020. The Impact of Negative Relevance Judgments on NDCG. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland) (CIKM ’20). Association for Computing Machinery, New York, NY, USA, 2037–2040. doi:10.1145/3340531.3412123
-
[8]
Robert M. Gray. 1984. Vector quantization. IEEE ASSP Magazine 1 (1984), 4–29. https://api.semanticscholar.org/CorpusID:116066292
work page 1984
Show all 36 references
-
[9]
Balázs Hidasi and Ádám Tibor Czapp. 2023. Widespread Flaws in Offline Evalu- ation of Recommender Systems. In Proceedings of the 17th ACM Conference on Recommender Systems (Singapore, Singapore) (RecSys ’23). Association for Com- puting Machinery, New York, NY, USA, 848–855. d...
2023
-
[10]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. 20, 4 (2002), 422–446
2002
-
[11]
Marius Kaminskas and Derek Bridge. 2016. Diversity, Serendipity, Novelty, and Coverage: A Survey and Empirical Analysis of Beyond-Accuracy Objectives in Recommender Systems. ACM Trans. Interact. Intell. Syst. 7, 1, Article 2 (Dec. 2016), 42 pages. doi:10.1145/2926720
2016 doi
-
[12]
Denis Kotkov, Jari Veijalainen, and Shuaiqiang Wang. 2016. Challenges of Serendipity in Recommender Systems. 251–256. doi:10.5220/0005879802510256
2016 doi
-
[13]
Ashishkumar Patel, Dharmesh Tank, and Pratik Modi. 2023. A Survey on Finding Novelty in the Recommender Systems by Adding Lower Similar Items in the Top-N List. In International Journal for Scientific Research and Development . India, 153–158
2023
-
[14]
Aleksandr Vladimirovich Petrov and Craig MacDonald. 2023. gSASRec: Reducing Overconfidence in Sequential Recommendation Trained with Negative Sampling. In RecSys. ACM, 116–128
2023
-
[15]
Petrov and Craig Macdonald
Aleksandr V. Petrov and Craig Macdonald. 2024. RecJPQ: Training Large- Catalogue Sequential Recommenders. In WSDM. ACM, 538–547
2024
-
[16]
Aleksandr Vladimirovich Petrov, Craig Macdonald, and Nicola Tonellotto. 2024. Efficient Inference of Sub-Item Id-based Sequential Recommendation Models with Millions of Items. In RecSys. ACM, 912–917
2024
-
[17]
Petrov, Craig Macdonald, and Nicola Tonellotto
Aleksandr V. Petrov, Craig Macdonald, and Nicola Tonellotto. 2025. Efficient Recommendation with Millions of Items by Dynamic Pruning of Sub-Item Em- beddings. In SIGIR. ACM, 2102–2111
2025
-
[18]
Petrov, Efi Karra Taniskidou, and Sean Murphy
Aleksandr V. Petrov, Efi Karra Taniskidou, and Sean Murphy. 2025. CountNet: Utilising Repetition Counts in Sequential Recommendation. In ECIR (3) (Lecture Notes in Computer Science, Vol. 15574) . Springer, 52–66
2025
-
[19]
Massimo Quadrana, Paolo Cremonesi, and Dietmar Jannach. 2018. Sequence- Aware Recommender Systems. ACM Comput. Surv. 51, 4 (2018), 66:1–66:36
2018
-
[20]
Tran, Jonah Samost, Maciej Kula, Ed H
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Samost, Maciej Kula, Ed H. Chi, and Mahesh Sathiamoorthy. 2023. Recommender Systems with Generative Retrieval. In NeurIPS
2023
-
[21]
Pengjie Ren, Zhumin Chen, Jing Li, Zhaochun Ren, Jun Ma, and Maarten de Rijke. 2019. RepeatNet: a repeat aware neural recommendation machine for session-based recommendation (AAAI’19/IAAI’19/EAAI’19). AAAI Press, Article 590, 8 pages. doi:10.1609/aaai.v33i01.33014806
2019 doi
-
[22]
Keshavan, Maheswaran Sathiamoorthy, Yilin Zheng, Lichan Hong, Lukasz Heldt, Li Wei, Devansh Tandon, Ed H
Anima Singh, Trung Vu, Nikhil Mehta, Raghunandan H. Keshavan, Maheswaran Sathiamoorthy, Yilin Zheng, Lichan Hong, Lukasz Heldt, Li Wei, Devansh Tandon, Ed H. Chi, and Xinyang Yi. 2024. Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations. In Rec...
2024
-
[23]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[25]
Viet-Anh Tran, Guillaume Salha-Galvan, Bruno Sguerra, and Romain Hennequin
-
[26]
Saúl Vargas and Pablo Castells. 2011. Rank and relevance in novelty and diversity metrics for recommender systems. In Proceedings of the Fifth ACM Conference on Recommender Systems (Chicago, Illinois, USA) (RecSys ’11). Association for Com- puting Machinery, New York, NY, USA,...
2011
-
[27]
Sheng, and Mehmet A
Shoujin Wang, Liang Hu, Yan Wang, Longbing Cao, Quan Z. Sheng, and Mehmet A. Orgun. 2019. Sequential Recommender Systems: Challenges, Progress and Prospects. In IJCAI. ijcai.org, 6332–6338
2019
-
[28]
Hiromu Yakura, Tomoyasu Nakano, and Masataka Goto. 2018. FocusMusicRec- ommender: A System for Recommending Music to Listen to While Working. In IUI. ACM, 7–17
2018
-
[29]
Hiromu Yakura, Tomoyasu Nakano, and Masataka Goto. 2022. An automated system recommending background music to listen to while working. User Model. User Adapt. Interact. 32, 3 (2022), 355–388
2022
-
[30]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma
-
[31]
Yixin Zhang, Lizhen Cui, Wei He, Xudong Lu, and Shipeng Wang. 2021. Behav- ioral data assists decisions: exploring the mental representation of digital-self. Int. J. Crowd Sci. 5, 2 (2021), 185–203
2021
-
[32]
Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization. In CIKM. ACM, 1893–1902
2020
-
[33]
Reza Jafari Ziarani and Reza Ravanmehr. 2021. Serendipity in Recommender Systems: A Systematic Literature Review. Journal of Computer Science and Technology 36, 2 (2021), 375–396. doi:10.1007/s11390-020-0135-9
2021 doi
-
[2019]
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Rep- resentations from Transformer. In CIKM. ACM, 1441–1450
-
[2021]
arXiv:2108.00644 [cs.IR] https://arxiv.org/abs/2108.00644
Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance. arXiv:2108.00644 [cs.IR] https://arxiv.org/abs/2108.00644
-
[2023]
In SIGIR
Attention Mixtures for Time-Aware Sequential Recommendation. In SIGIR. ACM, 1821–1826
-
[2024]
In RecSys
Transformers Meet ACT-R: Repeat-Aware and Sequential Listening Session Recommendation. In RecSys. ACM, 486–496
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.