REVIEW 4 major objections 4 minor 1 cited by
A Multistakeholder Approach to Value-Driven Co-Design of Recommender System Evaluation Metrics in Digital Archives
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that stakeholder values in digital archives can be translated into a four-stage research-funnel evaluation framework, with eight proposed metric directions.
desk verdict A useful and clearly written qualitative contribution with a real gap in the funnel-claim reporting: 'naturally align' outruns the coding evidence, but the setup is worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the research funnel, a four-stage model of scholarly work in digital archives (discovery, interaction, integration, and impact), adapted from an established model of information-seeking behavior. It carries the argument by providing a shared structure onto which stakeholder values and proposed metric directions are mapped. The second mechanism is the focus group procedure itself: 25 experts in five stakeholder groups discussed scenario-based prompts, and an abductive coding process formalized their concerns into the funnel and the evaluation setup.
What would settle it
A direct check would be to re-analyze the 25 focus group transcripts with a published codebook and inter-rater reliability measures; if the four funnel stages cannot be reliably identified, or if a log-based study of archive users shows no sequential progression through discovery, interaction, integration, and impact, the evaluation setup loses its foundation.
Extended reading notes
Core claim
The paper's central claim is that the diverse, sometimes conflicting values of digital-archive stakeholders can be organized into a single evaluation framework built on the research funnel. Stakeholders described their work in phases; the authors synthesized these into four sequential stages — discovery, interaction, integration, and impact — and mapped value priorities and metric directions onto each stage. The resulting setup contains eight proposed metric directions: Research Path Quality and Collection Representation for discovery; Contextual Appropriateness and Control Effectiveness for interaction; Metadata-Weighted Relevance and Document Relationship Insight for integration; Research Integration and Cross-Stakeholder Value Alignment for impact. The paper presents these as directions rather than validated metrics, grounded in a qualitative analysis of 25 expert transcripts.
Load-bearing premise
The load-bearing premise is that the four-stage research funnel faithfully represents how scholars actually work in digital archives, since the funnel is the scaffold onto which every proposed metric direction is mapped.
Editorial extensions
If this is right
- Digital-archive recommender systems should be evaluated stage by stage along the research journey rather than by click-through or immediate engagement.
- Collection Representation and Metadata-Weighted Relevance give evaluators concrete handles on structural biases inherited from historical digitization priorities.
- Research Integration would extend evaluation beyond the session, tracking whether recommended items end up in citations, teaching materials, or curated collections.
- Cross-Stakeholder Value Alignment makes trade-offs explicit by weighting satisfaction across stakeholder groups, a design that could transfer to other domains with long-horizon value creation.
Reading between the lines
- The funnel could be turned into a testable measurement model: define stage transition probabilities from log data and check whether scholars who progress through all four stages report higher perceived value.
- A likely friction point the paper leaves implicit is that Collection Representation and Research Path Quality may pull against each other, since serendipitous coverage can reduce topical coherence; a composite objective would need to make that trade-off explicit.
- The same value-to-metric translation could be applied to longitudinal domains such as education or health, where value also accumulates over extended engagement rather than single consumption events.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multistakeholder, value-driven approach to designing evaluation metrics for recommender systems in digital archives. It reports on five focus groups (25 experts from upstream, provider, system, consumer, and downstream stakeholder groups) conducted around value priorities identified in the authors' prior work: visibility/representation, expertise adaptation, and transparency/trust. The central empirical claim is that stakeholder concerns naturally align with a four-stage research funnel (discovery, interaction, integration, impact), and the paper derives eight metric directions (e.g., Research Path Quality, Collection Representation, Contextual Appropriateness, Control Effectiveness, Metadata-Weighted Relevance, Document Relationship Insight, Research Integration, Cross-Stakeholder Value Alignment) organized by funnel stage. The paper presents these as proposals and explicitly calls for future implementation and user studies.
Significance. If the funnel-alignment claim holds, this would be a useful transferable framework for evaluating recommender systems in cultural heritage and other process-oriented domains, addressing a real gap in multistakeholder RecSys evaluation. The paper is transparent about its limitations, labels the metric directions as proposals, and provides publicly shared anonymized data, consent forms, and discussion guides, which are strengths for reproducibility. However, the load-bearing link between the focus-group data and the four-stage funnel is currently under-evidenced: the analysis is described as abductive, with the funnel 'structured' by the authors and attributed to Ellis's model, yet no coding evidence (codebook, inter-rater agreement, excerpt counts, or per-participant stage reference rates) is reported. Without such evidence, the evaluation setup may be an interpretive scaffold rather than an empirically grounded co-design outcome.
major comments (4)
- [§3 (Analysis and Development of the Evaluation Setup) and §4.2] The central claim that stakeholder concerns 'naturally align' with the four-stage research funnel is not supported by the reported analysis. Section 3 states that 'our synthesis structured these narratives into the four-part research funnel,' and Section 4.2 attributes the funnel to Ellis's model [19], yet no codebook, inter-rater agreement, per-stage excerpt counts, or analysis of how many of the 25 participants referenced each stage is provided. This is a load-bearing gap: either the funnel is an empirical finding grounded in coded transcripts, in which case coding evidence must be reported, or it is an interpretive scaffold, in which case the abstract's 'naturally align' is too strong and Table 2's metric directions should be framed as theory-driven proposals rather than data-driven results.
- [§4.2 and Table 2] The stage-specific evidence consists of roughly six short quotes, predominantly from system (S1, S3, S4) and consumer (C1) stakeholders, with no direct quotes supporting the Discovery-stage claims about upstream stakeholders' emphasis on serendipity or providers' awareness of structural biases, despite those claims being central to the funnel alignment. The manuscript should either provide representative quotes and counts for each stakeholder group and stage or temper the claim that all stakeholder concerns align with all four stages.
- [§4.3 (Impact stage) and §4.4] The translation from values to metrics is incomplete for at least two of the proposed directions. Research Integration requires longitudinal tracking of citations and research outputs, and Cross-Stakeholder Value Alignment requires weighting competing stakeholder priorities, but the paper does not specify how these weights would be derived from the focus-group value data. Since the paper's stated contribution is 'translating diverse stakeholder values into an evaluation metric setup,' a worked example or formalization of at least one metric (e.g., how Collection Representation would be computed from collection composition and recommendation distribution, or how stakeholder values would set the multi-objective weights) would make the translation operational rather than merely directional.
- [§3 (Participatory Focus Groups) and §4.1] The discussion guide was built on the three priority areas identified in the authors' prior paper [2] (visibility/representation, expertise adaptation, transparency/trust), and the abductive coding used 'pre-established value categories from literature.' This creates a risk of confirmatory bias: the resulting evaluation setup may re-inscribe the authors' prior framework rather than reflect emergent stakeholder values. The paper should include a reflexivity statement and, ideally, triangulate the funnel alignment with an independent coder or a member-checking procedure, to guard against this circularity.
minor comments (4)
- [§3 (Analysis and Development of the Evaluation Setup)] The citation for abductive coding is given as [31], but reference [31] is a paper on news recommender metrics (Vandenbroucke and Smets), not a methodological source for abductive coding; a proper methodological citation is needed.
- [§4.2] Reference [19] (Ellis's model) is introduced only in the results section; it should be cited earlier in Section 3, where the research funnel is first described as the organizing structure.
- [§Abstract and §4.2] The phrase 'naturally align' in the abstract and Section 4.2 overstates the degree of empirical grounding; suggest rewording to 'were organized into' or 'were structured according to' to match the methodology described in Section 3.
- [§3 (Participatory Focus Groups)] The sentence 'Discussion questions centered on three topics from previously identified priority areas [2]' omits the word 'the' before 'previously identified priority areas'; also, the repository link is given but the paper does not state whether verbatim transcripts or only anonymized summaries are shared, which would clarify reproducibility claims.
Circularity Check
No circular derivation: value categories and funnel are disclosed inputs; metric directions are proposed translations, not predictions forced by construction.
full rationale
The paper's central claim is a translation of stakeholder values into candidate metric directions, not a quantitative prediction or theorem derivation. The value dimensions (visibility/representation, expertise adaptation, transparency/trust) were indeed taken from the authors' prior work [2] and used to structure the focus-group discussion guide, and the research funnel is explicitly attributed to Ellis's model [19] and to the authors' own synthesis ('our synthesis structured these narratives into the four-part research funnel'). This is an interpretive scaffold, not a concealed circular step: the paper discloses both the deductive coding and the funnel's source, and the output is a set of proposed evaluation directions (e.g., Collection Representation, Research Path Quality) that are grounded in participant quotes and existing RecSys evaluation literature. There is no equation or fitted parameter whose value is reused as a prediction; the claims are qualitative, and the proposed metrics are explicitly flagged as requiring future validation. The absence of a codebook, inter-rater agreement, and per-stage excerpt counts weakens the empirical warrant for the 'natural alignment' claim, but that is an evidence and reporting limitation that falls under correctness risk rather than circularity. Self-citations [1, 2] are used to motivate the study and supply starting categories, but they are not used to forbid alternatives or to justify the metric directions by authority; the new focus-group data and external citations carry the argument. Hence no step reduces by definition to its own input.
Assumptions & free parameters
assumptions (3)
- domain assumption The four-stage research funnel (discovery, interaction, integration, impact) is an appropriate representation of scholarly workflows in digital archives, adapted from Ellis's model [19].
- domain assumption The stakeholder taxonomy (upstream, provider, system, consumer, downstream) transfers from prior RecSys multistakeholder research to the digital archives domain.
- domain assumption The three value dimensions used to design the focus group questions (visibility and representation, expertise adaptation, transparency and trust) cover the relevant value space for the domain.
Cite this review
Pith. "Pith review of A Multistakeholder Approach to Value-Driven Co-Design of Recommender System Evaluation Metrics in Digital Archives." pith.science (2026). https://pith.science/paper/S5GUVLFC
@misc{pith2026250703556,
author = {Pith},
title = {Pith review of: A Multistakeholder Approach to Value-Driven Co-Design of Recommender System Evaluation Metrics in Digital Archives},
year = {2026},
howpublished = {\url{https://pith.science/paper/S5GUVLFC}},
note = {Machine review of arXiv:2507.03556}
}
read the original abstract
This paper presents the first multistakeholder approach for translating diverse stakeholder values into an evaluation metric setup for Recommender Systems (RecSys) in digital archives. While commercial platforms mainly rely on engagement metrics, cultural heritage domains require frameworks that balance competing priorities among archivists, platform owners, researchers, and other stakeholders. To address this challenge, we conducted high-profile focus groups (5 groups x 5 persons) with upstream, provider, system, consumer, and downstream stakeholders, identifying value priorities across critical dimensions: visibility/representation, expertise adaptation, and transparency/trust. Our analysis shows that stakeholder concerns naturally align with four sequential research funnel stages: discovery, interaction, integration, and impact. The resulting evaluation setup addresses domain-specific challenges including collection representation imbalances, non-linear research patterns, and tensions between specialized expertise and broader accessibility. We propose directions for tailored metrics in each stage of this research journey, such as research path quality for discovery, contextual appropriateness for interaction, metadata-weighted relevance for integration, and cross-stakeholder value alignment for impact assessment. Our contributions extend beyond digital archives to the broader RecSys community, offering transferable evaluation approaches for domains where value emerges through sustained engagement rather than immediate consumption.
Forward citations
Cited by 1 Pith paper
-
Multistakeholder Fairness in Tourism: What can Algorithms learn from Tourism Management?
A comparative literature review shows tourism management and computer science define multistakeholder fairness differently, and argues algorithmic design should adopt qualitative, participatory methods from tourism research.
Reference graph
Works this paper leans on
-
[19]
Wenqi Li, Pengyi Zhang, and Jun Wang. 2024. Analysing humanities scholars’ data seeking behaviour patterns using Ellis’ model. Information Research an international electronic journal 29, 2 (June 2024), 401–418. doi:10.47989/ir292835
-
[2]
Geiger, Georg Vogeler, and Do- minik Kowald
Florian Atzenhofer-Baumgartner, Bernhard C. Geiger, Georg Vogeler, and Do- minik Kowald. 2024. Value Identification in Multistakeholder Recommender Systems for Humanities and Historical Research: The Case of the Digital Archive Monasterium.net. arXiv:2409.17769 (Sept. 2024). doi:10.48550/arXiv.2409.17769 arXiv:2409.17769
-
[1]
Geiger, Christoph Trattner, Georg Vogeler, and Dominik Kowald
Florian Atzenhofer-Baumgartner, Bernhard C. Geiger, Christoph Trattner, Georg Vogeler, and Dominik Kowald. 2024. Challenges in Implementing a Recommender System for Historical Research in the Humanities. arXiv:2410.20909 (Oct. 2024). doi:10.48550/arXiv.2410.20909 arXiv:2410.20909
-
[3]
Ashmi Banerjee. 2023. Fairness and Sustainability in Multistakeholder Tourism Recommender Systems. In Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization . ACM, Limassol Cyprus, 274–279. doi:10.1145/3565472.3595607
arXiv 2023
-
[4]
Christine Bauer, Chandni Bagchi, Olusanmi A. Hundogan, and Karin Van Es. 2024. Where Are the Values? A Systematic Literature Review on News Recommender Systems. ACM Transactions on Recommender Systems 2, 3 (Sept. 2024), 1–40. doi:10.1145/3654805
doi:10.1145/3654805 2024
-
[5]
Robin Burke, Gediminas Adomavicius, Toine Bogers, Tommaso Di Noia, Dominik Kowald, Julia Neidhardt, Özlem Özgöbek, Maria Soledad Pera, Nava Tintarev, and Jürgen Ziegler. 2025. De-centering the (Traditional) User: Multistakeholder Evaluation of Recommender Systems. arXiv:2501.05170 (Jan. 2025). doi:10.48550/ arXiv.2501.05170 arXiv:2501.05170
-
[6]
Robin Burke, Gediminas Adomavicius, Toine Bogers, Tommaso Di Noia, Dominik Kowald, Julia Neidhardt, Özlem Özgöbek, Maria Soledad Pera, and Jürgen Ziegler
-
[7]
Zohreh Dehghani Champiri, Seyed Reza Shahamiri, and Siti Salwah Binti Salim
Show all 40 references
-
[8]
Alvise De Biasio, Andrea Montagna, Fabio Aiolli, and Nicolò Navarin. 2023. A systematic review of value-aware recommender systems. Expert Systems with Applications 226 (2023), 120131. doi:10.1016/j.eswa.2023.120131
2023
-
[9]
Yashar Deldjoo, Dietmar Jannach, Alejandro Bellogin, Alessandro Difonzo, and Dario Zanzonelli. 2024. Fairness in recommender systems: research landscape and future directions. User Modeling and User-Adapted Interaction 34, 1 (2024), 59–108. doi:10.1007/s11257-023-09364-z
2024 doi
-
[10]
Milena Dobreva, Andy O’Dwyer, and Pierluigi Feliciati. 2012. Introduction: user studies for digital library development . Facet, 1–18
2012
- [11]
-
[12]
Michael Färber, Melissa Coutinho, and Shuzhou Yuan. 2023. Biases in scholarly recommender systems: impact, prevalence, and mitigation. Scientometrics 128, 5 (2023), 2703–2736. doi:10.1007/s11192-023-04636-2
2023 doi
-
[13]
Chen Gao, Yu Zheng, Nian Li, Yinfeng Li, Yingrong Qin, Jinghua Piao, Yuhan Quan, Jianxin Chang, Depeng Jin, Xiangnan He, and Yong Li. 2023. A Survey of Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions. ACM Transactions on Recommender Systems 1...
2023 doi
-
[14]
Armin Haberl, Jürgen Fleiß, Dominik Kowald, and Stefan Thalmann. 2024. Take the aTrain. Introducing an interface for the Accessible Transcription of In- terviews. Journal of Behavioral and Experimental Finance 41 (2024), 100891. doi:10.1016/j.jbef.2024.100891
2024
-
[15]
Dietmar Jannach and Himan Abdollahpouri. 2023. A survey on multi-objective recommender systems. Frontiers in big Data 6 (2023), 1157899. https://doi.org/ 10.3389/fdata.2023.1157899
2023
- [16]
-
[17]
Anastasiia Klimashevskaia, Dietmar Jannach, Mehdi Elahi, and Christoph Trat- tner. 2024. A survey on popularity bias in recommender systems. User Modeling and User-Adapted Interaction 34, 5 (2024), 1777–1834. doi:10.1007/s11257-024- 09406-0
2024 doi
-
[18]
Emanuel Lacic, Dominik Kowald, Matthias Traub, Granit Luzhnica, Jörg Peter Simon, and Elisabeth Lex. 2015. Tackling Cold-Start Users in Recommender Sys- tems with Indoor Positioning Systems. In 9th ACM Conference on Recommender Systems. ACM. https://ceur-ws.org/Vol-1441/recsys...
2015
- [20]
-
[21]
Gary Marchionini. 2024. Information and library professionals’ roles and respon- sibilities in an AI -augmented world. Journal of the Association for Information Science and Technology 75, 8 (2024), 865–868. doi:10.1002/asi.24930
2024 doi
-
[22]
Gary Marchionini, Catherine Plaisant, and Anita Komlodi. 2003. The People in Digital Libraries: Multifaceted Approaches to Assessing Needs and Impact . The MIT Press, 119–160. doi:10.7551/mitpress/2424.003.0009
2003 doi
- [23]
-
[24]
Anna Marie Rezk, Auste Simkute, Ewa Luger, John Vines, Chris Elsden, Michael Evans, and Rhianne Jones. 2024. Agency Aspirations: Understanding Users’ Preferences And Perceptions Of Their Role In Personalised News Curation. In Proceedings of the CHI Conference on Human Factors ...
2024
-
[25]
Harald Semmelrock, Tony Ross-Hellauer, Simone Kopeinik, Dieter Theiler, Armin Haberl, Stefan Thalmann, and Dominik Kowald. 2025. Reproducibil- ity in machine-learning-based research: Overview, barriers, and drivers. AI Magazine 46, 2 (2025), e70002. https://doi.org/10.1002/aaai.70002
2025 doi
-
[26]
Faisal Shehzad, Maurizio Ferrari Dacrema, and Dietmar Jannach. 2025. A Worrying Reproducibility Study of Intent-Aware Recommendation Models. arXiv:2501.10143 (Jan. 2025). doi:10.48550/arXiv.2501.10143 arXiv:2501.10143
2025 doi
-
[27]
Smith, Aishwarya Satwani, Robin Burke, and Casey Fiesler
Jessie J. Smith, Aishwarya Satwani, Robin Burke, and Casey Fiesler. 2024. Rec- ommend Me? Designing Fairness Metrics with Providers. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . ACM, Rio de Janeiro Brazil, 2389–2399. doi:10.1145/3630106.3659044
2024
-
[28]
2024.Report on NORMalize: The Second Workshop on the Normative Design and Evaluation of Recommender Systems
Alain Dominique Starke, Sanne Vrijenhoek, Lien Michiels, Johannes Kruse, and Nava Tintarev. 2024.Report on NORMalize: The Second Workshop on the Normative Design and Evaluation of Recommender Systems . https://bora.uib.no/bora-xmlui/ handle/11250/3185749
2024
-
[29]
Jonathan Stray, Alon Halevy, Parisa Assar, Dylan Hadfield-Menell, Craig Boutilier, Amar Ashar, Chloe Bakalar, Lex Beattie, Michael Ekstrand, Claire Leibowicz, Connie Moon Sehat, Sara Johansen, Lianne Kerlin, David Vickrey, Spandana Singh, Sanne Vrijenhoek, Amy Zhang, McKane An...
2024
-
[30]
Lawrence Van Den Bogaert, David Geerts, and Jaron Harambam. 2024. Putting a Human Face on the Algorithm: Co-Designing Recommender Personae to Democratize News Recommender Systems. Digital Journalism 12, 8 (Sept. 2024), 1097–1117. doi:10.1080/21670811.2022.2097101
2024 arXiv
-
[31]
Hanne Vandenbroucke and Annelien Smets. 2024. It’s (not) all about that CTR: A Multi-Stakeholder Perspective on News Recommender Metrics. In 18th ACM Conference on Recommender Systems . ACM, Bari Italy, 999–1003. doi:10.1145/ 3640457.3688183
2024
-
[32]
Shoujin Wang, Qi Zhang, Liang Hu, Xiuzhen Zhang, Yan Wang, and Charu Aggarwal. 2022. Sequential/Session-based Recommendations: Challenges, Ap- proaches, Applications and Opportunities. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in ...
2022
-
[33]
Kathrin Wardatzky, Oana Inel, Luca Rossetto, and Abraham Bernstein. 2025. Whom do Explanations Serve? A Systematic Literature Survey of User Charac- teristics in Explainable Recommender Systems Evaluation. ACM Transactions on Recommender Systems (Feb. 2025), 3716394. doi:10.11...
2025 doi
-
[34]
Alan Jay Wecker, Tsvi Kuflik, Tsafrir Goldberg, Joel Lanir, and Tal Tabashi
-
[35]
Eva Zangerle and Christine Bauer. 2023. Evaluating Recommender Systems: Survey and Framework. Comput. Surveys 55, 8 (Aug. 2023), 1–38. doi:10.1145/ 3556536
2023
-
[36]
Zitong Zhang, Braja Gopal Patra, Ashraf Yaseen, Jie Zhu, Rachit Sabharwal, Kirk Roberts, Tru Cao, and Hulin Wu. 2023. Scholarly recommendation systems: a literature survey. Knowledge and Information Systems 65, 11 (2023), 4433–4478. doi:10.1007/s10115-023-01901-x
2023 doi
- [37]
-
[2015]
doi:10.1016/j.eswa.2014.09.017
A systematic review of scholar context-aware recommender systems.Expert Systems with Applications 42, 3 (2015), 1743–1758. doi:10.1016/j.eswa.2014.09.017
2015 doi
-
[2023]
In Adjunct Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization
Using Recommendations to Affect Social Change in Cultural Heritage: Should We and How?. In Adjunct Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization . ACM, Limassol Cyprus, 419–421. doi:10.1145/3563359.3596663
-
[2024]
Dagstuhl Report on Evaluation Perspectives of Recommender Systems: Driving Research and Education (2024)
Dagstuhl Seminar on Evaluation Perspectives of Recommender Systems: Multistakeholder and Multimethod Evaluation. Dagstuhl Report on Evaluation Perspectives of Recommender Systems: Driving Research and Education (2024). https://doi.org/10.4230/DagRep.14.5.58
2024 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.