{"id":"5a6fc849-8476-46c0-85f7-7954addbf6fe","arxiv_id":"2505.10839","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Users given 78 value-based feed controls expressed their preferences more precisely and used a wider variety of values than users limited to a single value system, supporting the value-library approach.","lead":"A browser extension re-ranks each user's X/Twitter feed using 78 values, such as humor, helpfulness, or calm, chosen by the user and detected by an AI model. Two user studies with 269 participants suggest that people can express their feed preferences more precisely when given this large library than when limited to one value framework.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'more precise' result is confounded: Full and Single conditions differed in onboarding and visible option set, so displaced weights may reflect presentation, not articulation precision.","rationale":"I read the paper in good faith: it is a plausible HCI contribution, the system is real, the two studies are honestly reported, and the classifier validation on 12 values is a reasonable start rather than evidence of bad faith. The reader's verdict of CONDITIONAL is appropriate. I partially agree with the reader's weakest_assumption: the unvalidated 78-value zero-shot classifier is a genuine limitation, since labeling noise also propagates into value merging and real-time re-ranking. However, I think the more directly load-bearing weakness for the central claim is the precision comparison itself. Even a perfectly accurate classifier would not rescue the 'more precise' conclusion if the Full and Single conditions differ in onboarding, option set size, and recommendation flow, and if 'precision' is inferred from weight displacement rather than measured directly. The paper's own coverage analysis is about library coverage, not articulation precision. This is an addressable empirical concern: a matched design and a direct measure of articulation quality would settle it. It does not invalidate the system or the qualitative findings, so I keep the reader's CONDITIONAL verdict unchanged.","tokens_in":28536,"tokens_out":6089,"duration_ms":72851,"concrete_test":"Re-analyze the existing quantitative logs after removing the five seeded onboarding values and all values shown on the dynamic recommendation pages, and restrict the Full condition to participants who did not open the recommendation panel after onboarding. If the displacement effect and the Full-vs-Single differences in Table A.5 persist for non-recommended values, the presentation confound is less likely. Stronger and decisive: run a preregistered matched experiment in which Full and Single conditions use identical onboarding—same number of initially visible values, same recommendation mechanism, only the underlying pool differs—and measure articulation precision directly by having blind coders judge whether each participant's configured value profile captures that participant's free-text description of the feed they wanted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim—that users can articulate desired values more precisely with the full library—rests on a comparison that is not internally clean. In the Quantitative User Study, Full participants completed a three-page onboarding flow with five seeded values and then dynamic recommendations before seeing the rest of the library, while Single participants were 'directly shown the list of values'. The two conditions therefore differ in visible option count, presentation order, and recommended subsets. The precision interpretation is built on the displacement result: mean absolute weights for 17 of 18 significantly differing values drop in the Full condition. But lower absolute weights are exactly what would be expected if users spread attention across more options, followed onboarding recommendations, or used sliders more conservatively—without any gain in precision. Table A.4 strengthens this worry: the five onboarding values ('Knowledge, informativeness', 'A world at Peace', 'Collectivism', 'Appreciation', 'Education and Entertainment') are among the most-used values in the Full condition, which indicates that presentation shapes selection. The qualitative comparison (10 of 12 prefer full) is also always ordered single-first then full, so order and familiarity are confounded. The 55-of-84 coverage result shows that the full library covers more missing values, which is evidence for coverage, not for more precise articulation. Thus the load-bearing claim 'users can articulate their desired values more precisely' is not yet established by the reported data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Alexandria, an extensible library of 78 values for social media feed ranking, constructed from six published value systems (Rokeach, Maslow, Hofstede, Stray et al., Weld et al., and Ge and Gretzel). Each value is operationalized as a zero-shot GPT-4o-mini classifier that rates whether a post weakly or strongly expresses the value, and the library is instantiated as a Chrome extension that re-ranks a user's X/Twitter 'For You' feed in real time based on user-selected upranking/downranking weights. The paper reports two user studies: a qualitative interview study (N=12) and a between-subjects quantitative experiment (N=257) comparing a condition with access to the full 78-value library against conditions with access to a single source value system. The central empirical claim is that users can articulate their desired values more precisely when given access to the full library, supported by the observation that mean absolute weights for 17 of 18 significantly differing values are lower in the full-library condition, as well as by qualitative preferences (10 of 12 participants preferred the full library) and by a gap analysis in which 55 of 84 missing-value requests in the single-system conditions had counterparts in the full library.","tokens_in":28641,"tokens_out":4945,"duration_ms":50094,"significance":"If the central claim holds, the paper makes a useful contribution to value-sensitive design and end-user algorithmic control: it demonstrates that a pluralistic, extensible value library can be operationalized with LLM classifiers and deployed today as a browser extension, without requiring platform-level changes. The paper's strengths include a real deployment on X/Twitter, a between-subjects human experiment, independent human annotation for a subset of classifiers, and a transparent limitations section that concedes classifier inconsistency and onboarding influence. The coverage result (fewer than 3% of tweets contain no library value) and the 55-of-84 gap-match result are credible evidence for the breadth of the library. However, the load-bearing precision claim is currently supported by a confounded comparison and by a measure (weight displacement) that is compatible with attention spreading rather than articulation precision. The validation of the classifiers also covers only 12 of 78 values, which limits the strength of the deployability claim.","major_comments":[{"comment":"The central claim that users 'articulate their desired values more precisely' with the full library is not supported by the current comparison because the Full and Single conditions differ in presentation and onboarding. Full participants completed a three-page onboarding flow with five seeded values ('A World at Peace', 'Knowledge, Informativeness', 'Collectivism', 'Appreciation', 'Education and Entertainment') and then received dynamic recommendations, whereas Single participants were 'directly shown the list of values'. Lower absolute weights in the Full condition (Table A.5) are exactly what would be expected if users distributed weight across more options or followed onboarding suggestions, not if they expressed preferences more precisely; Table A.4 shows that the five onboarding values are among the most-used values in the Full condition, with usage rates between 72.7% and 88.6%. To support the precision claim, the paper needs a measure of articulation precision that is independent of option count and presentation (for example, a matching task against an independently elicited preference inventory, or a within-subject crossover design), or the conclusion should be reframed as improved coverage and flexibility rather than precision.","section":"Quantitative User Study, Method and Results"},{"comment":"The validation of the LLM classifier covers only 12 of the 78 values, each with 30 posts, and the samples were stratified by the LLM's own labels. The same LLM pipeline is then used to compute the pairwise correlations that drive the value-merging step (r >= 0.6) and to power the re-ranking itself, so labeling error propagates into the library's construction for the 66 unvalidated values. The paper's own Limitations section concedes that zero-shot LLM classifiers 'may perform inconsistently across different values', and Table A.1 already shows a spread from MAE 0.28 to 0.75. If unvalidated classifiers are systematically unreliable, the re-ranking behavior and the merging-derived library composition could change substantially. Please validate all values or a larger representative sample, or restrict claims about deployable ranking quality to the validated subset and analyze the sensitivity of the merging step to label noise.","section":"Creating a Library of Values, Evaluating Model Performance / Table A.1"},{"comment":"The displacement analysis reports 18 values with p<0.05 from two-sample t-tests across 78 values but does not mention any multiple-comparison correction, unlike the one-sample analysis in the same section, which uses the Benjamini-Hochberg procedure. Under the null, roughly four of 78 comparisons would be expected to reach p<0.05 by chance. Moreover, 'absolute weight decreases' is also predicted by attention spreading across a larger option set and by conservative slider use, so even a corrected significant difference would not by itself establish more precise articulation. Please report corrected p-values (for example, BH-FDR), effect sizes, and a discussion that separates the attention-spread explanation from the precision interpretation. The small Full-condition sample size (N=44) makes this separation especially important.","section":"Quantitative User Study, Results / Table A.5"},{"comment":"The qualitative evidence cited for the precision claim is also confounded: every participant used the single value system first and the full library second, so the 10-of-12 preference could reflect increased familiarity, learning, or recency rather than the library itself. The 55-of-84 coverage result is evidence that the full library contains more candidate values, not that users can articulate their own preferences more precisely; coverage and precision are distinct constructs. A counterbalanced presentation order, or a measure of accuracy against an independently elicited preference set, is needed before 'more precise' can be claimed from this study.","section":"Qualitative User Study, Study Procedure; Results on Full Library Preference"}],"minor_comments":[{"comment":"The opening sentence in the table caption says 'We observe significant value displacement for 17 values' but the table lists 18 rows; please align the count and the narrative.","section":"Table A.5"},{"comment":"In the reference for Knijnenburg et al., 'ACM onference on Recommender Systems' should read 'ACM Conference on Recommender Systems'.","section":"References"},{"comment":"The legend label 'ORIGINATING VALUE SYSTEM' appears above the chart with six colors, but the figure does not provide an explicit mapping from colors to the six source value systems; please add a legend key.","section":"Figure 2"},{"comment":"The text states that the user's Twitter ID is hashed locally before sending data, but also that usage logs include value rankings and Twitter feeds; please clarify whether post content is deidentified before storage or whether only the ID is hashed.","section":"Ethical Considerations"}],"recommendation":"major_revision","confidential_remarks":"This is a promising systems paper with a real deployment and a serious attempt at empirical evaluation. The main risk is that the headline claim ('more precise articulation') is currently overinterpreted: the between-subjects comparison is confounded by onboarding and option count, and the displacement metric is compatible with a much weaker attention-spread explanation. I would not reject the paper, but the revision needs either a sharper analysis of the existing data (e.g., controlling for number of values selected, onboarding effects, FDR-corrected tests) or a targeted follow-up study with a precision measure that is independent of the interface. The classifier validation issue is also worth tightening, particularly because the same labels drive both library construction and evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The library is a real artifact: 78 values distilled from six established value systems, with LLM prompts and a working browser extension that re-ranks X/Twitter. That's a useful contribution for value-sensitive design and algorithmic choice. The operationalization pipeline is generalizable, and the human validation on the 12 sampled values is honest and reasonably good (MAE 0.45). The coverage result—55 of 84 missing values reported by single-system users have counterparts in the full library—is solid evidence for the library's breadth.\n\nThe load-bearing claim, though, is that users articulate preferences 'more precisely' with the full library. The stress-test note gets this right. The Full and Single conditions differ in two ways at once: Full users saw an onboarding flow with five seeded values plus dynamic recommendations before the full list, while Single users were shown their list directly. So the observed displacement of weights—lower absolute weights on 17 of 18 values—could just reflect presentation, option spread, or onboarding influence, not precision. Table A.4 adds to this worry: all five onboarding values are among the most-used in the Full condition. The qualitative comparison is always single-first-then-full, so order is confounded with familiarity. The paper's own Limitations section concedes that onboarding may influence selection, but the main text still sells the precision interpretation.\n\nTwo more soft spots, both addressable. The classifier is load-bearing for both library construction and re-ranking, yet it is validated on only 12 of 78 values; the paper itself concedes zero-shot LLM classifiers may perform inconsistently across different values. And despite the checklist saying reproduction artifacts exist, I don't see a link to code, data, or a commit hash.\n\nWho is this for? Researchers working on recommender-system values, feed customization, and end-user algorithmic control. They'll get a usable library and a clear pipeline. The precision claim needs rework or a more careful framing.\n\nRecommendation: send to peer review. The artifact and the coverage evidence deserve referee time. But the authors should be pushed to fix the comparison—matching onboarding, reporting per-system baselines, and releasing artifacts—before the 'more precise' result is treated as robust.","headline":"A genuinely useful value library, but the paper's central 'more precise articulation' claim is undercut by a confounded comparison that the authors should be pushed to fix.","tokens_in":29323,"tokens_out":1981,"would_cite":true,"duration_ms":20927,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a pluralistic library of 78 values, implemented as LLM-powered classifiers in a browser extension, lets users re-rank their social media feeds more precisely than any single value system can.","keywords":["value pluralism","feed re-ranking","LLM classifiers","social media","user control","browser extension","human values","recommender systems"],"falsifier":"Take the 66 unvalidated values, collect fresh human annotations on a stratified sample of posts not used in the merging step, and compare LLM labels to averaged human labels. If the mean absolute error on the full set exceeds the human disagreement baseline, or binary accuracy falls well below 81.2%, the library's ranking signal is not dependable. A second check would run a field experiment where users configure the same stated preferences in the full library and in a sham condition with only generic values, and measure whether feed composition changes as predicted.","tokens_in":28212,"feed_emoji":"📜","tokens_out":7112,"duration_ms":59621,"temperature":0.7,"pith_summary":"This paper sets out to prove that the human and societal values missing from engagement-driven social media ranking can be operationalized and deployed today, without waiting for platforms to change their algorithms. To do so it builds a library of 78 values drawn from six published value systems, translates each value into a concrete post-level definition, and turns those definitions into language-model classifiers that rate whether a post weakly or strongly expresses the value. These classifiers power a browser extension that re-ranks a user's X/Twitter feed in real time by weighting the values the user selects. Two studies, a qualitative interview study (N=12) and a quantitative experiment (N=257), support the paper's central empirical claim: users can articulate their desired values more precisely when they have the full library than when they are restricted to a single value system. If true, the claim matters because value-based ranking becomes an end-user feature today rather than a platform-level aspiration.","feed_headline":"A 78-value library lets users re-rank their own feeds","feed_subtitle":"A browser extension scores posts with LLMs, letting people uprank wisdom and downrank spam in real time.","key_machinery":"The value library itself is the central artifact: 78 values, each with a short definition, sourced from six value systems and reduced from 111 candidates by merging any pair whose LLM-labeled co-occurrence correlation reached $r \\geq 0.6$. The argument runs on an operationalization pipeline: each value definition is converted into a platform-agnostic prompt, a zero-shot LLM (GPT-4o-mini) rates each post's expression of every active value as $0$ (absent), $1$ (weak), or $2$ (strong), and each post $i$ receives score $s_i = \\sum_v r_{i,v} w_v$, where $w_v$ is the user's weight for value $v$; posts sort descending by $s_i$. The browser extension collects roughly 70 posts from the user's \"For You\" feed, labels them on a server, and reorders the feed in five to ten seconds, repeating whenever the user scrolls or changes value settings. The classifier validation reports binary accuracy of 81.2% and mean absolute error of 0.45 against averaged human labels on 12 values, better than the average human-vote disagreement of 0.61.","core_discovery":"The central claim is that value pluralism is practically achievable: a library of 78 values, rather than any single value system, gives users enough vocabulary to say what they want from a feed. The paper constructs the library by combining six value systems, filtering to post-level constructs, adapting definitions for a zero-shot LLM labeler, and merging correlated values (Pearson $r \\geq 0.6$), then re-ranks posts by the dot product of the LLM's value ratings and the user's weights. In the controlled study, participants with the full library selected more specific values and applied lower absolute weights to any one value, while 65.5% of the values that single-system users reported missing had a counterpart in the full library; 10 of 12 interview participants preferred the full library. The authors conclude that the values criticized as missing from social media ranking can be operationalized and deployed today through end-user tools.","pith_inferences":["The observed \"value displacement\" (weights on a given value shrink when more values are available) suggests users are not simply adding preferences with a bigger menu; they are substituting specific values for broad ones. A direct test would ask users to describe their ideal feed in free text and measure which condition's selected values better predict their description.","Because the merging step uses the same LLM labels that later do the ranking, classifier noise is baked into the library's structure; validating the 66 unvalidated values would reveal whether the 78-value set is stable under better labels or whether some merges should be undone.","The five-to-ten-second re-ranking latency and the dependence on a server-side LLM API mean the current deployment transfers users' feed data to a third party; a local-model variant would be the natural next deployment test, and it would also reveal how much of the perceived control depends on the stock LLM's labeling skill.","If value-based control spreads, the same mechanism that lets users escape engagement-driven content can let them build value-aligned filter bubbles; a testable design response is to pair value re-ranking with a diversity nudge and measure exposure diversity over weeks."],"forward_implications":["Value-based feed ranking can be deployed today as an end-user tool: the paper's extension re-ranks X/Twitter feeds in five to ten seconds without any platform-level change.","A large value library lets users express preferences that a single value system cannot: in the study, 65.5% of missing-value requests had a counterpart in the full library.","The operationalization pipeline is generalizable: any proposed value system with post-level constructs can be translated into ranking objectives and added to the library.","Platform designers can borrow the value-control interface as a complement to engagement ranking, and decentralized platforms can offer server-side value ranking.","The paper's governance discussion implies that an open, community-extensible value library will need moderation and merging rules to avoid incoherent or conflicting values."],"supporting_citations":[{"why":"Supplies the Rokeach Value Survey, one of six source value systems whose entries seed the library.","marker":"Rokeach 1973"},{"why":"Supplies the cultural-dimensions value system used as a second source of general human values.","marker":"Hofstede 1984"},{"why":"Supplies Maslow's hierarchy-of-needs values as a third general-purpose source.","marker":"Maslow and Lewis 1987"},{"why":"Supplies the community-values taxonomy from Reddit, a domain-specific source for feed-ranking values.","marker":"Weld, Zhang, and Althoff 2022"},{"why":"Supplies the recommender-system values taxonomy, the domain-specific source for platform-level ranking objectives.","marker":"Stray et al. 2022"},{"why":"Supplies the Weibo value co-creation taxonomy, a domain-specific source of social-media interaction values.","marker":"Ge and Gretzel 2018"},{"why":"Establishes the LLM-labeling approach for social-science constructs that the operationalization pipeline adapts.","marker":"Jia et al. 2024"},{"why":"Provides the motivating argument that societal values should be embedded into social media algorithms.","marker":"Bernstein et al. 2023"},{"why":"Supports the premise that LLMs can code social-science constructs with agreement comparable to expert annotators.","marker":"Ziems et al. 2024"}],"fun_headline_variants":["78 values let you re-rank your own feed","Pick from 78 values to reshape your feed","Your feed, your values: a library of 78","User control: 78 values for feed re-ranking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that one fully automated language-model prompt can correctly judge, for every one of the 78 values, how strongly a post expresses that value, even though the paper checks this on only 12 values with 30 posts each and uses the same automated labels to decide which values to merge.","fun_headline_variants_meta":{"raw":{"variants":["78 values let you re-rank your own feed","Pick from 78 values to reshape your feed","Your feed, your values: a library of 78","User control: 78 values for feed re-ranking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000146,"raw_usage":{"total_tokens":1174,"prompt_tokens":932,"completion_tokens":242,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":180}},"tokens_in":548,"tokens_out":242,"duration_ms":3026,"temperature":1.0,"reasoning_tokens":180,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:02:26.312188+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 66 unvalidated values, collect fresh human annotations on a stratified sample of posts not used in the merging step, and compare LLM labels to averaged human labels. If the mean absolute error on the full set exceeds the human disagreement baseline, or binary accuracy falls well below 81.2%, the library's ranking signal is not dependable. A second check would run a field experiment where users configure the same stated preferences in the full library and in a sham condition with only generic values, and measure whether feed composition changes as predicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Rokeach Value Survey, one of six source value systems whose entries seed the library."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the cultural-dimensions value system used as a second source of general human values."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Maslow's hierarchy-of-needs values as a third general-purpose source."},{"cited_title":"X.; and Althoff, T","cited_arxiv_id":null,"evidence_quote":"Supplies the community-values taxonomy from Reddit, a domain-specific source for feed-ranking values."},{"cited_title":"Building Human Values into Recommender Systems: An Interdisciplinary Synthesis","cited_arxiv_id":"2207.10192","evidence_quote":"Supplies the recommender-system values taxonomy, the domain-specific source for platform-level ranking objectives."},{"cited_title":"S.; Mai, M","cited_arxiv_id":null,"evidence_quote":"Establishes the LLM-labeling approach for social-science constructs that the operationalization pipeline adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the motivating argument that societal values should be embedded into social media algorithms."}],"review_version":1}