{"id":"3bbcf6d4-2a52-4ac8-a136-ca706bca96c8","arxiv_id":"2607.10590","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Across five subjective tasks and five open-source LLMs, demographic prompting improves human agreement only for 1–3 high-signal, directionally coherent attributes and degrades under the full attribute set.","lead":"LLM agreement with human annotators peaks when prompts use only one to three coherent demographic attributes and falls when the full attribute set is supplied. The result gives practitioners a concrete rule for persona prompting instead of dumping every available demographic cue.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified to the over-specification threshold claim; the combinatorial evidence is robust within the stated scope.","rationale":"The reader correctly isolates the combinatorial over-specification finding as the strongest, most robust claim and correctly treats the Fisher ¯ρ coherence measure as a secondary, correlational diagnostic whose validity is limited by small-n rank correlations and English-only data. That diagnostic is not load-bearing for the primary claim: the peak-and-degrade pattern is established by direct enumeration and bootstrap significance tests independent of any dataset-side signal analysis. The paper’s own Limitations section already flags the exact external-validity concerns (model scale, language, distillation artifact for DeepSeek). Because no internal flaw threatens the central combinatorial result, the CONDITIONAL verdict with high confidence remains appropriate; no adjustment is warranted.","tokens_in":32623,"tokens_out":495,"duration_ms":4714,"concrete_test":"Re-run the full combinatorial sweep for Toxicity and Offensiveness on one larger open model (e.g., Llama-3.1-70B or Qwen2.5-32B) using the identical three prompt paraphrases and quadratic-weighted κ. If the peak remains at 1–3 attributes and the full set still underperforms the compact optimum by a statistically significant margin, the threshold generalizes beyond the 3–12 B regime; otherwise the claim is scale-limited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (alignment peaks at 1–3 high-signal attributes and degrades under the full set) rests on exhaustive enumeration of combinations across five tasks and five models (Figures 1–5; Tables 12–16). The pattern is consistent: the full-attribute configuration is never best and is frequently a statistically significant degradation relative to the compact optimum (paired bootstrap, B=10 000). The reader’s weakest assumption correctly flags that Fisher-averaged ¯ρ is only a correlational diagnostic for RQ2/RQ3 and does not underwrite the combinatorial result itself. No internal inconsistency, selection artifact, or metric failure undermines the over-specification threshold on the tested English 3–12 B models. The claim is therefore secure within the paper’s own scope; remaining risks are external generalization (larger proprietary models, non-English data, non-persona prompting styles), which the authors already list as limitations.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies how supplying annotator demographic attributes as prompt cues affects LLM–human agreement on five subjective tasks (toxicity, sentiment, politeness, offensiveness, emotion), using five open-source models (3B–12B). It enumerates essentially all attribute combinations from single-attribute through full-attribute prompts, and reports three findings: (i) agreement typically peaks with one to three high-signal attributes and degrades under the full set (an over-specification threshold); (ii) raw demographic influence on human labels (SHAP) does not predict which attributes help LLMs, whereas jointly considering lexical learnability (LSVC word×demographic interactions) and directional coherence (Fisher-averaged Spearman ¯ρ of subgroup weight vectors) better organizes when prompting helps or hurts; (iii) specialized-neuron activation proportion correlates with alignment gains only under coherent signals, and high activation volume (notably DeepSeek) does not imply steerability. Alignment is measured primarily by quadratic-weighted Cohen’s κ (accuracy for emotion), with bootstrap tests and three prompt paraphrases.","tokens_in":32857,"tokens_out":964,"duration_ms":22326,"significance":"If the over-specification pattern holds more broadly, the work supplies a concrete, actionable constraint for persona-style demographic prompting: more attributes are not better, and full-attribute prompts are often actively harmful. The exhaustive combinatorial design across five tasks and five models is a genuine advance over prior single-attribute or all-vs-none comparisons, and the three-level dataset-side diagnostic (magnitude / learnability / coherence) plus the first application of specialized-neuron probing to demographic alignment are useful contributions for both practitioners and interpretability work. Strengths include paired bootstrap CIs (B=10,000), multi-paraphrase averaging, explicit significance markers in the main tables, an honest limitations section (English-only, small models, small-n rank correlations, correlational neurons, distillation caveat for DeepSeek), and practical task/model recommendations in Appendix J. The central combinatorial claim is well supported within the stated scope; the explanatory RQ2/RQ3 framework is more provisional but still informative.","major_comments":[{"comment":"Figures 1–5 and the accompanying narrative report the best κ (or accuracy) among all combinations of a given size k. For tasks with many attributes (e.g., Toxicity, n=8), C(8,3)=56 and C(8,4)=70, so the intermediate-k peaks are maxes over large candidate sets, while the full-attribute point is a single configuration. This selection asymmetry can inflate the apparent “peak at 1–3” even if the degradation of the full set is real. Please either (a) also report mean/median (and quantiles) of κ over all configs of size k, or (b) apply a multiple-comparison-aware procedure when declaring a size-k optimum, and state clearly that the over-specification claim rests primarily on full-set degradation vs. compact optima rather than on the precise location of the max.","section":null},{"comment":"§4.2 / Tables 1–2: the Spearman correlations that underwrite the three-level framework are computed over only 5–9 attributes per task (as few as five for Politeness and Offensiveness). At this n, rank swaps move ρ substantially; several “significant” cells rest on very small samples. The Limitations section already flags this, but the Abstract and §5 still present learnability+coherence as a principal finding on equal footing with the combinatorial result. Please either aggregate evidence more robustly (e.g., task-pooled or model-pooled tests, bootstrap of the rank correlations themselves) or demote the language so that RQ2 is framed as a diagnostic hypothesis supported by consistent directional patterns, not as a firmly established predictor.","section":null},{"comment":"§4.2 and Appendix G: two concrete cases sit awkwardly with the joint learnability+coherence story. On Offensiveness, race has low Fisher ¯ρ (+0.058) yet is the only attribute that significantly improves any model (Qwen); on Emotion, age drives the largest gains for Mistral/Qwen despite ranking below country/field_of_study on both LSVC accuracy and Fisher ¯ρ. These counterexamples do not refute the framework, but they show that coherence+learnability is neither necessary nor sufficient in every model–task cell. The main text should discuss these cases explicitly and state what residual factors (architecture, baseline strength, subgroup granularity) remain after the three-level account.","section":null}],"minor_comments":[],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that they actually ran the full combinatorial space of demographic attributes (single through all) on five subjective tasks and five open models, and the pattern is clean: alignment peaks at one to three high-signal attributes and the full set is never best, often a significant drop. Prior work only compared single attributes or all-versus-none, so this over-specification threshold is new and well-supported by the tables and bootstrap tests.\n\nThey do the rest carefully. Three prompt paraphrases, quadratic-weighted κ with CIs, SHAP for magnitude, LinearSVC word×demo interactions for learnability, and Fisher-averaged Spearman on the interaction weights for directional coherence. The second finding lands: raw SHAP importance does not predict which attributes help the LLM; you need both learnability and whether subgroups point the same way. The neuron probing (top-10 specialized activations vs baseline) is the first application of that method to demographic alignment; it only correlates with gains when the signal is coherent, and DeepSeek’s high-volume/low-steerability result is a useful caution. Limitations section is honest about English-only data, 3–12B open models, distillation effects, small-n rank correlations, and purely correlational neuron analysis.\n\nSoft spots are real but proportionate. The Fisher ¯ρ diagnostic is correlational and may not fully capture what a persona prompt can steer; the Spearman tables for Politeness/Offensiveness rest on five attributes, so treat the exact ρ values as trends. No code release, and DeepSeek’s behavior may be a distillation artifact. None of that undercuts the combinatorial peak-and-degrade result on the models and tasks they tested.\n\nThis is for people building LLM-as-annotator or persona pipelines who need concrete guidance on how many attributes to include. The math and citation pattern look solid; no circularity. I would send it to peer review and would bring it to reading group. Worth citing if you work on sociodemographic prompting or evaluation of subjective tasks.","headline":"Solid combinatorial study: more demographic attributes in the prompt reliably hurt LLM–human agreement after 1–3 high-signal ones; the rest is useful diagnostics with known correlational limits.","tokens_in":33447,"tokens_out":499,"would_cite":true,"duration_ms":6410,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLM-human agreement peaks with one to three demographic attributes and falls when the full set is used.","keywords":["demographic prompting","LLM-human alignment","over-specification","directional coherence","neuron probing","subjective NLP tasks","persona prompts"],"falsifier":"Run the same full combinatorial prompting and coherence analysis on a held-out subjective task or a substantially larger proprietary model; if full-attribute prompts then outperform the one-to-three-attribute peak, or if high-coherence attributes no longer predict alignment gains, the over-specification threshold and the coherence claim fail.","tokens_in":33519,"feed_emoji":"📉","tokens_out":690,"duration_ms":8665,"temperature":0.7,"pith_summary":"This paper asks how much demographic detail you should put into a prompt when you want a language model to match a particular group's judgments on subjective tasks such as toxicity, sentiment, politeness, offensiveness, and emotion. The authors run every combination of available demographic attributes across five open-source models and five datasets. They find that agreement with human labels is usually highest when the prompt names only one to three high-signal attributes; stuffing in the full demographic profile reliably hurts. Simply knowing which attributes most influence human labels is not enough to choose the right ones. What matters is whether each attribute carries a learnable and directionally coherent lexical signal that a single persona can exploit. Neuron-level measurements further show that specialized activation tracks alignment gains only when that signal is coherent; more activated neurons alone do not make a model more steerable. The practical takeaway is that demographic prompting is not a one-size-fits-all lever: less is often more, and attribute quality matters more than quantity.","feed_headline":"More demographic attributes can hurt LLM-human agreement","feed_subtitle":"Agreement peaks at one to three high-signal cues; full profiles degrade performance across five tasks.","key_machinery":"The three-level diagnostic of attribute signal quality: magnitude (SHAP importance of demographics for human labels), learnability (LinearSVC kappa on word-by-demographic interaction features), and directional coherence (Fisher-averaged Spearman correlation of subgroup lexical weights). This framework, together with combinatorial enumeration of every attribute subset and specialized-neuron probing against a no-persona baseline, is what carries the argument that more demographic detail is not better.","core_discovery":"Alignment between LLM predictions and human annotations consistently peaks with one to three high-signal demographic attributes and degrades under the full attribute set, establishing a clear over-specification threshold. The magnitude of an attribute's influence on human labels does not predict alignment gains; both learnability of word-demographic patterns and the directional coherence of those patterns must be considered jointly. Specialized neuron activation correlates with better alignment only when the underlying signal is coherent, and activation volume alone does not imply steerability.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM-human agreement peaks with 1-3 demographic cues then falls","Full demographic sets degrade LLM annotation alignment","Over-specifying attributes hurts LLM-human agreement","High-signal cues beat full demographic profiles for alignment","Demographic prompting gains reverse past three attributes"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The paper assumes that the Fisher-averaged rank correlation of word-demographic interaction weights is a valid measure of the kind of directional coherence a persona prompt can actually exploit, and that this measure generalizes beyond the five English datasets and five mid-size open models tested.","fun_headline_variants_meta":{"raw":{"variants":["LLM-human agreement peaks with 1-3 demographic cues then falls","Full demographic sets degrade LLM annotation alignment","Over-specifying attributes hurts LLM-human agreement","High-signal cues beat full demographic profiles for alignment","Demographic prompting gains reverse past three attributes"]},"model":"grok-4.5","effort":"low","cost_usd":0.006316,"raw_usage":{"total_tokens":1616,"prompt_tokens":749,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":63160000,"prompt_tokens_details":{"text_tokens":749,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":810,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":749,"tokens_out":57,"duration_ms":8535,"temperature":1.0,"reasoning_tokens":810,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T10:36:12.318209+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same full combinatorial prompting and coherence analysis on a held-out subjective task or a substantially larger proprietary model; if full-attribute prompts then outperform the one-to-three-attribute peak, or if high-coherence attributes no longer predict alignment gains, the over-specification threshold and the coherence claim fail.","supporting_citations":[],"review_version":1}