{"id":"d1f0d6f5-ba1d-42ed-9189-fc8bbce62476","arxiv_id":"2412.10558","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Larger language models resist misleading answer hints better than smaller ones while still following legitimate instructions.","lead":"Bigger language models were less likely to be fooled by a fake hint injected into a multiple-choice question, keeping more accuracy than their smaller counterparts. The study also shows these models still follow genuine instructions, suggesting the effect is not simple stubbornness.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central \"consistent scaling\" claim is contradicted by the paper's own appendix: on Logiqa2 all four larger models show larger Relative Accuracy Drop, and on MathQA Llama/Gemma reverse; Figure 2 aggregates away these per-benchmark failures.","rationale":"The reader's weakest assumption is that the model pairs do not isolate parameter scale because they differ in architecture, training data, and instruction-tuning. That is a real external-validity concern, but the more load-bearing problem is internal: even if scale were perfectly isolated, the paper's own per-benchmark results do not support the unconditional claim. Table 10 (Logiqa2) shows all four larger models with larger Relative Accuracy Drop, and Table 11 (MathQA) reverses for Llama and Gemma. The phrase \"consistently observe\" in Section 4.1 is therefore false as written. No significance testing, confidence intervals, or effect sizes are reported, so the aggregate Figure 2 may average away genuine per-benchmark reversals. Additionally, many Deception conditions drive accuracy to near chance, so small fluctuations near the floor can produce large relative-drop differences that are not meaningful resilience. The memorization control is thoughtful, and the overall methodology is transparent enough to be checked, but it does not repair the mismatch between the headline and the appendix tables. I would keep the reader's CONDITIONAL verdict, but the condition should include reporting per-benchmark statistics and restricting or revising the claim to the benchmarks that actually show the direction, rather than only addressing the scale-isolation confound.","tokens_in":20090,"tokens_out":6428,"duration_ms":57545,"concrete_test":"Re-run the Deception experiment with the released prompts and code and report per-benchmark Relative Accuracy Drop for each of the four family pairs, with bootstrap 95% confidence intervals. Then run a paired sign test across the 40 family-benchmark comparisons. If the sign test does not reject the null of equal probability of improvement at p < 0.05, or if Logiqa2 and MathQA still reverse, the unconditional scaling claim should be withdrawn or explicitly restricted to the benchmarks where the direction holds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim in Section 4.1 is that \"within each model family, we consistently observe that the Relative Accuracy Drop is smaller for larger models.\" The per-benchmark tables in Appendix E contradict this. Recomputing Relative Accuracy Drop = (Original - Altered) / Original from Table 10 (Logiqa2): Llama 8B to 70B gives (0.55-0.29)/0.55 = 0.473 versus (0.71-0.32)/0.71 = 0.549; Gemma 2B to 9B gives 0.750 versus 0.923; Phi Mini to Medium gives 0.316 versus 0.444; Mistral 7B to Mixtral gives 0.615 versus 0.678. All four larger models are less resilient on this benchmark. Table 11 (MathQA) also reverses for Llama (0.793 to 1.0) and Gemma (0.714 to 0.955). Thus the scaling trend is not consistent; it exists only in the aggregate mean, and the paper provides no significance test, confidence interval, or effect size to establish that the aggregate is not driven by a few benchmarks or by floor effects. Several Deception conditions push accuracy to near zero (GPQA and MathQA, for example), so the relative-drop comparison can be dominated by small absolute differences near chance. Because the central claim is stated unconditionally, this is a load-bearing empirical failure independent of the family-confounding issue raised by the reader.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates eight open-source instruction-tuned LLMs from four families (Llama-3.1-8B/70B, Gemma-2-2B/9B, Phi-3 Mini/Medium, Mistral-7B/Mixtral-8x22B) on multiple-choice benchmarks, altering prompts by adding a false hint (Deception), a true hint (Guidance), a directive to answer incorrectly (Directive Instruction), or removing the question (Context Removal). Its central claim is that larger models show smaller Relative Accuracy Drop under Deception and are therefore more resilient to misleading in-context information, and that this resilience is not due to ignoring hints or to memorization/contamination. The paper also discusses world-model interpretations and includes qualitative generation examples.","tokens_in":20360,"tokens_out":9999,"duration_ms":85518,"significance":"If the main claim were established, the paper would be a useful empirical contribution on how scale affects susceptibility to misleading in-context information, with implications for model evaluation and safety. The experimental setup is transparent and inexpensive: it reuses standard benchmarks, uses open models, and the Guidance control (near-perfect accuracy with true hints) is a clean way to show that larger models do not simply ignore prompt content. The Context Removal and overfitting comparison is a constructive idea for probing memorization. However, the evidence as reported does not support the strong, unconditional phrasing in Section 4.1: the paper's own per-benchmark tables contain reversals, and no confidence intervals or significance tests are provided. The causal attribution to parameter scale is also not established by the chosen model pairs.","major_comments":[{"comment":"The central claim that \"within each model family, we consistently observe that the Relative Accuracy Drop is smaller for larger models\" is contradicted by the paper's own Appendix E. From Table 10 (Logiqa2), the Relative Accuracy Drop is larger for the larger model in every family: Llama 8B→70B gives (0.55−0.29)/0.55=0.473 vs. (0.71−0.32)/0.71=0.549; Gemma 2B→9B gives 0.750 vs. 0.923; Phi Mini→Medium gives 0.316 vs. 0.444; Mistral 7B→Mixtral gives 0.615 vs. 0.678. Table 11 (MathQA) reverses for Llama (0.793 vs. 1.000) and Gemma (0.714 vs. 0.955). Because the claim is stated unconditionally and Figure 2 plots only the aggregate mean (with an undefined shaded \"deviation\" and no confidence intervals or significance tests), the reversals are load-bearing: the result currently holds in the aggregate, not consistently. The authors should report per-benchmark effect sizes with uncertainty, pre-specify the summary statistic, or substantially weaken the claim. Floor effects make this particularly important: in GPQA and MathQA the Deception condition drives several accuracies to or near zero, so the relative-drop ratio is unstable.","section":"Section 4.1, Figure 2, Tables 10-11"},{"comment":"The headline metric, Relative Accuracy Drop = (Original Accuracy − Altered Accuracy) / Original Accuracy, is normalized by baseline accuracy. Since larger models generally have higher original accuracy, this metric partially builds in the paper's conclusion: even with exactly equal absolute drops, the larger model will have a smaller relative drop. Absolute drops (Figure 7) are more favorable to the claim but are also not universal; for example, Table 11 (MathQA) shows Llama 8B's absolute drop is 0.23 while Llama 70B's is 0.40, and Table 7 (HellaSwag) shows Phi Medium with a lower original accuracy than Phi Mini, so the baseline ordering itself is not consistent. The analysis should report both metrics with confidence intervals and justify the normalization choice rather than presenting it as the only natural standardization.","section":"Section 3.5, Equations for Accuracy Drop and Relative Accuracy Drop"},{"comment":"The paper says the model pairs are chosen \"to isolate the effect of scale on model performance,\" but the pairs differ in more than parameter count. Mistral-7B is a dense model while Mixtral-8x22B is a mixture-of-experts model with a different total parameter count and different active-parameter behavior; the Llama, Gemma, and Phi pairs also differ in training data, instruction-tuning recipes, tokenizers, and possibly architecture details. Parameter scale is therefore confounded with family-specific design choices, so the observed differences cannot be causally attributed to scale alone. The paper should reframe the results as within-family capacity trends, or use a controlled comparison (e.g., same architecture and data with different widths/depths) to support the \"as models scale\" language in the conclusion.","section":"Section 3.3, Models"},{"comment":"The memorization control does not directly test the alternative explanation it targets. The Context Removal experiment and the overfitted-Llama/DCLM-7B comparison show that both a contaminated model and a clean model can answer some MMLU items from the answer choices alone, but they do not show that test-set contamination does not contribute to the larger models' smaller Deception drops. In particular, the overfitted model is not evaluated under the Deception condition, and the comparison is not made between small and large models with matched contamination. The overfitting experiment also uses Llama-3.1-8B only, not the larger models whose resilience is at issue. The paper concedes \"we cannot entirely dismiss the possibility that memorization contributes,\" but Section 5 concludes the resilience is \"unlikely due to memorization.\" This overreach should be fixed by either adding a direct deception experiment on contaminated vs. clean models or by removing the causal claim.","section":"Section 4.3, Context Removal and Overfitting Details"}],"minor_comments":[{"comment":"Table 1 lists nine benchmarks, but Appendix E and Figures 6-7 include Logiqa2, which is never described in the benchmark table; this omission should be fixed for reproducibility.","section":"Table 1 and Appendix E"},{"comment":"The arrow labels in the Logiqa2 panel do not match Table 10: Phi-Mini is labeled −39.7% and Mixtral −17.7%, whereas Table 10 implies drops of approximately 18% and 40%, respectively; the label-to-model mapping appears shifted.","section":"Figure 7, Logiqa2 panel"},{"comment":"The metric example contains missing spaces: \"would be6.25%and 8.33%\" should read \"would be 6.25% and 8.33%\".","section":"Section 3.5, metric example"},{"comment":"Figure 5 shows training epochs up to 30, while Appendix B states the training loop was executed for 50 epochs; the relationship between these numbers should be clarified, and the caption contains the typo \"ALlama-3.1-8B-Instruct\".","section":"Figure 5 caption and Appendix B"},{"comment":"Section 4.2 correctly acknowledges that Gemma is an outlier for instruction-following, but the abstract and Section 5 present instruction-following improvements without that qualifier; the softer claim should be carried through the summary sections.","section":"Section 4.2 and Abstract"}],"recommendation":"major_revision","confidential_remarks":"The reversals in Appendix E are strong enough that the editors should require a revised statistical presentation before considering acceptance. The underlying experiment is valuable, and the Guidance control for the 'ignoring hints' alternative is well designed, so I would not reject the manuscript at this stage. The authors should be asked to either provide a rigorous per-benchmark analysis with confidence intervals or restrict the central claim to the aggregate trend."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read on arXiv:2412.10558.\n\nThe headline—larger models are more resilient to deceptive in-context hints—is a real empirical question, and the evaluation design is mostly sound. Credit where due: reusing existing multiple-choice benchmarks with uniform prompt formatting, adding controls for ignoring hints (truthful hints, directive instructions), and a memorization control using an overfitted model versus DCLM-7B is thoughtful. The result that larger models still use correct hints and follow instructions rules out the simple \"they just ignore the prompt\" story. As a measurement contribution, this is useful.\n\nBut the central claim is overstated. The paper says \"consistently\" within each family, but its own appendix contradicts that. On Logiqa2, all four larger models show larger Relative Accuracy Drop, not smaller. On MathQA, Llama and Gemma reverse. The aggregate Figure 2 averages these away. There are no confidence intervals, significance tests, or effect sizes, so we can't tell if the aggregate trend is real or driven by a couple of benchmarks and floor effects. Several Deception conditions push accuracy near zero (GPQA, MathQA), where relative drop comparisons are dominated by small absolute differences. That's a load-bearing problem for the paper's unconditioned claim.\n\nThere's also a metric issue. Relative Accuracy Drop normalizes by baseline accuracy; larger models have higher baselines, so part of the apparent resilience is built into the metric. The authors report absolute drops too, and those mostly agree, but the narrative relies on the relative metric.\n\nThe family confounding is real but not disqualifying—Llama 8B vs 70B differ in architecture, data, and training, so \"scale\" isn't cleanly isolated. That's a standard limitation, but it should be stated more cautiously.\n\nThe memorization control is clever but indirect: it tests whether models can answer without the question, not whether contamination changes deception susceptibility.\n\nVerdict: This paper deserves serious peer review. The question is timely, the controls are above average, and the per-benchmark data are actually in the appendix, which is honest. But the main claim needs to be reframed as a tendency with exceptions, backed by statistics, and the metric choice needs defending. I'd conditional-accept if I were editor.","headline":"Larger models may resist deception on average, but the paper's 'consistent' scaling claim doesn't survive its own per-benchmark tables.","tokens_in":20861,"tokens_out":2045,"would_cite":false,"duration_ms":18242,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Larger language models are harder to fool than their smaller counterparts.","keywords":["deception","misinformation","relative accuracy drop","in-context information","model scaling","world model","instruction following"],"falsifier":"Train or obtain a small and a large model that share the same architecture, data, and instruction-tuning pipeline and differ only in width, depth, or parameter count; if the larger model's Relative Accuracy Drop under a false hint is not smaller than the smaller model's, the claimed scaling law is refuted. A cheaper check is to re-run the deception experiment on a matched pair where the small model was trained on exactly the same data as the large one.","tokens_in":19890,"feed_emoji":"🤖","tokens_out":4959,"duration_ms":43403,"temperature":0.7,"pith_summary":"Large language models must weigh what they learned during training against information written in the prompt. The paper claims that, within the same model family, the larger model keeps more of its accuracy when the prompt contains an incorrect hint, and that this relative advantage is not explained by ignoring hints or by memorization. If true, scaling makes models harder to fool by injected misinformation while still leaving them responsive to genuine instructions.","feed_headline":"Bigger language models are harder to fool","feed_subtitle":"Across four model families, larger models resist false hints while still following real instructions.","key_machinery":"The central measurement is the Relative Accuracy Drop, the accuracy loss under a prompt alteration divided by the original accuracy; it allows drops to be compared across models and benchmarks of differing difficulty. The evaluation procedure standardizes every benchmark to the MMLU prompt format and applies four alterations — a deceptive hint, a truthful hint, a directive instruction to answer incorrectly, and removal of the question — so that each comparison isolates one way a model can be steered. The contrast between deceptive and directive conditions is what lets the paper claim scale improves both skepticism and instruction following.","core_discovery":"Across four open-weight model families, pairing each small model with a larger sibling, the paper finds that larger models show a smaller Relative Accuracy Drop — the fractional loss defined as (original accuracy minus altered accuracy) divided by original accuracy — when a false answer hint is appended to multiple-choice questions. Control experiments show all models exploit truthful hints nearly perfectly, and larger models follow explicit wrong-answer instructions at least as well as smaller ones, ruling out the idea that resilience comes from disregarding prompt content. A contamination experiment, comparing a model with no possible exposure to the benchmark against one deliberately overfitted on the test set, finds both stay above chance when the question is removed, suggesting the resilience reflects inference from choices and world knowledge rather than rote memorization.","pith_inferences":["Inference: the paper's hints are uniform and simple; realistic misinformation is more nuanced, so the size advantage may shrink or reverse on contextually believable false hints, and a targeted study is needed.","Inference: because the smaller models' accuracy under deception often falls near chance, the relative drop metric may partly reflect a floor effect; absolute drop plots in the paper show the same direction, but the metric's normalization should be stress-tested on models of matched competence.","Inference: one testable extension would be to repeat the deception experiment with the hint presented as an authoritative source (e.g., a domain expert states...) to see whether larger models become more gullible to authority framing.","Inference: the same prompt-alteration suite could be applied to open-ended generations with judge-based correctness, rather than multiple choice, to see whether the scaling pattern generalizes."],"forward_implications":["Larger open-weight models will be proportionally less degraded by injected false hints on multiple-choice benchmarks.","The resilience is not bought by ignoring prompts: accurate hints still raise accuracy nearly to ceiling in all models, so the improvement is in how hints are screened.","Instruction-following improves with scale on this setup, meaning larger models can be directed to wrong answers when explicitly asked, even as they resist unsupported hints.","Models without benchmark contamination behave like overfitted ones when the question is removed, so question-removal accuracy is not evidence of memorization; future benchmark audits should control for choice-only inference.","Scaling is a partial, not complete, defense against misinformation, and the findings motivate studying malicious hints in open-ended generation."],"supporting_citations":[{"why":"Supplies the MMLU benchmark and its prompt format, which is used as the unified structure for all experiments and for the context-removal control.","marker":"(Hendrycks et al., 2021)"},{"why":"Provides the evaluation harness that computes per-choice log-likelihoods and reports accuracy across the benchmark suite.","marker":"(Gao et al., 2024)"},{"why":"Supplies DCLM-7B, a model with no prior exposure to the benchmark, used as the contamination-free control in the memorization experiment.","marker":"(Li et al., 2024a)"},{"why":"Supplies LoRA, the parameter-efficient fine-tuning method used to overfit a model on the evaluation set to simulate severe data contamination.","marker":"(Hu et al., 2021)"},{"why":"Contributes GPQA, a graduate-level science benchmark, as one of the multiple-choice tasks in the evaluation suite.","marker":"(Rein et al., 2023)"},{"why":"Contributes PIQA, a physical commonsense reasoning benchmark, as part of the multiple-choice evaluation suite.","marker":"(Bisk et al., 2019)"}],"fun_headline_variants":["Large LLMs shrug off false hints","Too big to fool: large models resist deception","Scale shields LLMs from deceptive prompts","Size matters: big models resist false clues","Larger LLMs ignore false hints without ignoring real ones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that comparing Llama-8B against Llama-70B, Gemma-2B against Gemma-9B, Phi-mini against Phi-medium, and Mistral-7B against Mixtral-8x22B isolates the effect of parameter count alone, even though the paired models differ in architecture, training data, tokenizer, and fine-tuning recipe.","fun_headline_variants_meta":{"raw":{"variants":["Large LLMs shrug off false hints","Too big to fool: large models resist deception","Scale shields LLMs from deceptive prompts","Size matters: big models resist false clues","Larger LLMs ignore false hints without ignoring real ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000921,"raw_usage":{"total_tokens":3874,"prompt_tokens":794,"completion_tokens":3080,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":410,"completion_tokens_details":{"reasoning_tokens":3012}},"tokens_in":410,"tokens_out":3080,"duration_ms":21651,"temperature":1.0,"reasoning_tokens":3012,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:51:12.749105+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train or obtain a small and a large model that share the same architecture, data, and instruction-tuning pipeline and differ only in width, depth, or parameter count; if the larger model's Relative Accuracy Drop under a false hint is not smaller than the smaller model's, the claimed scaling law is refuted. A cheaper check is to re-run the deception experiment on a matched pair where the small model was trained on exactly the same data as the large one.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes GPQA, a graduate-level science benchmark, as one of the multiple-choice tasks in the evaluation suite."}],"review_version":1}