{"id":"43597259-da10-4c5d-9896-478ec0bbe42d","arxiv_id":"2606.12071","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM judges produce a novelty mirage by preferring model-generated research questions over author-anchored references from real papers, while domain experts prefer the references.","lead":"The paper shows that LLM judges overrate the novelty of AI-generated research questions compared to questions reconstructed from real published papers. Human experts reach the opposite conclusion, raising doubts about using LLMs to assess scientific novelty.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Author-anchored RQ reconstruction may embed hindsight bias, undermining its use as neutral reference for novelty judgments.","rationale":"The reader's weakest_assumption directly identifies the load-bearing premise. No other internal inconsistency (e.g., in the reported LLM vs. expert divergence) is visible from the given material; the concern is structural to the benchmark rather than a data or statistical flaw. Full methods would allow the concrete_test above.","tokens_in":1733,"tokens_out":338,"duration_ms":14445,"concrete_test":"From the RQ-Bench construction section, sample 20 papers; have two domain experts independently reconstruct RQs from the identical background/gap/contribution excerpts (blinded to original paper); compute average semantic similarity (e.g., embedding cosine) between the two expert versions and the paper's version. If mean similarity <0.65 or inter-expert agreement is low, the reference points are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (LLM judges produce a novelty mirage while experts correctly prefer author-anchored RQs) requires that the reconstructed references are appropriate anchors for testing novelty assessment. The construction pulls RQs from a paper's own cited background, gaps, and contributions; this risks circularity because the reconstruction necessarily incorporates the paper's realized framing and outcomes. If experts favor these RQs partly because they align with published work rather than because they are objectively more novel, the observed discrepancy does not isolate LLM judging failure. The abstract explicitly flags that these are not the only valid RQs, yet provides no quantitative check on reconstruction fidelity or inter-expert agreement.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces RQ-Bench, a benchmark built from recent arXiv papers in which author-anchored research questions (RQs) are reconstructed from each paper's cited background, gaps, and contributions. These serve as reference points against which model-generated RQs are evaluated via standalone LLM judging, comparative LLM judging, and human expert evaluation. The central finding is that LLM judges consistently rate model-generated RQs as highly novel (producing a 'novelty mirage' that strengthens under comparative evaluation), whereas domain experts prefer the author-anchored references; the paper also notes that many generated RQs are narrow or source-bound, a dimension often missed by LLMs unless explicitly prompted.","tokens_in":1876,"tokens_out":659,"duration_ms":16530,"significance":"If the results hold after methodological clarification, the work is significant because it supplies a concrete, paper-grounded benchmark for testing LLM novelty assessment—an increasingly common use case in scientific ideation. The explicit contrast between LLM and expert judgments, together with the observation that LLMs overlook narrowness/source-boundedness, offers a falsifiable empirical probe into LLM judging reliability. The benchmark construction itself is a constructive contribution, though its value depends on the validity of the reference RQs.","major_comments":[{"comment":"RQ-Bench construction (described in the abstract and methods): the reconstruction of author-anchored RQs from a paper's own cited background, gaps, and contributions risks hindsight bias, as the references necessarily incorporate the realized framing and outcomes of the published work. This is load-bearing for the central claim, because the discrepancy between LLM and expert judgments is interpreted as evidence that experts correctly identify higher novelty in the references; without a quantitative check on reconstruction fidelity or inter-expert agreement, the observed preference may partly reflect alignment with published work rather than objective novelty.","section":"RQ-Bench construction"},{"comment":"Evaluation results (abstract and results section): the abstract reports contradictory outcomes between LLM and human judgments but supplies no sample sizes, statistical tests, controls, or details on how RQs were generated and reconstructed. This prevents verification of the data support for the claim that LLM judges produce a novelty mirage while experts reach the opposite conclusion.","section":"Evaluation results"},{"comment":"Analysis of narrow/source-bound RQs (results/discussion): the observation that many generated RQs are narrow or source-bound and that LLM judges miss this dimension unless explicitly tested is central to the critique of LLM judging, yet the manuscript provides no quantitative measures, inter-rater statistics, or concrete examples to substantiate the claim.","section":"Analysis of narrow/source-bound RQs"}],"minor_comments":[{"comment":"The abstract would be strengthened by including at least one key quantitative result (e.g., mean novelty scores or preference rates) alongside the qualitative description of the mirage.","section":"Abstract"},{"comment":"Notation for 'standalone' versus 'comparative' LLM judging should be defined explicitly on first use to aid readability.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed feedback. We address each major comment below, indicating planned revisions where appropriate.","responses":[{"response":"We acknowledge the risk of hindsight bias in RQ reconstruction. The author-anchored RQs are intentionally derived from each paper's own cited background, gaps, and contributions to serve as concrete, paper-specific reference points rather than claims of maximal novelty. We will expand the methods section with a fuller description of the reconstruction protocol and add an explicit discussion of this limitation and its implications for interpretation. We will also report inter-expert agreement from the human evaluation if not already quantified. The core empirical result—the divergence between LLM and expert judgments on identical RQs—remains informative even if absolute novelty is not claimed.","revision_made":"partial","referee_comment":"[RQ-Bench construction] the reconstruction of author-anchored RQs from a paper's own cited background, gaps, and contributions risks hindsight bias, as the references necessarily incorporate the realized framing and outcomes of the published work. This is load-bearing for the central claim, because the discrepancy between LLM and expert judgments is interpreted as evidence that experts correctly identify higher novelty in the references; without a quantitative check on reconstruction fidelity or inter-expert agreement, the observed preference may partly reflect alignment with published work rather than objective novelty."},{"response":"We agree that the abstract and results section require these details for verifiability. In the revision we will insert sample sizes (number of papers, generated RQs, and evaluated items), statistical tests comparing LLM vs. expert preferences, and expanded methodological descriptions of both RQ generation and reconstruction procedures.","revision_made":"yes","referee_comment":"[Evaluation results] the abstract reports contradictory outcomes between LLM and human judgments but supplies no sample sizes, statistical tests, controls, or details on how RQs were generated and reconstructed. This prevents verification of the data support for the claim that LLM judges produce a novelty mirage while experts reach the opposite conclusion."},{"response":"We will strengthen the results and discussion sections by adding quantitative counts of narrow or source-bound RQs, inter-rater agreement statistics for this classification, and representative examples. This will make the analysis of LLM limitations more concrete and reproducible.","revision_made":"yes","referee_comment":"[Analysis of narrow/source-bound RQs] the observation that many generated RQs are narrow or source-bound and that LLM judges miss this dimension unless explicitly tested is central to the critique of LLM judging, yet the manuscript provides no quantitative measures, inter-rater statistics, or concrete examples to substantiate the claim."}],"tokens_in":1562,"tokens_out":567,"duration_ms":23700,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that LLM judges give high novelty scores to model-generated research questions and even more so in head-to-head comparisons, while domain experts prefer the reconstructed author-anchored versions from actual papers. This mismatch is the core empirical observation.\n\nThe paper builds RQ-Bench from recent arXiv papers by pulling reference RQs out of the cited background, gaps, and contributions sections. It then runs standalone and comparative LLM judgments against human expert ratings on the same items. The finding that generated RQs often turn out narrow or source-bound, and that LLMs miss this unless prompted, adds a practical detail. The direct contrast with experts extends prior LLM-as-judge work without simply repeating it.\n\nThe reconstruction step is the clearest weak point. Because the references are derived from the published paper itself, they may already embed the final framing and outcomes; experts could favor them for that alignment rather than for superior novelty. The abstract notes these are not the only valid RQs but gives no numbers on reconstruction fidelity, inter-expert agreement, or sample sizes. Without those, the size of the LLM-expert gap is hard to judge.\n\nThe work is aimed at groups building AI tools for scientific ideation or running benchmarks on novelty assessment. It raises a usable caution even if the methods need tightening. The paper deserves peer review so the construction details and quantitative controls can be checked.","headline":"LLMs rate generated RQs as more novel than author-anchored ones from real papers while experts disagree, but the reference construction risks hindsight bias.","tokens_in":2337,"tokens_out":353,"would_cite":false,"duration_ms":10875,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLM judges rate model-generated research questions as more novel than the original author-anchored questions from real papers.","keywords":["LLM judges","research questions","scientific novelty","benchmark","AI evaluation","novelty assessment","human expert evaluation","arXiv papers"],"falsifier":"A controlled study in which domain experts, using the same instructions as the LLM judges, rate the novelty of model-generated questions higher than or equal to the author-anchored references would falsify the reported discrepancy.","tokens_in":2655,"feed_emoji":"⚖️","tokens_out":634,"duration_ms":14008,"temperature":0.7,"pith_summary":"The paper examines whether LLMs can serve as reliable judges of scientific novelty by focusing on research questions rather than full ideas. It constructs RQ-Bench from recent arXiv papers, reconstructing author-anchored reference questions directly from each paper's cited background, identified gaps, and stated contributions. Standalone and comparative LLM evaluations both assign high novelty scores to questions generated by models, an effect that strengthens under direct comparison. Domain experts evaluating the same pairs reach the reverse conclusion and favor the author-anchored questions. The work also notes that many generated questions are narrow or tightly bound to their source material, a limitation LLM judges tend to overlook unless explicitly prompted.","feed_headline":"LLM judges rate AI-generated RQs as more novel than real ones","feed_subtitle":"Domain experts instead favor the original author-anchored questions, exposing a mismatch in how novelty is perceived.","key_machinery":"RQ-Bench, a benchmark that reconstructs author-anchored research questions from real papers' cited background, gaps, and contributions to serve as reference points for novelty judgments.","core_discovery":"LLM judges produce a novelty mirage by consistently rating model-generated research questions as highly novel, with the bias intensifying in comparative settings, while domain experts instead prefer the author-anchored reference questions reconstructed from the cited background, gaps, and contributions of actual papers; generated questions are frequently narrow or source-bound, a shortcoming that LLM judges miss unless the dimension is tested directly.","pith_inferences":["Automated novelty assessment may reinforce incremental or source-derived questions rather than genuinely new directions.","The benchmark approach could be extended to evaluate other dimensions of research ideation such as feasibility or empirical promise.","Hybrid evaluation pipelines that combine LLM screening with targeted human review may be required for reliable scientific novelty checks."],"forward_implications":["Standalone LLM judging rates model-generated RQs as highly novel.","Comparative LLM judging increases the preference for model-generated RQs over author-anchored ones.","Domain experts reach the opposite conclusion and favor the author-anchored reference questions.","Many model-generated RQs are narrow or source-bound, a dimension LLM judges miss unless explicitly tested."],"fun_headline_variants":["LLMs rate generated RQs more novel than paper originals","Expert views clash with LLM novelty ratings on RQs","LLM judging favors model RQs while experts pick originals","Generated RQs score high with LLMs but low with experts"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The reconstructed author-anchored research questions from cited background, gaps, and contributions serve as appropriate and representative reference points for testing novelty judgments.","fun_headline_variants_meta":{"raw":{"variants":["LLMs rate generated RQs more novel than paper originals","Expert views clash with LLM novelty ratings on RQs","LLM judging favors model RQs while experts pick originals","Generated RQs score high with LLMs but low with experts"]},"model":"grok-4.3","cost_usd":0.007267,"raw_usage":{"total_tokens":3358,"prompt_tokens":687,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":72674500,"prompt_tokens_details":{"text_tokens":687,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2606,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":687,"tokens_out":65,"duration_ms":14446,"temperature":1.0,"reasoning_tokens":2606,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T07:39:13.944527+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled study in which domain experts, using the same instructions as the LLM judges, rate the novelty of model-generated questions higher than or equal to the author-anchored references would falsify the reported discrepancy.","supporting_citations":[],"review_version":1}