{"id":"34e099ce-df8c-4143-ad21-24e8f04964d0","arxiv_id":"2605.26937","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Proposes open knowledge evaluation using open-ended elicitation prompts and releases the BeQu benchmark of 10,000 entities with reference corpora to characterize LLMs' naturally expressed parametric knowledge.","lead":"The paper introduces open knowledge evaluation, a paradigm that assesses LLMs by analyzing responses to open-ended prompts like 'Tell me everything you know about X' instead of fixed questions. A smart generalist might read it to understand how current benchmarks may miss what models actually know and how a new benchmark could change evaluation practices.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Open-ended elicitation may still sample only a prompt-shaped subset of parametric knowledge, even with prompt-format analysis.","rationale":"The reader's weakest_assumption matches the load-bearing point exactly; the proposed test would directly test whether the observed knowledge is robust to elicitation choice, which is required for the new paradigm to be less biased than question-based benchmarks.","tokens_in":1686,"tokens_out":311,"duration_ms":18817,"concrete_test":"For the 1000 most frequent BeQu entities, run three additional elicitation templates (e.g., \"List all facts you know about X\", \"What are the key attributes of X?\", \"Describe X in detail\") and measure Jaccard overlap of verified statements against the original template; if average overlap <0.6 or drops further for larger models, the representativeness assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that open-ended prompts (e.g., \"Tell me everything you know about X\") yield a representative sample of what the model actually stores, rather than a biased subset shaped by generation priors or surface-form sensitivity. The abstract states the paradigm shift rests on this (paragraph 2), and the paper reports analyzing prompt format effects. However, format variation alone does not establish completeness: different phrasings could systematically suppress entire classes of facts (e.g., negative or low-probability statements) without the analysis detecting it, leaving the \"naturally express\" characterization unanchored.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that existing LLM knowledge benchmarks suffer from availability bias because they rely on predefined questions selected by benchmark designers. It proposes a new 'open knowledge evaluation' paradigm that instead evaluates the knowledge models naturally express in response to open-ended elicitation prompts (e.g., 'Tell me everything you know about M.L. King'). The paradigm is instantiated with the BeQu benchmark, which consists of 10,000 entities paired with reference corpora for statement verification. The authors evaluate a range of LLMs and analyze the effects of reasoning effort, model scale, prompt format, and knowledge domain, with data and leaderboard made available.","tokens_in":1796,"tokens_out":444,"duration_ms":33356,"significance":"If the central assumption holds, this work offers a meaningful shift in how parametric knowledge in LLMs is assessed, potentially providing a more unbiased characterization by focusing on what models choose to surface rather than what designers query. The scale of the BeQu benchmark (10,000 entities) and the public release of data and leaderboard are positive for reproducibility and further research. The analysis of multiple factors (scale, prompt format, etc.) adds to its potential impact if the method's validity is established.","major_comments":[{"comment":"Abstract, paragraph 2: The core claim that open-ended prompts allow characterization of knowledge models 'naturally express' is load-bearing but rests on the untested assumption that such responses are representative of parametric knowledge rather than shaped by prompt phrasing or generation biases. The paper reports analyzing prompt format effects, but this does not rule out systematic suppression of entire classes of facts (e.g., low-probability or negative statements), as highlighted in the stress-test note. A concrete test or discussion of this completeness is needed to support the paradigm shift.","section":"Abstract, paragraph 2"}],"minor_comments":[{"comment":"Abstract: The abstract mentions 'a broad range of language models' but does not specify which models or scales are evaluated; this detail would help readers assess the scope.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed review and the opportunity to clarify the scope and limitations of the open knowledge evaluation paradigm. We address the single major comment below.","responses":[{"response":"We acknowledge that prompt-format experiments alone do not fully establish completeness with respect to suppressed fact classes. The manuscript already notes generation biases in the stress-test discussion, but we agree a more explicit treatment is warranted. We will expand the Limitations section with a dedicated paragraph on potential systematic omissions (low-probability facts, negative statements, etc.) and outline concrete follow-up tests (e.g., targeted probing for withheld knowledge categories) that future work could perform. This addition will temper the paradigm-shift language while preserving the core contribution of shifting from designer-chosen queries to model-elicited statements.","revision_made":"yes","referee_comment":"[Abstract, paragraph 2] Abstract, paragraph 2: The core claim that open-ended prompts allow characterization of knowledge models 'naturally express' is load-bearing but rests on the untested assumption that such responses are representative of parametric knowledge rather than shaped by prompt phrasing or generation biases. The paper reports analyzing prompt format effects, but this does not rule out systematic suppression of entire classes of facts (e.g., low-probability or negative statements), as highlighted in the stress-test note. A concrete test or discussion of this completeness is needed to support the paradigm shift."}],"tokens_in":1368,"tokens_out":307,"duration_ms":17332,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is a shift from fixed-question benchmarks to open-ended prompts that let models surface what they know, paired with the BeQu dataset of 10,000 entities and reference corpora for verification.\n\nThe work does a clean job naming the availability bias in current tests and then delivering a usable benchmark plus leaderboard. Releasing the data and running checks on scale, reasoning effort, prompt format, and domain gives people something concrete to try and compare against.\n\nThe soft spot sits in the core assumption. Open prompts may still reflect generation biases rather than a neutral sample of stored knowledge, and varying prompt formats alone does not rule out systematic gaps in what gets expressed. The abstract shows they looked at format effects, but without deeper error analysis or direct comparisons it is not clear how much this reduces bias overall.\n\nThis is for researchers who build or critique knowledge evaluations in NLP. Readers who want an alternative to standard QA sets will find the dataset and analyses worth examining.\n\nIt deserves peer review. The resources are public, the methodological point is clear, and referees can assess the empirical claims directly.","headline":"Paper introduces open elicitation for LLM knowledge eval via BeQu benchmark but the representativeness of open prompts remains lightly tested.","tokens_in":2257,"tokens_out":293,"would_cite":false,"duration_ms":32554,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Open knowledge evaluation measures what LLMs naturally express rather than testing only on preset questions.","keywords":["LLM knowledge evaluation","parametric knowledge","availability bias","open knowledge evaluation","elicitation prompts","BeQu benchmark","knowledge benchmarking"],"falsifier":"A study showing that models correctly answer many more facts in direct questions than they volunteer unprompted in open descriptions of the same entities would indicate the open method under-samples knowledge.","tokens_in":2573,"feed_emoji":"","tokens_out":589,"duration_ms":24430,"temperature":0.7,"pith_summary":"Existing knowledge benchmarks for large language models rely on questions chosen by benchmark creators, which introduces availability bias by only testing knowledge that designers decide to query. The paper introduces open knowledge evaluation as an alternative that assesses the knowledge models surface on their own when given open-ended prompts such as requests to tell everything known about an entity. This paradigm is put into practice with the BeQu benchmark, which covers 10,000 entities and uses reference corpora to verify the accuracy of model statements. Analysis using this benchmark examines how model scale, reasoning effort, prompt formats, and knowledge domains influence what is expressed. The result is a method focused on naturally occurring knowledge rather than retrieval of specific answers.","feed_headline":"Open prompts reveal what LLMs know beyond preset questions","feed_subtitle":"New method evaluates models on the knowledge they choose to express in broad responses instead of only human-selected queries.","key_machinery":"Open knowledge evaluation paradigm, which elicits model statements via open-ended prompts about entities and verifies them against reference corpora.","core_discovery":"The paper establishes that evaluating LLMs by their responses to open-ended elicitation prompts, rather than predefined questions, provides a characterization of the knowledge models naturally express, implemented via the BeQu benchmark of 10,000 entities paired with reference corpora for statement verification.","pith_inferences":["The method could help surface training data gaps by identifying topics where models remain silent in open responses.","It might complement traditional benchmarks to give a fuller view of model capabilities.","Future tests could check whether models hold knowledge they consistently fail to volunteer without specific prompting."],"forward_implications":["Larger models express more knowledge when responding to open elicitation prompts.","Increased reasoning effort changes the volume and accuracy of knowledge that surfaces.","Different prompt formats alter which knowledge gets expressed by the model.","Knowledge expression patterns vary across different domains.","The BeQu benchmark supports systematic comparison of these effects across many models."],"fun_headline_variants":["Open prompts evaluate what LLMs naturally know","BeQu assesses LLM knowledge from broad replies","Knowledge benchmarking with open-ended prompts","Evaluating expressed LLM knowledge beyond queries"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Responses to open-ended prompts provide a representative and unbiased sample of a model's parametric knowledge rather than being shaped by prompt phrasing or generation biases.","fun_headline_variants_meta":{"raw":{"variants":["Open prompts evaluate what LLMs naturally know","BeQu assesses LLM knowledge from broad replies","Knowledge benchmarking with open-ended prompts","Evaluating expressed LLM knowledge beyond queries"]},"model":"grok-4.3","cost_usd":0.005304,"raw_usage":{"total_tokens":2455,"prompt_tokens":613,"num_sources_used":0,"completion_tokens":49,"cost_in_usd_ticks":53040500,"prompt_tokens_details":{"text_tokens":613,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1793,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":613,"tokens_out":49,"duration_ms":21470,"temperature":1.0,"reasoning_tokens":1793,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T18:00:50.771543+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A study showing that models correctly answer many more facts in direct questions than they volunteer unprompted in open descriptions of the same entities would indicate the open method under-samples knowledge.","supporting_citations":[],"review_version":1}