{"id":"3bc03a3a-5616-41a6-8339-2809a24b58cb","arxiv_id":"2604.16258","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM-generated competency questions exhibit distinct profiles in readability, relevance, and complexity that vary by model type and use case.","lead":"This paper introduces quantitative measures to compare competency questions generated by various open and closed LLMs across domains, focusing on readability, relevance to input text, and structural complexity. Smart generalists and ontology engineers might read it to understand how generative AI can scale requirement elicitation and the distinct output profiles of different models.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Validity of new quantitative measures as proxies for CQ utility in ontology engineering remains untested","rationale":"The reader's weakest assumption directly identifies the missing link between the new metrics and downstream utility. The work is a descriptive cross-model comparison; its interpretive claim about 'distinct profiles shaped by the use case' therefore inherits the unvalidated proxy status of the metrics. Full-text methods/results sections would be needed to confirm whether any indirect support (e.g., inter-metric consistency or qualitative examples) exists, but the abstract alone leaves this gap load-bearing.","tokens_in":1753,"tokens_out":320,"duration_ms":27652,"concrete_test":"Sample 50 generated CQs across two use cases; obtain blind ratings from 5 ontology engineers on practical utility (e.g., clarity for domain experts, coverage of requirements); compute Spearman rank correlation between these ratings and the paper's quantitative scores—if <0.4, the measures do not support the claimed profiles.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that LLMs exhibit distinct generation profiles shaped by use case—rests on the introduced measures (readability, relevance to input text, structural complexity) being meaningful indicators. The paper defines these via standard NLP metrics and applies them across models and scenarios, but supplies no correlation analysis, expert validation, or downstream task evaluation showing that differences in these scores predict better ontology requirements, expert preference, or improved engineering outcomes. Observed profile differences could therefore be artifacts of the chosen metrics rather than substantive, use-case-driven distinctions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a cross-domain empirical study of competency questions (CQs) generated by LLMs for ontology engineering. It defines quantitative measures of readability, relevance to input text, and structural complexity, applies them to outputs from open models (KimiK2-1T, Llama3.1-8B, Llama3.2-3B) and closed models (Gemini 2.5 Pro, GPT-4.1) across multiple use cases, and concludes that LLMs exhibit distinct generation profiles shaped by the use case.","tokens_in":1853,"tokens_out":514,"duration_ms":57047,"significance":"If the proposed measures can be shown to correlate with downstream ontology-engineering outcomes, the work would offer practical guidance on model selection for automated CQ generation and help broaden access to ontology engineering. The systematic inclusion of both open and closed models together with a multi-domain design is a clear strength.","major_comments":[{"comment":"The central claim that LLMs display use-case-shaped generation profiles rests on the three newly introduced quantitative measures (readability, relevance, structural complexity). The manuscript defines these via standard NLP metrics but supplies no validation—such as correlation with expert CQ quality ratings, inter-annotator agreement, or performance on a downstream ontology task—demonstrating that differences in the scores predict actual utility. Without this, observed profile differences risk being artifacts of the chosen proxies rather than substantive distinctions (see the skeptic note on untested validity of the measures).","section":"Section introducing the quantitative measures"},{"comment":"The results and analysis sections lack essential experimental details required to assess robustness: number of CQs generated per use case and model, temperature or sampling settings, number of independent runs, statistical tests used to declare 'distinct profiles,' and any data-exclusion criteria. These omissions make it impossible to determine whether the reported differences are reliable or merely reflect sampling variability.","section":"Results and analysis"}],"minor_comments":[{"comment":"The abstract refers to 'well defined use cases and scenarios' without enumerating them; a short list or reference to the specific domains would improve readability and allow readers to judge generalizability.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The work is an early empirical exploration; the journal may wish to request that the authors add at least a small expert-validation study or downstream-task correlation before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.3","letter":"The core contribution is an empirical comparison of how several LLMs turn use-case text into competency questions. They apply the same set of measures across open models like Llama 3.1/3.2 and Kimi and closed ones like GPT-4.1 and Gemini 2.5, then report that the outputs show different profiles depending on the domain and scenario. That gives a concrete picture of variation that was not previously quantified for this task in ontology engineering.","headline":"This paper runs a cross-model comparison of LLM-generated competency questions using new metrics for readability, relevance, and complexity, but does not test whether those metrics track actual usefulness in ontology work.","tokens_in":2357,"tokens_out":177,"would_cite":false,"duration_ms":39675,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-05-10T08:35:54.669346+00:00","model_set":{"reader":"grok-4.3"},"falsifier":null,"supporting_citations":[],"review_version":1}