{"id":"4cda8605-c5e2-4b0c-9932-268ad154f2f5","arxiv_id":"2504.21605","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The authors build an RDF-based vocabulary called SQARE for structured LLM quality assessments and apply it to a 28-question, two-model, two-language knowledge-conflict study.","lead":"This paper proposes an RDF vocabulary for recording multilingual LLM evaluation results, demonstrated on 28 fire safety questions in German and English. If the approach is adopted, evaluation datasets could be shared and queried in a standardized way, though the demonstration's empirical claims rest on a small, manually labeled sample.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sufficiency claim exceeds what the demonstration can establish, because the 28-question fire-safety corpus is neither diverse nor adversarially sampled and the validation labels are produced by a single un-rubricated annotator, so 'vocabulary sufficient for every facet' is not yet shown.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that outcome. The reader's weakest assumption is about the reliability of the binary correctness labels with no rubric or inter-annotator agreement; my concern is broader and distinct: I identify the sufficiency claim as the strongest claim, and I argue that even with perfectly reliable labels, the claim as stated is under-supported because the evidence is self-referential. The reader's weakest_assumption focuses on the labels; my load-bearing concern focuses on the coverage/sufficiency claim, which is stated in the abstract and Section 5. These are different weaknesses, so agreement_with_reader is partial. My recommendation that the verdict stays CONDITIONAL is consistent with the reader's high-confidence conditional judgment: the framework is a useful proposal with real limitations, and the empirical/sufficiency claims are not yet established. I do not escalate to REJECT because the schema is posted publicly, the construction is coherent, and the core vocabulary contribution (14 classes, 57 properties, OWL/SHACL constraints) is presented as a proposal that can be improved; the appropriate action is for the authors to either narrow the claim or strengthen the evidence before the coverage claim is accepted. The concrete test I propose is feasible: the first part is an independent corpus extension that would take a few days with standard prompting; the second part is a two-annotator reliability check that can be done immediately on the existing data with a published rubric. I considered whether the 'context dominance' finding (89-93% replication of incorrect context, Section 5) could be the central concern, but that finding depends directly on the labels and is well within the label-reliability issue; the sufficiency claim is more load-bearing because it is the paper's stated validation of the entire research task. I also considered whether the small sample with wide confidence intervals (which the authors acknowledge, e.g., German incomplete CI of ±25 pp in Section 4.1) could be the main concern; it is a real limitation of the empirical generalization but not of the framework itself, while the sufficiency claim is a claim about the framework's coverage, which is the central contribution. Quote locations: Abstract 'demonstrating that our vocabulary was sufficient to express every assessment facet encountered in the 28-question study'; Section 4 'Validation assessed correctness per fire safety standards and context expectations'; Section 5 'All such findings are reflected in the vocabulary. Hence, such findings can be generated using SPARQL queries, which validates our research task'; Section 4.1 'built the 2x2 contingency table' and 'McNemar's exact test... Cohen's kappa'.","tokens_in":3938,"tokens_out":2863,"duration_ms":24846,"concrete_test":"Re-run the demonstration on an independent, adversarially sampled corpus of at least 50 questions spanning two new domains (e.g., medical dosage and legal citation), with two additional prompt templates and one additional model. If the authors or an independent annotator find any assessment facet that cannot be expressed in SQARE (e.g., a partial-credit answer, a temporally contingent fact, a citation-grade answer, or a multi-step reasoning trace), the sufficiency claim in Section 5 fails as stated. Additionally, to test labeling reliability, have a second annotator independently apply a written rubric (published in the repo) to the existing 28 questions and report Cohen's kappa; if kappa is below 0.8, the quantitative findings in Tables 1 and 2 and every SPARQL-queryable finding derived from is_valid are conditional on the original annotator's subjective labels.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central artifact claim (Abstract, Section 5) is that the SQARE vocabulary was sufficient to express every assessment facet encountered in the 28-question study, and that all reported findings are therefore reproducible via SPARQL. Two conditions must hold for this to be load-bearing: (1) the facet set actually encountered must be a meaningful test of vocabulary coverage, and (2) the labels and observations stored in RDF must be reliable. Condition (1) is weak because the corpus is a single domain (fire safety), a single prompt template (zero-shot, system-first per Section 4), 28 questions, two models, and two languages. The facets 'encountered' are generated by the authors' own pipeline, so a self-consistency loop is present: the vocabulary is declared sufficient for the data that the authors themselves constructed and annotated. No independent facet inventory, negative cases, or out-of-domain questions are reported. The strongest support for the sufficiency claim would be evidence that the vocabulary was derived before seeing the data, or a systematic enumeration of the facet space with a completeness argument, or a demonstration on an independent corpus. Condition (2) is also weak: Section 4 states validation 'assessed correctness per fire safety standards and context expectations,' but no rubric, no inter-annotator agreement, and no handling of ambiguous answers is reported. Since every quantitative finding (accuracy gaps, McNemar p-values, kappa values in Tables 1 and 2) is conditional on these binary labels, inconsistent labeling would change the stored observations and hence the SPARQL-queryable findings. The reader's weakest assumption captures this second condition; my concern adds that even with perfect labels, the sufficiency claim remains under-supported because the facet space is defined by the same authors who claim coverage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SQARE, an RDF-based vocabulary (14 classes, 57 properties) for representing multilingual LLM evaluation results, aligned with FAIR principles, PROV, and Dublin Core. The authors demonstrate the framework on a fire-safety domain study with 28 questions, four context conditions (complete, incomplete, conflicting, no-context), two models (GPT-4o-mini, Gemini-2.0-Flash), and two languages (German, English). Responses and manual validation labels are stored in RDF, and SPARQL queries are used to analyze context prioritization, knowledge leakage, and multilingual differences. The central claim is that the vocabulary was sufficient to express every assessment facet encountered in the 28-question study, and that all reported findings can be reproduced via SPARQL queries.","tokens_in":4121,"tokens_out":3642,"duration_ms":38269,"significance":"If the claims hold, the paper contributes a reusable, queryable, FAIR-aligned schema for LLM evaluation data, which would be a useful infrastructure resource for the NLP and semantic-web communities. The authors also publish the schema at a persistent URL and accompany the statistical comparisons with exact McNemar tests and Newcombe confidence intervals, which are appropriate tools for paired binary outcomes. However, the empirical demonstration is narrow (28 hand-selected questions from one domain, one prompt template, two models, two languages), the validation labels come from a single un-rubricated annotator, and the sufficiency claim is supported only by a self-consistency check on the authors' own data. The contribution is therefore best viewed as a promising but not yet fully validated resource.","major_comments":[{"comment":"The claim that the vocabulary was \"sufficient to express every assessment facet encountered in the 28-question study\" is not established by the reported demonstration. The 28 questions are all from the fire-safety domain, use a single zero-shot system-first prompt template, and cover only two models and two languages. The facets themselves are identified by the authors' own pipeline, so the statement in Section 5 that \"All such findings are reflected in the vocabulary\" is a self-consistency check rather than an independent validation of coverage. No systematic enumeration of the facet space, no adversarial examples, and no out-of-domain questions are provided. The claim should be weakened to \"covered all facets observed in this dataset,\" or supplemented with a completeness argument or a demonstration on an independent corpus.","section":"Abstract and Section 5"},{"comment":"All quantitative findings in Tables 1 and 2 rest on binary correctness labels described only as \"assessed correctness per fire safety standards and context expectations.\" No rubric is provided, no inter-annotator agreement is reported, and ambiguous answers are not discussed. Since the central empirical patterns (context-dominance rates of 89-93%, language differences, McNemar p-values, Cohen's kappa) are all conditional on these labels, the paper should supply a detailed labeling rubric, independent annotations with agreement statistics, or at minimum a sensitivity analysis treating the labels as uncertain. Without this, the empirical layer cannot support the strength of the conclusions drawn.","section":"Section 4, Data Collection and Analysis"},{"comment":"The statistical reporting overstates what the data can show. For contexts where the sum of discordant pairs b+c is below 5, the table marks the McNemar p-value as \"-\", yet the text states that \"McNemar's test is non-significant (p>0.05) in all German contexts\"; those cells do not provide evidence for non-significance because the test was not performed. Similarly, the only significant result (English no-context, p=0.0039) is based on 9 discordant pairs, and the Newcombe CI for the accuracy difference is wide ([-49.4, -14.8] percentage points). The text should explicitly distinguish \"not tested\" from \"not significant,\" and should present the confidence intervals as the primary evidence of effect size rather than relying on the dichotomous p-value.","section":"Section 4.1 and Tables 1-2"}],"minor_comments":[{"comment":"The heading \"T able 1\" contains a typographical error; it should read \"Table 1.\"","section":"Section 4"},{"comment":"The statement that models replicate incorrect information \"at rates of 89-93%\" is not directly traceable to a specific table row or calculation; please specify how these rates are derived from the reported contingency tables or provide the underlying query results.","section":"Section 4, Key Findings"},{"comment":"The phrase \"low agreement under manipulated prompts (e.g., kappa = 0.143 de, undefined in en)\" is imprecise: Table 2 shows kappa = 0 for English incomplete, while kappa is undefined for the English complete condition. Please clarify which cell is meant.","section":"Section 4.1 and Table 2"},{"comment":"The reference to the \"online appendix (Git repository)\" is not accompanied by a direct URL in the main text; the purl.org/sqare/repo# link in the footnote should be moved to a prominent location and verified for accessibility.","section":"Section 4, Experimental Setup"},{"comment":"The assertion that RDF provides \"advantages unattainable with flat CSV tables\" is plausible but not demonstrated; a concrete example showing a SPARQL query that is difficult or impossible to express over the equivalent CSV representation would strengthen the argument.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is probably better framed as a resource paper presenting a new vocabulary plus a small case study, rather than as a rigorous validation of either the vocabulary's completeness or the empirical findings. The authors should also consider whether the empirical claims about context prioritization and language differences are necessary to the contribution; if the vocabulary is the main artifact, the statistical layer could be substantially de-emphasized. The lack of a rubric and inter-annotator agreement for the correctness labels is a reproducibility concern that should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: this paper's contribution is the SQARE RDF vocabulary, and that part is worth engaging with. The empirical demonstration attached to it is thin and the abstract oversells it, but the artifact is real.\n\nWhat's actually new: the T-Box with 14 classes, 57 properties, OWL/SHACL constraints, aligned with PROV and Dublin Core, published at the PURL. The four cited prior works study knowledge conflicts but do not provide a standard representation. A machine-readable schema for storing LLM evaluation records, queryable with SPARQL, is a reasonable step toward FAIR-compliant sharing. The design is sensible: questions, answers, validation results, materials, language-tagged literals. I can see this being useful if the community picks it up.\n\nThe paper also does a few things well. The four-condition setup (complete, incomplete, conflicting, no context) is a clean way to probe context reliance. The authors report Newcombe confidence intervals and acknowledge when gaps are too wide to interpret. That is honest.\n\nThe soft spots are real and proportionate. The sufficiency claim in the abstract and Section 5—that the vocabulary expressed every facet encountered—is validated only against the authors' own 28-question, single-domain, two-model, two-language dataset. That is a self-consistency loop, not a completeness argument. Even with perfect labels, the facet space is defined by the same people who claim coverage. The validation labels are produced by the authors without a published rubric, inter-annotator agreement, or discussion of ambiguous answers. Every accuracy number, McNemar p-value, and kappa in Tables 1 and 2 is conditional on those labels. The abstract's 'critical patterns' language is too strong for a sample this size; the patterns are suggestive at best.\n\nNone of this kills the framework. The schema can stand on its own as a proposal, and the paper is clear about what it does and does not show. What it needs is a proper validation story: a rubric or reliability statistics, and either a more diverse corpus or a careful scope statement that says 'sufficient for these facets, not all facets.'\n\nIf I were the editor, I would send this to review. The vocabulary is a concrete, citable artifact, and the limitations are fixable. My recommendation: accept with revisions if the authors either ship the rubric and reliability data or soften the generality claims to match the evidence. The behavioral findings are a demonstration, not the main event.\n\nWorth a quick look if you care about FAIR evaluation sharing.","headline":"A real RDF vocabulary for LLM evaluation records, but the sufficiency claim is self-referential and the empirical support is thinner than the abstract suggests.","tokens_in":4829,"tokens_out":2269,"would_cite":true,"duration_ms":22317,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes an RDF vocabulary for LLM quality assessments and claims it captured every assessment facet in a 28-question German/English fire-safety study, with all headline findings recoverable by SPARQL queries.","keywords":["RDF","LLM evaluation","knowledge conflicts","multilingual assessment","SPARQL","FAIR data","context adherence","fire safety"],"falsifier":"Re-annotate the same 28 questions' model responses under a published rubric with at least two independent annotators; if labels change on enough items, the McNemar p-values, accuracy gaps, and kappa values in Tables 1 and 2 will shift or lose significance. Alternately, take any one finding stated in the paper and attempt to express it as a SPARQL query against the released RDF dataset; a single finding that cannot be expressed would refute the completeness claim.","tokens_in":3588,"feed_emoji":"🕸️","tokens_out":4430,"duration_ms":42277,"temperature":0.7,"pith_summary":"This paper argues that LLM quality assessments—especially evaluations of how models handle missing, conflicting, or absent context across languages—should be stored as RDF graphs instead of flat result tables. To that end it introduces an RDF vocabulary whose classes and properties record questions, answers, reference materials, and validation outcomes, and it demonstrates the vocabulary on 28 fire-safety questions answered by two LLMs in German and English under four context conditions. The central claim is that the vocabulary was sufficient to express every assessment facet the study produced, and that the study's findings about context adherence and language-specific performance can be regenerated by running SPARQL queries over the stored data. If that claim holds, evaluation datasets become queryable, interoperable, and reusable without bespoke scripts for each experiment.","feed_headline":"RDF vocabulary captured every facet of a 28-question LLM test","feed_subtitle":"Two models, four context conditions, two languages: the authors say SPARQL queries can recover every finding.","key_machinery":"The central object is the RDF T-Box for LLM evaluation: 14 classes such as :Question, :Answer, :ValidationResult, and :Material, plus 57 properties such as :hasGivenFor, :hasUsedMaterial, and :hasValidationResult, constrained with OWL/SHACL and aligned to PROV-O and Dublin Core for FAIR compliance. Multilingual content is stored as language-tagged literals, and binary correctness is attached to answers through :isValid, which is what the paired statistical comparisons use. This vocabulary turns each experimental response into a graph that can be queried with SPARQL, letting the authors recover context-adherence rates, language differences, and validation outcomes without custom analysis code.","core_discovery":"The paper's discovery is a demonstration of completeness: the proposed RDF schema, together with its OWL/SHACL constraints and FAIR alignments, is said to capture every assessment facet that arose in their 28-question, two-model, two-language study. Experimentally, the study found that both models reproduced incorrect provided context at high rates (89–93%) rather than falling back on training knowledge; that English responses handled incomplete information better while German responses showed stronger no-context baseline knowledge; and that in the only statistically testable English contrast, GPT-4o-mini outperformed Gemini-2.0-Flash by 32.1 percentage points in the no-context condition (McNemar p=0.0039). The authors take the fact that these findings can be produced through SPARQL queries as validation of the research task, namely representing LLM assessment data in a semantically rich, queryable form.","pith_inferences":["The paper does not test the vocabulary's sufficiency outside its own 28-question study; an obvious extension would be to audit the schema against an independent corpus of evaluation facets (e.g., from other benchmarks) and count how many require new properties.","Because the binary correctness labels are the statistical backbone, the numerical findings are conditional on a single manual annotation pass; adding a second annotator and measuring agreement would turn the framework into a more defensible evaluation standard.","The strong context-adherence result (89–93% replication of wrong context) suggests a concrete downstream test: a retrieval-augmented system that feeds unsanitized context could inherit these errors, so the queryable RDF store could be used to flag documents that trigger model over-reliance."],"forward_implications":["Any researcher who adopts the vocabulary can publish LLM evaluation results as RDF and let others reproduce the headline analyses via the same SPARQL queries, rather than re-analyzing raw outputs.","The four-condition protocol (complete, incomplete, conflicting, no-context) provides a reusable template for knowledge-conflict testing in other domains, with the RDF layer making cross-study comparison straightforward.","In practice, the measured context dominance implies that LLM responses in this domain should not be trusted to override explicitly provided wrong context; the graph representation makes such failure cases auditable.","If the sufficiency claim transfers to larger question sets, the vocabulary could serve as a standard target for dumping and comparing multilingual LLM evaluations, supporting the FAIR goals the paper emphasizes."],"supporting_citations":[{"why":"Supplies the factuality-assessment use case that the proposed RDF framework generalizes.","marker":"[1]"},{"why":"Documents the knowledge-graph and hallucination landscape the paper positions itself against.","marker":"[2]"},{"why":"Reports LLM bias toward provided context even when incorrect, the behavior the experiment quantifies.","marker":"[3]"},{"why":"Analyzes knowledge-conflict behavior the framework is designed to represent.","marker":"[4]"}],"fun_headline_variants":["RDF schema captures every facet of 28-question LLM eval","SPARQL queries recover all findings from 28-question LLM test","Multilingual LLM eval: RDF vocabulary proven complete","RDF framework covers all facets in two-language LLM study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every statistical result in the paper rests on the authors' manually assigned correct/incorrect labels, described only by circumstance as following fire safety standards and context expectations, with no published rubric and no measured agreement between annotators.","fun_headline_variants_meta":{"raw":{"variants":["RDF schema captures every facet of 28-question LLM eval","SPARQL queries recover all findings from 28-question LLM test","Multilingual LLM eval: RDF vocabulary proven complete","RDF framework covers all facets in two-language LLM study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1604,"prompt_tokens":846,"completion_tokens":758,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":683}},"tokens_in":462,"tokens_out":758,"duration_ms":7132,"temperature":1.0,"reasoning_tokens":683,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:58:30.646128+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the same 28 questions' model responses under a published rubric with at least two independent annotators; if labels change on enough items, the McNemar p-values, accuracy gaps, and kappa values in Tables 1 and 2 will shift or lose significance. Alternately, take any one finding stated in the paper and attempt to express it as a SPARQL query against the released RDF dataset; a single finding that cannot be expressed would refute the completeness claim.","supporting_citations":[{"cited_title":"In: International Workshop on the Semantic Web (2024), https://api.semanticscholar.org/CorpusID:274281581","cited_arxiv_id":null,"evidence_quote":"Supplies the factuality-assessment use case that the proposed RDF framework generalizes."}],"review_version":1}