{"id":"a8fc2b1f-bc85-422a-885b-d8d033be2af5","arxiv_id":"2608.05867","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors release the DISO defence/security ontology collection and a new OAEI track with eight matching tasks and a silver-standard reference alignment.","lead":"This paper introduces a collection of more than 60 public ontologies for defence and national security, and a new evaluation benchmark for software that aligns such ontologies. The benchmark includes eight matching tasks and a manually checked reference set of mappings between concepts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The silver standard is author-curated with no independent inter-annotator validation; until its precision is externally verified, OAEI rankings derived from this track are provisional.","rationale":"The reader's weakest assumption—that the manual validation of the silver standard is accurate and unbiased—is exactly the load-bearing point. I found no more central concern. The task-selection procedure is transparently tied to LogMap, but the paper acknowledges this and does not overclaim. The authors also wisely avoid reporting precision and recall against their own silver standard, and the ontology corpus and mapping pipeline are reproducible with released code and a Zenodo archive. The single gate for the paper's value as a benchmark is whether the silver-standard reference is trustworthy enough for third-party evaluation. The concern is not that the authors are careless; it is that the required evidence—independent annotation, inter-annotator agreement, and a clear rule for which unique mappings are included—is explicitly deferred to future work. Because the paper says this is planned but not done, the conditional verdict stands: accept the resource and its description, but keep the reference alignment provisional until external validation is provided.","tokens_in":18119,"tokens_out":3871,"duration_ms":39502,"concrete_test":"Run an independent double-annotation study before OAEI 2026 release: draw a stratified random sample of at least 50 mappings per task from the current silver standard and at least 50 from the set of rejected unique/consensus mappings; have two external annotators with defence/security ontology expertise, blinded to the mappings' provenance and to the authors' labels, classify each mapping as correct or incorrect. Compute Cohen's kappa and compare external labels to the authors' labels. If kappa is below 0.8, or if external labels change per-task silver precision by more than 5 percentage points, release the track as a provisional reference and report OAEI results against both the current and corrected silver standards.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a reusable benchmark whose value depends on the reliability of the silver-standard reference alignment. That reference is created in Section 3.2 by merging manually validated Con-2 consensus mappings and 'correct unique mappings' from the participating systems. The paper states that validation was 'performed primarily by the authors' and that external expert validation plus inter-annotator agreement analysis is only planned for OAEI 2026. This is load-bearing because every downstream comparison—OAEI participants, leaderboard, track results—computes precision and recall against this reference; incorrect silver mappings penalize correct systems, and missing mappings deflate recall. Moreover, several authors are co-developers of matcher families used to construct the consensus (e.g., LogMap, AML/Matcha), and no blinding protocol is reported, so the curation could systematically favor outputs of these systems. Table 6 shows the manual labels are not trivial: in JC3IEDM-Brick only 67% of Con-2 mappings were judged correct, while unique-mapping correctness varies from 0% to 97% across systems. The selection rule for 'a subset of unique mappings' merged into the silver standard is also unspecified; if it is not the full set of correct unique mappings, the reference is arbitrary in a way that affects rankings. None of this implies the resource is useless—the corpus, pipeline, and task selection are reproducible—but the benchmark's validity as a reference standard is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces DISO, a curated collection of 60+ publicly available ontologies relevant to the defence and national security domain, together with a new OAEI evaluation track consisting of eight matching tasks. The track is built by running several existing ontology alignment systems, aggregating their outputs into family-based consensus alignments, and manually curating a silver-standard reference alignment. The paper reports ontology statistics, consensus and per-system mapping counts, manual validation results, reasoning-based ontology compatibility analysis, and describes three released repositories (diso, diso-mappings, diso-oaei) that support reproducibility and community participation.","tokens_in":18328,"tokens_out":4392,"duration_ms":47530,"significance":"If the resource is valid, this is a valuable contribution: a reproducible, documented ontology collection and a benchmark for ontology matching in a domain where evaluation datasets are scarce. The paper has clear strengths: the corpus and pipeline are publicly released with permanent repositories, the collection criteria are explicit, the task-selection process is described, and the authors honestly acknowledge the silver-standard's limitations. The analysis of logical errors across systems and the silver standard also provides a useful preliminary view of the difficulty of the tasks. However, the central benchmark claim depends on the reliability of the manually constructed silver-standard, which has not yet been externally validated; this tempers the immediate trustworthiness of the track but does not undermine the corpus and pipeline contributions.","major_comments":[{"comment":"The silver-standard construction is underspecified. The text says the silver standard is created by merging the correct Con-2 mappings and the correct unique mappings, but the final counts do not match the sum of these two sets. For example, in UCO-STIX, 171 Con-2 mappings at 91% correctness (approximately 156) plus approximately 88 correct unique mappings (43*0.84 + 516*0.10) sum to roughly 244, yet the reported silver total is 194. The paper's earlier phrase 'a subset of the unique mappings' is never quantified or given a selection rule. Since the silver standard is the reference for all future participant evaluations, this missing rule makes the reference partly arbitrary and materially affects precision/recall scores. The authors should specify exactly how the subset was chosen, or provide the complete set of correct unique mappings and justify any exclusion.","section":"Section 3.2, Table 6"},{"comment":"The manual validation is described as 'performed primarily by the authors,' and the paper reports no inter-annotator agreement and no external expert validation. This is load-bearing because the silver standard is the ground truth of the track, and several of the evaluated systems are developed by the same research groups as the authors (e.g., LogMap, AML/Matcha). Without a blinding protocol or independent adjudication, the reference may systematically incorporate the biases of those matchers. The paper should either provide inter-annotator agreement statistics on a sample, include validation by independent domain experts, or clearly label the current reference as provisional and specify how bias will be controlled in the OAEI 2026 update.","section":"Section 3.2, Manual assessment"},{"comment":"The authors acknowledge that the silver standard is derived from the evaluated systems, so its recall is bounded by those systems and it will favor systems resembling them. They therefore choose not to report precision/recall values and instead show Jaccard distances. This is a reasonable caution, but the track's utility for future OAEI participants depends on a clearer quantification of this limitation. The paper should provide an explicit estimate of the reference's coverage (for instance, the union-of-systems upper bound, or the number of human-added mappings beyond any system output) and discuss how participants' precision/recall should be interpreted given this bounded recall.","section":"Section 3.2, Mapping comparison"},{"comment":"Table 7 reports that the silver standard leads to unsatisfiable classes in several tasks, yet the text states that 'for the DISO-OAEI 2026 track, the silver standard has been treated to fix logical errors and annotate incoherent mappings.' The paper does not report the corrected statistics or clarify which version is actually shipped in the repositories. This creates an inconsistency between the analyzed silver standard and the benchmark that participants will use. The authors should state clearly which version is released, describe the fixing procedure, and report the logical-error counts for the final released version.","section":"Section 3.2, Table 7 and Section 4"}],"minor_comments":[{"comment":"The caption contains a typo: 'ALOD2VEc' should be 'ALOD2Vec'.","section":"Table 6 caption"},{"comment":"Table 5 lists only six systems (AML, BertMap, BertMapLt, LogMap, LogMapLt, Matcha), while the text states that LogMapLLM, ALOD2Vec, ATMatcher, Fine-TOM, and KGMatcher were also executed but were unable to produce mappings on all tasks. To make the comparison complete, the table should include a row or column indicating which systems failed on which tasks, or an explicit note about their omission.","section":"Table 5"},{"comment":"The axes are labelled MDS-1 and MDS-2, but the text never defines MDS; please add a sentence explaining that these are the first two multidimensional scaling dimensions of the Jaccard distance matrix.","section":"Figure 2"},{"comment":"Table 1 lists the cluster 'mid-level' with size 12 but does not show the subcluster 'cco-modules', which is used in Table 3 and in the description of the Facility ontology. Please add the missing subcluster to Table 1 or note explicitly that 'cco-modules' is a subcluster within 'mid-level'.","section":"Table 1 and Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely a good fit for a resource or evaluation track venue, and the authors should be credited for shipping reproducible repositories and for explicitly acknowledging the silver-standard's provenance. The main risk is that the benchmark's validity depends on a reference whose construction and validation are not yet fully specified; this is fixable within the manuscript's scope. I would not recommend rejection, because the corpus and pipeline stand on their own. However, the authors must address the selection rule for unique mappings and provide some form of external or inter-annotator validation evidence before the silver standard can be treated as a reference for participant rankings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the DISO collection and eight matching tasks fill a real gap, and the paper deserves a serious referee. The main weakness, the author-curated silver standard, is real and the authors mostly own up to it in the text.\n\nWhat's genuinely new: a curated set of 63 defence/security ontologies, a bulk alignment analysis over 1,653 pairs, and a new OAEI track with eight tasks that mix classes, properties, and instances. The three repositories are shipped with code, data, and a reproducible make workflow, which is exactly the kind of thing that should be credited. The family-based consensus voting is a sensible attempt to reduce bias from multiple variants of the same matcher. The reasoning checks (Table 7) are a nice addition: they show real unsatifiability problems in several tasks, including in the silver standard itself, which is useful information for the OAEI community.\n\nSoft spots, in proportion. The silver standard is constructed from consensus mappings and a subset of unique mappings, manually validated primarily by the authors. There is no inter-annotator agreement yet, and no external validation. That matters, because this reference is what future OAEI participants will be measured against. But the authors explicitly call it a partial reference, say recall is bounded by what the systems can collectively discover, and deliberately avoid reporting precision/recall of the participating systems against it. That is the right call and it takes some of the edge off the circularity concern.\n\nTwo smaller issues. The selection rule for the subset of unique mappings merged into the silver standard is not specified precisely. Table 6 shows big variation in unique-mapping correctness across systems and tasks, so the subset choice can affect the reference. Also, since several authors are co-developers of matcher families used in the consensus, a blinded validation protocol would have been stronger. Neither is fatal, but they should be addressed in the review.\n\nWho this is for: ontology matching researchers, especially OAEI participants, and people working on defence knowledge graph interoperability. The track itself is the contribution, not a new matching algorithm. My recommendation: send to peer review. Ask the authors to document the unique-mapping selection rule and to commit to external expert validation and IAA as part of the OAEI 2026 campaign. With those, the benchmark becomes much more solid.","headline":"A useful, reproducible new OAEI track with an honestly-labeled provisional silver standard; the resource is worth a serious referee, but don't treat the reference alignment as settled yet.","tokens_in":18911,"tokens_out":1661,"would_cite":true,"duration_ms":20425,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper assembles a curated collection of 60+ defence and security ontologies and a new eight-task evaluation track for aligning them.","keywords":["Ontology Matching","Ontology Alignment","Defence and National Security","Knowledge Graphs","Evaluation","OAEI","Silver Standard","Consensus Alignment"],"falsifier":"Take the STIX-D3FEND task, have a panel of defence and cybersecurity experts who did not participate in the paper independently classify each silver-standard mapping as correct or incorrect, and compare: if a substantial number of silver mappings is rejected, or if the experts identify many obviously correct mappings the systems all missed, the track's reference alignments cannot support reliable evaluation.","tokens_in":17895,"feed_emoji":"🛡️","tokens_out":9284,"duration_ms":84817,"temperature":0.7,"pith_summary":"The paper aims to make defence and national security ontologies interoperable by giving researchers the corpus and the measuring stick to align them. It assembles DISO, a documented collection of 63 publicly available OWL ontologies covering cybersecurity, situation awareness, smart environments, and other subdomains, and it runs a bulk alignment exercise over 1,653 ontology pairs to find where these ontologies actually intersect. From that analysis it selects eight non-trivial matching tasks and packages them as a new track in the Ontology Alignment Evaluation Initiative, complete with consensus alignments produced by several state-of-the-art systems and a manually verified silver-standard reference alignment. A sympathetic reader would care because, if the resource is sound, alignment systems can finally be benchmarked on a realistic, open, defence-specific test bed instead of only on generic or biomedical ontologies.","feed_headline":"A new benchmark maps 60+ defence and security ontologies","feed_subtitle":"A curated 63-ontology corpus and a silver-standard alignment let any team test matching systems on defence tasks.","key_machinery":"The load-bearing artifact is the DISO network of ontologies, a curated collection of 63 publicly available OWL ontologies organised into 11 clusters. The argument runs through a three-stage pipeline: a bulk LogMap alignment over 1,653 ontology pairs identifies candidate matching tasks; family-based voting aggregates the outputs of eight system families into consensus alignments, with each family counted once so that systems entering multiple variants cannot dominate; and manual curation of the vote-2 consensus plus mappings unique to a single family produces the silver-standard reference alignment. The family-based voting is what makes the consensus reproducible and less biased, and the manual step is what turns agreement into an asserted ground truth.","core_discovery":"On its own terms, the paper establishes that the defence and national security domain already has a large, loosely connected ecosystem of public ontologies, and that this ecosystem can be turned into a reproducible evaluation setting. LogMap found at least one mapping in 647 of 1,653 ontology pairs, showing substantial overlap despite very different modelling styles. The eight selected tasks split into three subdomains—cybersecurity (UCO-STIX, STIX-D3FEND), situation awareness (JC3IEDM against mIO!, Brick, and Facility), and smart environments (ThinkHome-Brick, Brick-SmartEnv, CityOWL-Brick)—and system outputs on them differ sharply: for JC3IEDM-mIO!, LogMap produced 234 mappings while AML produced 34. The authors therefore built a family-based consensus alignment and manually merged it with correct system-unique mappings to form a silver standard for each task. They also show that merging these ontologies with either system or silver-standard mappings makes many classes unsatisfiable, meaning no individual can belong to them under the combined axioms, so the track includes a genuine logical-reasoning challenge, and the silver standard has been treated to fix such errors for the 2026 campaign.","pith_inferences":["A natural next step the authors do not spell out is to score systems not just on precision and recall against the silver standard but also on how few unsatisfiable classes their mappings introduce; the paper's own coherence data make such a combined metric feasible.","If the DISO collection grows the way biomedical repositories have, the track could become the empirical basis for deciding which upper-level or mid-level ontologies should anchor a shared defence and security ontology stack.","The sharp disagreement among systems on tasks like JC3IEDM-mIO! suggests that label-based matching is insufficient for many defence mappings; testing whether structure-aware or instance-aware matchers close that gap would be a direct use of the released pipeline.","Re-running the same consensus-plus-validation protocol on ontologies whose labels are not in English would reveal whether the observed system disagreements are a property of the defence domain or an artefact of the systems' English-centric features."],"forward_implications":["OAEI participants can benchmark their alignment systems against eight defence and security tasks spanning three subdomains, with a silver standard to score against.","The DISO repository and the diso-mappings pipeline allow anyone to reproduce the consensus alignments and extend the collection with new ontologies or new matchers.","The coherence results show that even strong alignment systems create unsatisfiable classes when their mappings are merged with the source ontologies, so future systems in this domain should treat logical consistency as part of the task.","Because the silver standard is explicitly partial, its recall is bounded by what the participating systems can discover, so it should be used as a lower-bound reference rather than a complete gold standard.","The planned leaderboard and candidate-ranking task will let machine-learning-based matchers enter without requiring them to output full alignments."],"supporting_citations":[{"why":"LogMap is the alignment system used for the bulk analysis over 1,653 ontology pairs and for selecting the eight track tasks.","marker":"[27]"},{"why":"The OAEI results paper defines the annual evaluation campaign that the new track joins.","marker":"[49]"},{"why":"AML is one of the state-of-the-art systems whose outputs are aggregated into the consensus alignments.","marker":"[15]"},{"why":"BERTMap and its variants form a system family whose mappings contribute to the consensus and validation.","marker":"[21]"},{"why":"Matcha produces large mapping sets that dominate the unique-mapping validation in several tasks.","marker":"[16]"},{"why":"This prior OAEI disease and phenotype track motivates manual validation by showing that consensus alignments can contain errors.","marker":"[20]"},{"why":"The GUARD project's phase 1 report supplies the detailed analysis behind the choice of the eight ontology pairs.","marker":"[26]"},{"why":"JC3IEDM is the command-and-control information exchange data model that appears in three of the eight matching tasks.","marker":"[43]"},{"why":"D3FEND is the cybersecurity countermeasures knowledge graph used as the target in the STIX-D3FEND task.","marker":"[30]"}],"fun_headline_variants":["New OAEI track tests alignment on 63 defence ontologies","Defence ontology alignment: 8 tasks, silver-standard mappings","Curated benchmark for defence and security ontology matching","63 defence ontologies mapped in new evaluation track","Silver-standard alignments for 60+ defence ontologies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark rests on the authors' own manual judgments that the silver-standard mappings are correct; if those judgments are wrong or biased, every score computed against the track inherits the error.","fun_headline_variants_meta":{"raw":{"variants":["New OAEI track tests alignment on 63 defence ontologies","Defence ontology alignment: 8 tasks, silver-standard mappings","Curated benchmark for defence and security ontology matching","63 defence ontologies mapped in new evaluation track","Silver-standard alignments for 60+ defence ontologies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000322,"raw_usage":{"total_tokens":1805,"prompt_tokens":931,"completion_tokens":874,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":794}},"tokens_in":547,"tokens_out":874,"duration_ms":8341,"temperature":1.0,"reasoning_tokens":794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T21:59:54.597177+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the STIX-D3FEND task, have a panel of defence and cybersecurity experts who did not participate in the paper independently classify each silver-standard mapping as correct or incorrect, and compare: if a substantial number of silver mappings is rejected, or if the experts identify many obviously correct mappings the systems all missed, the track's reference alignments cannot support reliable evaluation.","supporting_citations":[{"cited_title":"In: The Semantic Web - ISWC 2011 - 10th International Semantic Web Conference","cited_arxiv_id":null,"evidence_quote":"LogMap is the alignment system used for the bulk analysis over 1,653 ontology pairs and for selecting the eight track tasks."},{"cited_title":"In: 20th International Workshop on Ontology Matching (OM)","cited_arxiv_id":null,"evidence_quote":"The OAEI results paper defines the annual evaluation campaign that the new track joins."},{"cited_title":"Semantic Web16(2), SW–233304 (2025)","cited_arxiv_id":null,"evidence_quote":"AML is one of the state-of-the-art systems whose outputs are aggregated into the consensus alignments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Matcha produces large mapping sets that dominate the unique-mapping validation in several tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This prior OAEI disease and phenotype track motivates manual validation by showing that consensus alignments can contain errors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The GUARD project's phase 1 report supplies the detailed analysis behind the choice of the eight ontology pairs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"JC3IEDM is the command-and-control information exchange data model that appears in three of the eight matching tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"D3FEND is the cybersecurity countermeasures knowledge graph used as the target in the STIX-D3FEND task."}],"review_version":1}