REVIEW 3 major objections 2 minor 29 references
When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking
T0 review · 3 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read Z-score aggregation resolves conflicting rank metrics by ranking DualE highest for tail prediction and FMS highest for relation prediction in KGC models.
desk verdict Z-score emerges as the balanced aggregator for conflicting KGC metrics after five tests and removals, but the abstract supplies no equations or data details to verify the Pareto claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Pareto-optimal analysis applied to seven metric aggregators evaluated on five tests (consistency, stability, independence, robustness, generalizability) using LOMO and LOGO cross-validation procedures.
What would settle it
A new KGC dataset or model set in which another aggregator produces rankings that match held-out performance better than Z-score rankings do.
Extended reading notes
Core claim
By treating KGC benchmarking as a multi-criteria decision-making task and evaluating seven aggregators through the five tests with LOMO and LOGO removals, the meta-analysis finds Z-score to be the aggregator that best balances the criteria and produces stable model rankings, with DualE leading tail prediction and FMS leading relation prediction.
Load-bearing premise
The five tests combined with LOMO and LOGO removals are sufficient and unbiased for judging which aggregator produces reliable rankings across the selected models and datasets.
Editorial extensions
If this is right
- Model comparisons become consistent across metrics and datasets instead of depending on which metric is reported.
- Selective reporting of favorable metrics is reduced because one aggregator is shown to dominate on the reliability tests.
- Researchers gain evidence-based guidance for choosing an aggregator rather than defaulting to single metrics.
- Test-sensitivity results indicate that consistency and stability tests remain stable under model removals while generalizability and independence vary most.
- The same framework can be reused to compare future aggregators or new KGC models.
Reading between the lines
- The approach could transfer to other machine-learning domains where multiple metrics produce conflicting leaderboards.
- If Z-score rankings better predict downstream task success, they could replace ad-hoc metric selection in papers.
- Extending the tests to include runtime or memory cost would address whether the top-ranked models remain practical.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reframes KGC model evaluation as an MCDM problem and conducts a meta-analysis of seven aggregators evaluated on five tests (consistency, cross-dataset stability, metric independence, robustness under noise, generalizability). Tests are averaged over LOMO and LOGO removals; Pareto analysis identifies Z-score as most balanced, ranking DualE highest on tail prediction and FMS highest on relation prediction. A sensitivity analysis indicates consistency and stability are removal-invariant while generalizability and independence are more sensitive.
Significance. If the empirical results hold, the work supplies an evidence-based procedure for choosing metric aggregators in KGC benchmarking, directly addressing the documented problem of conflicting rank-based metrics (MRR, Hits@k, MR) across datasets and thereby reducing opportunities for selective reporting.
major comments (3)
- [§4] §4 (Tests and removals): The five tests plus LOMO/LOGO averaging are asserted to be jointly sufficient for ranking aggregators, yet no operational definitions, independence checks among the five dimensions, or power analysis for rank differences are supplied; without these the Pareto front can be sensitive to test selection.
- [§3.2] §3.2 (Data collection): The model and dataset collection used both to instantiate the five tests and to produce the final aggregator rankings creates circular dependence; the manuscript does not report an external hold-out collection or pre-registered protocol that would break this dependence.
- [§5.1] §5.1 (Pareto analysis): The claim that Z-score is Pareto-optimal rests on averaged scores whose variance across LOMO/LOGO folds is not reported; without fold-level dispersion or statistical tests it is unclear whether the reported dominance is robust or an artifact of the particular removal scheme.
minor comments (2)
- The abstract states conclusions but supplies neither the list of seven aggregators nor the concrete datasets; these should appear in the abstract or a prominent table.
- Notation for the five tests is introduced without a compact summary table that would allow readers to see at a glance which test addresses which failure mode of isolated metrics.
Simulated Author's Rebuttal
We thank the referee for the constructive comments, which identify important areas for clarification and strengthening. We respond point-by-point to the major comments below.
read point-by-point responses
-
Referee: [§4] §4 (Tests and removals): The five tests plus LOMO/LOGO averaging are asserted to be jointly sufficient for ranking aggregators, yet no operational definitions, independence checks among the five dimensions, or power analysis for rank differences are supplied; without these the Pareto front can be sensitive to test selection.
Authors: Operational definitions for each of the five tests appear in the subsections of §4 and are grounded in MCDM literature. We acknowledge that explicit pairwise independence checks among the dimensions and a formal power analysis for rank differences are not provided. In revision we will add a dedicated paragraph in §4 discussing test interdependencies, the rationale for joint sufficiency, and the practical constraints on power analysis given the model collection size. The LOMO/LOGO averaging already supplies empirical robustness across model subsets; we will emphasize this as partial mitigation of sensitivity to test selection. revision: partial
-
Referee: [§3.2] §3.2 (Data collection): The model and dataset collection used both to instantiate the five tests and to produce the final aggregator rankings creates circular dependence; the manuscript does not report an external hold-out collection or pre-registered protocol that would break this dependence.
Authors: The concern about circular dependence is valid. While LOMO and LOGO removals evaluate aggregator behavior on diverse subsets, they do not fully eliminate dependence on the original collection. We did not employ an external hold-out set or pre-register the protocol. In the revised manuscript we will add an explicit limitations paragraph acknowledging this issue and recommending that future work use hold-out collections or pre-registration to further validate aggregator rankings. revision: yes
-
Referee: [§5.1] §5.1 (Pareto analysis): The claim that Z-score is Pareto-optimal rests on averaged scores whose variance across LOMO/LOGO folds is not reported; without fold-level dispersion or statistical tests it is unclear whether the reported dominance is robust or an artifact of the particular removal scheme.
Authors: We agree that the absence of fold-level variance and statistical tests weakens the robustness claim for the Pareto front. Although the sensitivity analysis examines removal invariance for individual tests, we did not report dispersion across folds for the aggregator scores or apply statistical comparisons. In revision we will augment §5.1 with standard deviations across LOMO/LOGO folds and non-parametric tests comparing aggregator performances to substantiate that Z-score dominance is not an artifact of the removal scheme. revision: yes
Circularity Check
No significant circularity detected
full rationale
The paper defines five explicit tests (consistency, cross-dataset stability, metric independence, robustness under noise, generalizability) and applies them via LOMO/LOGO removals to rank seven aggregators on a fixed collection of KGC models and datasets, then selects Z-score via Pareto analysis. This is an empirical comparison of aggregator behavior against independent criteria; the output ranking is not equivalent to the input definitions or data by construction, nor does any quoted step reduce via self-citation, fitted parameter, or ansatz smuggling. The derivation remains self-contained against external benchmarks of aggregator quality.
Assumptions & free parameters
assumptions (1)
- domain assumption The five tests plus LOMO/LOGO removals are sufficient and unbiased measures of aggregator quality
Cite this review
Pith. "Pith review of When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking." pith.science (2026). https://pith.science/paper/26RZWJJK
@misc{pith2026260610287,
author = {Pith},
title = {Pith review of: When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking},
year = {2026},
howpublished = {\url{https://pith.science/paper/26RZWJJK}},
note = {Machine review of arXiv:2606.10287}
}
abstract
Evaluating Knowledge Graph Completion (KGC) models remains challenging because standard assessment relies on isolated rank-based metrics such as MRR, Hits$@$k, and Mean Rank, which often produce conflicting model orderings across datasets. A model that leads on MRR may trail on Hits@1, and strong performance on one dataset may not generalize to another. This fragmentation hinders comparison, enables selective reporting, and obscures real progress. We reframe KGC evaluation as a Multi-Criteria Decision-Making (MCDM) problem and present a meta-analysis of seven aggregators across five tests: consistency, cross-dataset stability, metric independence, robustness under noise, and generalizability. Each test is averaged over leave-one-model-out (LOMO) and leave-one-group-out (LOGO) removals so that reliability reflects aggregator behavior across diverse model subsets. Across tail $(h,r,?)$ and relation $(h,?,t)$ prediction, Pareto-optimal analysis identifies Z-score as the most balanced aggregator, which ranks DualE highest for tail prediction and FMS (Flow-Modulated Scoring) highest for relation prediction. A test-sensitivity analysis using the same removals shows that consistency and stability are largely removal-invariant, while generalizability and independence are the most sensitive. The framework resolves evaluation inconsistencies and offers evidence-based guidance for aggregator selection and model benchmarking in KGC.
Figures
Figures from the paper (25 more)
Reference graph
Works this paper leans on
-
[1]
Z. Sun, Z. H. Deng, J.-Y. Nie, J. Tang, Rotate: Knowledge graph embedding by relational rotation in complex space, in: ICLR, 2019
2019
-
[2]
Y. Wang, S. Broscheit, R. Gemulla, A relational tucker decomposition for multi-relational link prediction, arXiv preprint arXiv:1902.00898 (2019)
work page Pith review arXiv 1902
-
[3]
L. Yao, C. Mao, Y. Luo, Kg-bert: Bert for knowledge graph comple- tion, in: Proc. of AAAI, 2019
2019
-
[4]
Y. Wei, Q. Huang, Y. Zhang, J. Kwok, Kicgpt: Large language model with knowledge in context for kgc, in: Proc. of EMNLP, 2023
2023
-
[5]
H. Gul, A. G. Naim, A. A. Bhat, Muco-kgc: Multi-context-aware knowledge graph completion, in: Pacific-Asia Conference on Knowl- edge Discovery and Data Mining, Springer, 2025, pp. 3–15
2025
-
[6]
H. Li, B. Yu, Y. Wei, K. Wang, R. Y. Da Xu, B. Wang, Kermit: Knowl- edge graph completion of enhanced relation modeling with inverse transformation, Knowledge-Based Systems 324 (2025) 113500
2025
-
[7]
Rossi, D
A. Rossi, D. Barbosa, D. Firmani, A. Matinata, P. Merialdo, Knowl- edge graph embedding for link prediction: A comparative analysis, TKDD (2021)
2021
-
[8]
Toutanova, D
K. Toutanova, D. Chen, Observed versus latent features for knowledge base and text inference, in: Proceedings of the 3rd workshop on continuous vector space models and their compositionality, 2015, pp. 57–66
2015
Show all 29 references
-
[9]
Kadlec, O
R. Kadlec, O. Bajgar, J. Kleindienst, Knowledge base completion: Baselines strike back, in: Proceedings of the 2nd Workshop on Representation Learning for NLP, 2017, pp. 69–74
2017
-
[10]
Ruffinelli, S
D. Ruffinelli, S. Broscheit, R. Gemulla, You can teach an old dog new tricks! on training knowledge graph embeddings (2020)
2020
-
[11]
Z. Sun, Q. Zhang, W. Hu, C. Wang, M. Chen, F. Akrami, C. Li, A benchmarking study of embedding-based entity alignment for knowl- edge graphs, arXiv preprint arXiv:2003.07743 (2020)
2003
-
[12]
A. Rao, N. A. Krishnan, C. R. Rivero, Using model calibration to evaluate link prediction in knowledge graphs, in: Proceedings of the ACM Web Conference 2024 (WWW ’24), Association for Computing Machinery, New York, NY, USA, 2024, pp. 2042–2051. doi:10.1145/3589334.3645506
2024 doi
-
[13]
M. K. Egger, W. Ma, D. Mottin, P. Karras, I. Bordino, F. Gullo, A. Anagnostopoulos, Relik: A reliability measure for knowledge graph embeddings, in: Proceedings of the ACM Web Conference 2024 (WWW ’24), Association for Computing Machinery, New York, NY, USA, 2024. doi:10.1145/...
2024 doi
-
[14]
H. Gul, A. G. Naim, A. A. Bhat, Kg-edas: A meta-metric framework for evaluating knowledge graph completion models, arXiv preprint arXiv:2508.15357 (2025). H. Gul et al.: Preprint submitted to Elsevier Page 18 of 26 When Metrics Disagree: A Meta-Analysis of KGC Model Benchmarking
2025
-
[15]
J. Xu, S. Zhang, H. Xie, H. Zhang, K. Miao, Q. Fu, Knowledge graph completion based on a hierarchical graph attention network with structural information, Knowledge-Based Systems (2025) 115164
2025
-
[16]
Gyani, A
J. Gyani, A. Ahmed, M. A. Haq, Mcdm and various prioritization methods in ahp for css: A comprehensive review, IEEE Access 10 (2022) 33492–33511
2022
-
[17]
E. K. Zavadskas, V . Podvezko, Integrated determination of objective criteria weights in mcdm, International Journal of Information Technology & Decision Making 15 (2016) 267–283
2016
-
[18]
Z. Sun, S. Vashishth, S. Sanyal, P. Talukdar, Y. Yang, A re- evaluation of knowledge graph completion methods, in: D. Jurafsky, J. Chai, N. Schluter, J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for C...
2020
-
[19]
URL: https://aclanthology.org/2020.acl-main.489/. doi: 10. 18653/v1/2020.acl-main.489
2020
-
[20]
Pezeshkpour, Y
P. Pezeshkpour, Y. Tian, S. Singh, Revisiting evaluation of knowledge base completion models, arXiv preprint arXiv:2002.00967 (2020)
2002
-
[21]
Mohamed, S
A. Mohamed, S. Parambath, Z. Kaoudi, A. Aboulnaga, Popularity agnostic evaluation of knowledge graph embeddings, in: J. Peters, D. Sontag (Eds.), Proceedings of the 36th Conference on Uncer- tainty in Artificial Intelligence (UAI), volume 124 of Proceedings of Machine Learning ...
2020
-
[22]
S. Moon, Y. Ko, How sharp and bias-robust is a model? dual evalua- tion perspectives on knowledge graph completion, in: Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining, WSDM ’26, Association for Computing Machinery, New York, NY, USA, 2...
2026 doi
-
[23]
Habiburahman, K
M. Habiburahman, K. R. S. Wiharja, M. Fikriansyah, Beyond benchmarks: Assessing knowledge graph completion methods on non-benchmark employee data, in: 2024 International Conference on Data Science and Its Applications (ICODSA), volume 11, IEEE, 2024, pp. 28–33. doi: 10.1109/ic...
2024 doi
-
[24]
H. Peng, L. Yao, S. Yang, Z. Zhang, J. Shi, X. Liao, H. Xiong, Causal inference-based knowledge graph completion using multi- head attention for fault diagnosis, Knowledge-Based Systems (2026) 116031
2026
-
[25]
Z. Cao, C. Luo, Link prediction for knowledge graphs based on extended relational graph attention networks, Expert Syst. Appl. 259 (2025)
2025
-
[26]
Z. Cao, Q. Xu, Z. Yang, X. Cao, Q. Huang, Dual quaternion knowledge graph embeddings, Proceedings of the AAAI Conference on Artificial Intelligence 35 (2021) 6894–6902
2021
-
[27]
Ferrari, G
I. Ferrari, G. Frisoni, P. Italiani, G. Moro, C. Sartori, Comprehensive analysis of knowledge graph embedding techniques benchmarked on link prediction, Electronics 11 (2022)
2022
-
[28]
Rossi, D
A. Rossi, D. Barbosa, D. Firmani, A. Matinata, P. Merialdo, Knowl- edge graph embedding for link prediction: A comparative analysis, ACM Trans. Knowl. Discov. Data 15 (2021)
2021
-
[29]
Zhang, J
Z. Zhang, J. Cai, Y. Zhang, J. Wang, Learning hierarchy-aware knowledge graph embeddings for link prediction, in: Proc. of AAAI, 2020. A. Base Results This supplementary document provides the complete baseline and ablation results for the multi-criteria decision- making (MCDM)...
2020
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.