REVIEW 5 major objections 7 minor 44 references
ClarAVy: A Tool for Scalable and Accurate Malware Family Labeling
T0 review · 5 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read ClarAVy beats the prior best malware family labeler by 8-12 points using a sparse Bayesian aggregator, and scales to 40 million scan reports.
desk verdict Solid applied ML paper: SparseIBCC gives a real accuracy gain for AV-based malware family labeling, but the headline numbers are coverage-dependent and the runtime section needs correction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
SparseIBCC — the paper's sparse implementation of the Independent Bayesian Classifier Combination (VB-IBCC) algorithm — is the load-bearing component. VB-IBCC approximates the Dawid-Skene EM iteration with variational Bayes; SparseIBCC keeps the same statistical model but represents each antivirus product's confusion matrix only over families that co-occur, and it computes each file's class posterior only over families present in that file's scan report. This is what turns a 50,000-class, 40-million-file inference problem into something that runs in about an hour on a 128-thread machine, making the reported accuracy and scale possible.
What would settle it
Take the subset of MOTIF files whose true family is absent from every antivirus detection — the 18.64% the paper identifies — and measure ClarAVy's accuracy on that subset against ground truth: it would necessarily be zero, because the tool can only output a family that appeared in the scan. A simpler test is to construct a synthetic scan report with the true family deliberately withheld from all vendors and confirm that ClarAVy never outputs it, showing that the accuracy ceiling is set by antivirus coverage rather than by the aggregation method.
Extended reading notes
Core claim
The central discovery, on the paper's own terms, is that extreme-multiclass crowdsourcing models can be made to work for malware at web scale by exploiting two sparsity patterns. First, nearly all pairs of malware families never co-occur in scan reports, so the per-antivirus confusion matrices of a variational Bayesian Dawid-Skene model can be stored sparsely at about 1/200 of the dense memory. Second, the posterior for a file can be restricted to the families that actually appear in that file's scan report, shrinking an O(NL) problem to O(NK) and cutting memory by a further factor of 500-5000. Together with manually crafted rules for parsing 103 antivirus products and for resolving trivial, sibling, and parent-child aliases, these changes let ClarAVy label 39,747,485 scan reports in about 37 hours while beating the prior leading tool by 8 points on MOTIF and 12 on MalPedia.
Load-bearing premise
The method assumes that a file's true family is always one of the family names that appears in that file's antivirus scan report, and the paper reports that 18.64% of MOTIF files violate this assumption.
Editorial extensions
If this is right
- Security analysts can get family labels for tens of millions of files in about a day of compute, making whole-corpus labeling practical rather than a research luxury.
- Downstream malware classifiers trained on ClarAVy labels should see better label quality than those trained on prior aggregation tools, since ClarAVy beats the leading baseline by 8-12 percentage points on two benchmarks.
- The confidence score gives a precision lever: on the MOTIF set, keeping only scans with confidence at or above 70% yields roughly 90% label accuracy.
- Antivirus coverage, not aggregation method, sets the ceiling: 18.64% of MOTIF scans cannot be correctly labeled from AV reports alone no matter how the votes are combined.
- Most remaining mislabels are near misses (variants of the true family or catch-all names used by vendors), so improving vendors' family vocabularies would do more than further algorithmic tuning.
Reading between the lines
- A natural extension the paper does not test: the sparse-confusion representation should transfer to other extreme-multiclass crowdsourcing tasks, such as fine-grained image annotation or labeling rare scientific entities, wherever class co-occurrence is rare.
- One testable prediction from the paper's own case studies is that richer family vocabularies from antivirus vendors would raise accuracy more than any further refinement of the aggregation algorithm.
- A hybrid that adds static or dynamic malware analysis for files whose true family never appears in the scan report could push past the AV-coverage ceiling; the paper stops at AV-only labeling.
- The 52,371-family taxonomy and 4,472 alias pairs generated from 40 million reports are a reusable asset for the malware research community, but the paper does not analyze the structure of that alias graph.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. ClarAVy is an antivirus-detection-based malware family labeling tool. The paper describes three components: improved AV detection parsing, a two-tier alias resolution scheme (trivial, sibling, and parent-child aliases), and SparseIBCC, a sparse variational Bayesian implementation of the IBCC/Dawid-Skene label aggregation model, plus a confidence-score model for practical use. The evaluation compares ClarAVy with five prior tools on MOTIF and MalPedia, using externally derived ground-truth labels, and reports 64.16% and 56.88% family-labeling accuracy, respectively, which is 8 and 12 percentage points above the best prior tool. The paper also reports scaling to 39,747,485 VirusTotal reports and analyzes runtime as a function of dataset size.
Significance. The central comparison is meaningful because the MOTIF and MalPedia labels are not AV-derived, so the reported gains are not circular. The paper is also transparent about the main limitation: Section 6.4 quantifies how often AV detections lack or nearly lack the true family. The SparseIBCC memory analysis is a concrete contribution that makes large-scale aggregation feasible, and the authors state that the tool will be released. If the reported numbers hold, ClarAVy would be a new practical baseline for AV-based family labeling. However, the headline accuracy is an upper bound conditioned on AV family coverage, the runtime-scaling evidence in Section 6.3 is internally inconsistent, and the confidence-score evaluation needs clarification before the practical claims are fully supported.
major comments (5)
- [§5.5 and §6.4] The assumption t_i ∈ c_i in Section 5.5 restricts the posterior to families appearing in a file's AV scan report, and Section 6.4 reports that 18.64% of MOTIF scans contain no detection with the true family and another 11.53% contain it only once. This means the reported 64.16% and 56.88% accuracies are coverage-dependent ceilings rather than properties of the aggregation method alone: on corpora with sparser AV family coverage the achievable accuracy is lower by construction, and the comparison with prior tools is partially determined by how often the coverage assumption holds. Please report accuracy both on the full test sets and on the subset of scans for which the true family appears at least once, and discuss how the ceiling would shift under lower coverage.
- [§6.3 and Figure 7] The text states that Figure 7 includes a quadratic curve of best fit and then concludes that SparseIBCC's runtime complexity is 'sub-linear in practice.' A quadratic fit is not evidence of sub-linear scaling; at best it is a local fit whose shape depends on the range, and at worst it contradicts the claim. Please report the fitted functional form with its parameters and goodness of fit, or measure scaling explicitly (e.g., wall-clock time per iteration or per report as N grows), and reconcile the statement with the empirical curve.
- [§5.6 and §6.2] The confidence-score model in Section 5.6 is trained on the combined MOTIF and MalPedia sets with five-fold cross-validation, but Section 6.2 does not state whether Figures 5 and 6 evaluate accuracy on held-out cross-validation folds or on the training data. If the latter, the high accuracy at confidence thresholds is optimistic. Please state explicitly which predictions were used and, if necessary, report the held-out curves.
- [§6.1 and Sections 4.2–4.3, 5.2] The evaluation uses default values for a large number of manually selected parameters (S, T, E, C, M, plus the 934 hand-written parsing rules and curated alias/placeholder lists), but no sensitivity analysis is reported. The 8–12 point improvement over AVClass is the central claim, and it is not clear how stable this margin is to reasonable changes in these thresholds (for example, S ∈ [0.9, 0.99] or T ∈ [500, 2000]). Please add a sensitivity analysis for the new alias-resolution thresholds and, at minimum, report the accuracy range across plausible settings.
- [§6.1, MalPedia cleaning] The MalPedia benchmark is cleaned by the authors before evaluation: generic family names are removed and unresolved aliases are fixed, but the procedure is not quantified and the cleaned labels are not released. Because the comparison in Table 2 depends on this cleaning—and on how each tool's alias mapping relates to it—the 12-point advantage on MalPedia may be sensitive to the cleaning decisions. Please publish the cleaned label mapping, quantify how many files and labels were changed, and verify that the results are stable under reasonable alternative cleaning choices.
minor comments (7)
- [Abstract] The sentence 'Automated tools using that label malware using antivirus detections lack accuracy and/or scalability' appears to contain a typo; please rephrase.
- [Section 2, item 3] The word 'aggretation' should be 'aggregation'.
- [§5.6] The phrase 'Shannon’s entropy the detected families' is missing 'of'; it should read 'Shannon’s entropy of the detected families.'
- [§6.2] The word 'recieve' should be 'receive.'
- [§6.1.6] The spelling 'ClarA Vy' is used inconsistently in the running text; please standardize to 'ClarAVy.'
- [References] Reference [29], cited as AVClass by Sebastián and Caballero, points to a GitHub URL for 'avclassplusplus'; please verify that the reference and URL match the intended tool.
- [Table 2] The main accuracy numbers are reported as point estimates; adding standard errors or confidence intervals would help readers judge the stability of the differences, even though the observed gaps are large.
Circularity Check
No significant circularity: the central family-labeling accuracy claims rest on external ground-truth benchmarks, not on fitted inputs or self-citation chains.
full rationale
The paper's headline accuracy results (64.16% on MOTIF, 56.88% on MalPedia) are evaluated against external, non-AV-derived ground-truth labels: MOTIF uses open-source reporting and MalPedia uses open-source reporting plus YARA rules. SparseIBCC is an unsupervised variational Bayesian aggregator over AV detection tokens; it is not trained on MOTIF or MalPedia family labels, and the parsing rules, alias lists, and AV-relationship clusters are inherited from prior VirusShare-based work rather than from the evaluation sets. The t_i in c_i assumption in Section 5.5 is an explicit modeling constraint that restricts the posterior to families present in a given scan; the paper itself discloses in Section 6.4 that 18.64% of MOTIF scans are impossible to label under this constraint because the true family never appears in any AV detection. That is a transparent limitation, not a hidden equivalence between the method's inputs and outputs, and it applies to all AV-aggregation baselines compared in the paper. The confidence-score XGBoost is trained on MOTIF/MalPedia with five-fold cross-validation, but it is a post-hoc calibration module and does not generate the family labels used for the headline comparison. Self-citations to the authors' prior ClarAVy work are incremental and non-load-bearing: the core claim of improved labeling accuracy is independently benchmarked against AVClass, EUPHONY, AVClass++, Sumav, and TagClass using external labels. No derivation step reduces to its own inputs, and no load-bearing argument is justified solely by a self-citation. The paper is self-contained against external benchmarks, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (7)
- sibling alias co-occurrence threshold S =
0.95 (default)
- sibling alias minimum count T =
1000 (default)
- parent-child alias edit threshold E =
0.6 (default)
- parent-child alias co-occurrence threshold C =
0.5 (default)
- tag vote threshold M =
M=5 for BEH/FILE, M=1 for VULN/PACK/GRP
- manually curated rule and list counts =
934 AV parsing rules, 250+ APT names, placeholder family lists
- SparseIBCC Dirichlet hyperparameters =
not specified in paper
assumptions (6)
- domain assumption MOTIF and MalPedia ground truth labels are correct and independent of antivirus outputs.
- domain assumption For every file, the true family appears among the family tokens in its AV detections (t_i in c_i).
- domain assumption Antivirus detections are generated independently given the true family (VB-IBCC conditional independence).
- domain assumption The known AV product relationship graph in Figure 3 correctly identifies correlated engines, so correlated votes can be merged.
- domain assumption The 934 parsing rules and curated lists cover AV detection structures well enough that parsing errors do not dominate family labeling.
- standard math The VB-IBCC variational updates as described by Simpson et al. are correct and applicable.
Cite this review
Pith. "Pith review of ClarAVy: A Tool for Scalable and Accurate Malware Family Labeling." pith.science (2026). https://pith.science/paper/J5H42LYV
@misc{pith2026250202759,
author = {Pith},
title = {Pith review of: ClarAVy: A Tool for Scalable and Accurate Malware Family Labeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/J5H42LYV}},
note = {Machine review of arXiv:2502.02759}
}
abstract
Determining the family to which a malicious file belongs is an essential component of cyberattack investigation, attribution, and remediation. Performing this task manually is time consuming and requires expert knowledge. Automated tools using that label malware using antivirus detections lack accuracy and/or scalability, making them insufficient for real-world applications. Three pervasive shortcomings in these tools are responsible: (1) incorrect parsing of antivirus detections, (2) errors during family alias resolution, and (3) an inappropriate antivirus aggregation strategy. To address each of these, we created our own malware family labeling tool called ClarAVy. ClarAVy utilizes a Variational Bayesian approach to aggregate detections from a collection of antivirus products into accurate family labels. Our tool scales to enormous malware datasets, and we evaluated it by labeling $\approx$40 million malicious files. ClarAVy has 8 and 12 percentage points higher accuracy than the prior leading tool in labeling the MOTIF and MalPedia datasets, respectively.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
[n. d.]. VirusShare.com - Because Sharing is Caring. https://virusshare.com/, Last accessed on 2024-11-17
work page 2024
-
[2]
Ulrich Bayer, Paolo Milani Comparetti, Clemens Hlauschek, Christopher Kruegel, and Engin Kirda. 2009. Scalable, behavior-based malware clustering. In NDSS 2009, 16th Annual Network and Distributed System Security Symposium . http: //www.eurecom.fr/publication/2783
work page 2009
- [3]
-
[4]
Marcus Botacin, Felipe Duarte Domingues, Fabrício Ceschin, Raphael Machnicki, Marco Antonio Zanata Alves, Paulo Lício de Geus, and André Grégio. 2022. Antiviruses under the microscope: A hands-on perspective. Computers & Security 112 (2022), 102500
work page 2022
-
[5]
Bob Carpenter. 2008. Multilevel bayesian models of categorical data annotation. Unpublished manuscript 17, 122 (2008), 45–50
work page 2008
-
[6]
Colon Osorio, Hongyuan Qiu, and Anthony Arrott
Fernando C. Colon Osorio, Hongyuan Qiu, and Anthony Arrott. 2015. Segmented sandboxing - A novel approach to Malware polymorphism detection. In 2015 10th International Conference on Malicious and Unwanted Software (MALW ARE). 59–68. https://doi.org/10.1109/MALWARE.2015.7413685
-
[7]
Alexander Philip Dawid and Allan M Skene. 1979. Maximum likelihood esti- mation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics) 28, 1 (1979), 20–28
1979
-
[8]
Karsten Hahn. [n. d.]. Malware Naming Hell Part 1: Taming the mess of AV detection names. https://www.gdatasoftware.com/blog/2019/08/35146-taming- the-mess-of-av-detection-names, Last accessed on 2024-11-17
work page 2019
Show all 44 references
-
[9]
Richard Harang and Ethan M. Rudd. 2020. SOREL-20M: A Large Scale Benchmark Dataset for Malicious PE Detection. arXiv:2012.07634 [cs.CR]
2020 arXiv
-
[10]
Wenyi Huang and Jack Stokes. 2016. MtNet: A Multi-Task Neural Network for Dynamic Malware Classification. In International Conference on Detection of Intrusions. https://doi.org/10.1007/978-3-319-40667-1_20
2016 doi
-
[11]
Bissyandé, Yves Le Traon, Jacques Klein, and Lorenzo Cavallaro
Médéric Hurier, Guillermo Suarez-Tangil, Santanu Kumar Dash, Tegawendé F. Bissyandé, Yves Le Traon, Jacques Klein, and Lorenzo Cavallaro. 2017. Euphony: Harmonious Unification of Cacophonous Anti-Virus Vendor Labels for Android Malware. In 2017 IEEE/ACM 14th International Conf...
2017 doi
-
[12]
Yongkang Jiang, Gaolei Li, and Shenghong Li. 2023. TagClass: A Tool for Extract- ing Class-Determined Tags from Massive Malware Labels via Incremental Parsing. In 2023 53rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). 193–200. https://doi...
2023
-
[13]
Yongkang Jiang, Gaolei Li, Shenghong Li, Ying Guo, and Kai Zhou. 2024. Crowd- sourcing Malware Family Annotation: Joint Class-Determined Tag Extraction and Weakly-Tagged Sample Inference. IEEE Transactions on Network and Service Management (2024), 1–1. https://doi.org/10.1109/...
2024
-
[14]
Robert J Joyce, Dev Amlani, Charles Nicholas, and Edward Raff. 2023. Motif: A malware reference dataset with ground truth family labels. Computers & Security 124 (2023), 102921
2023
-
[15]
Joyce, Edward Raff, and Charles Nicholas
Robert J. Joyce, Edward Raff, and Charles Nicholas. 2021. Rank-1 Similarity Matrix Decomposition For Modeling Changes in Antivirus Consensus Through Time. In Proceedings of the Conference on Applied Machine Learning in Information Security. 54–69
2021
-
[16]
Joyce, Edward Raff, Charles Nicholas, and James Holt
Robert J. Joyce, Edward Raff, Charles Nicholas, and James Holt. 2023. MalDICT: Benchmark Datasets on Malware Behaviors, Platforms, Exploitation, and Packers. In Proceedings of the Conference on Applied Machine Learning in Information Security. 105–121
2023
-
[17]
Sangwon Kim, Wookhyun Jung, KyungMin Lee, HyungGeun Oh, and Eui Tak Kim. 2022. Sumav: Fully automated malware labeling. ICT Express 8, 4 (2022), 530–538
2022
-
[18]
Yura Kurogome. 2019. AVCLASS++: Yet Another Massive Malware Labeling Tool. (2019). https://github.com/killvxk/avclassplusplus Black Hat Europe
2019
-
[19]
Peng Li, Limin Liu, Debin Gao, and Michael K. Reiter. 2010. On Challenges in Evaluating Malware Clustering. InRecent Advances in Intrusion Detection, Somesh Jha, Robin Sommer, and Christian Kreibich (Eds.). 238–255
2010
-
[20]
Abedelaziz Mohaisen and Omar Alrawi. 2013. Unveiling Zeus: Automated Classification of Malware Samples. In Proceedings of the 22nd International Conference on World Wide Web (Rio de Janeiro, Brazil) (WWW ’13 Compan- ion). Association for Computing Machinery, New York, NY, USA,...
2013
-
[21]
Aziz Mohaisen and Omar Alrawi. 2014. AV-Meter: An Evaluation of Antivirus Scans and Labels. In Detection of Intrusions and Malware, and Vulnerability Assess- ment - 11th International Conference, DIMV A 2014, Egham, UK, July 10-11, 2014. Proceedings (Lecture Notes in Computer ...
2014 doi
-
[22]
Aziz Mohaisen, Omar Alrawi, Matt Larson, and Danny McPherson. 2014. Towards a Methodical Evaluation of Antivirus Scans and Labels. In Information Security Applications, Yongdae Kim, Heejo Lee, and Adrian Perrig (Eds.). Cham, 231–241
2014
-
[23]
Aziz Mohaisen, Omar Alrawi, and Manar Mohaisen. 2015. AMAL: High-fidelity, behavior-based automated malware analysis and classification. Computers & Security 52 (2015), 251 – 266. https://doi.org/10.1016/j.cose.2015.04.001
2015 doi
-
[24]
Todd K Moon. 1996. The expectation-maximization algorithm. IEEE Signal processing magazine 13, 6 (1996), 47–60
1996
-
[25]
Manjunath
Lakshmanan Nataraj, Shanmugavadivel Karthikeyan, Grégoire Jacob, and B. Manjunath. 2011. Malware Images: Visualization and Automatic Classification. (07 2011). https://doi.org/10.1145/2016904.2016908
2011
-
[26]
Tirth Patel, Fred Lu, Edward Raff, Charles Nicholas, Cynthia Matuszek, and James Holt. 2023. Small Effect Sizes in Malware Detection? Make Harder Train/Test Splits! Proceedings of the Conference on Applied Machine Learning in Information Security (2023). https://arxiv.org/abs/...
2023 arXiv
-
[27]
D Plohmann, M Clauss, Steffen Enders, and Elmar Padilla. 2017. Malpedia: A Collaborative Effort to Inventorize the Malware Landscape. In The Journal on Cybercrime & Digital Investigations, Vol. 3. https://doi.org/10.18464/cybin.v3i1.17
2017 doi
-
[28]
Y. Qiao, X. Yun, and Y. Zhang. 2016. How to Automatically Identify the Homology of Different Malware. In 2016 IEEE Trustcom/BigDataSE/ISPA. 929–936. https: //doi.org/10.1109/TrustCom.2016.0158
2016
-
[29]
Marcos Sebastián and Juan Caballero. 2023. AVClass. (2023). https://github.com/ killvxk/avclassplusplus
2023
-
[30]
Marcos Sebastián, Richard Rivera, Platon Kotzias, and Juan Caballero. 2016. AV- class: A Tool for Massive Malware Labeling. InResearch in Attacks, Intrusions, and Defenses, Fabian Monrose, Marc Dacier, Gregory Blanc, and Joaquin Garcia-Alfaro (Eds.). Cham, 230–253
2016
-
[31]
Silvia Sebastián and Juan Caballero. 2020. AVClass2: Massive Malware Tag Extraction from AV Labels. CoRR abs/2006.10615 (2020). arXiv:2006.10615 https: //arxiv.org/abs/2006.10615
2020 arXiv
-
[32]
Maximilien Servajean, Alexis Joly, Dennis Shasha, Julien Champ, and Esther Pacitti. 2017. Crowdsourcing Thousands of Specialized Labels: A Bayesian Active Training Approach. IEEE Transactions on Multimedia 19, 6 (2017), 1376–1391. https://doi.org/10.1109/TMM.2017.2653763
2017
-
[33]
Edwin Simpson, Stephen Roberts, Ioannis Psorakis, and Arfon Smith. 2013. Dy- namic bayesian combination of multiple imperfect classifiers. Decision making and imperfection (2013), 1–35
2013
-
[34]
Vaibhav B Sinha, Sukrut Rao, and Vineeth N Balasubramanian. 2018. Fast dawid- skene: A fast vote aggregation scheme for sentiment classification. arXiv preprint arXiv:1803.02781 (2018)
2018 arXiv
-
[35]
Roundy, and Nicolas Christin
Kyle Soska, Chris Gates, Kevin A. Roundy, and Nicolas Christin. 2017. Automatic Application Identification from Billions of Files. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Hali- fax, NS, Canada) (KDD ’17). Associati...
2017
-
[36]
Daniele Ucci, Leonardo Aniello, and Roberto Baldoni. 2019. Survey of machine learning techniques for malware analysis. Computers & Security 81 (2019), 123 –
2019
-
[37]
VirusTotal. [n. d.]. File statistics during last 7 days. https://www.virustotal.com/ en/statistics/, Last accessed on 2024-11-17
2024
-
[38]
Foster, and Michelle L
Daniel Votipka, Seth Rabin, Kristopher Micinski, Jeffrey S. Foster, and Michelle L. Mazurek. 2020. An Observational Investigation of Reverse Engineers’ Processes. In 29th USENIX Security Symposium (USENIX Security 20) . USENIX Association, 1875–1892. https://www.usenix.org/con...
2020
-
[39]
Mayuri Wadkar, Fabio Di Troia, and Mark Stamp. 2020. Detecting malware evolution using support vector machines. Expert Systems with Applications 143 (2020), 113022. https://doi.org/10.1016/j.eswa.2019.113022
2020
-
[40]
Limin Yang, Arridhana Ciptadi, Ihar Laziuk, Ali Ahmadzadeh, and Gang Wang
-
[41]
Yuchen Zhang, Xi Chen, Dengyong Zhou, and Michael I Jordan. 2014. Spectral methods meet EM: A provably optimal algorithm for crowdsourcing. Advances in neural information processing systems 27 (2014)
2014
-
[42]
Shuofei Zhu, Jianjun Shi, Limin Yang, Boqin Qin, Ziyi Zhang, Linhai Song, and Gang Wang. 2020. Measuring and Modeling the Label Dynamics of On- line Anti-Malware Engines. In 29th USENIX Security Symposium (USENIX Secu- rity 20). USENIX Association, 2361–2378. https://www.useni...
2020
-
[147]
https://doi.org/10.1016/j.cose.2018.11.001
2018 doi
-
[2021]
In 2021 IEEE Security and Privacy Workshops (SPW)
BODMAS: An open dataset for learning based temporal analysis of PE malware. In 2021 IEEE Security and Privacy Workshops (SPW) . IEEE, 78–84
2021
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.