REVIEW 3 major objections 5 minor 62 references
OptimOTU: Taxonomically aware OTU clustering with optimized thresholds and a bioinformatics workflow for metabarcoding data
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read OptimOTU clusters metabarcoding sequences using per-taxon genetic-distance thresholds learned from reference taxonomy, instead of one global cutoff, and claims this yields OTUs that better match species boundaries.
desk verdict Useful methods paper with a real software contribution, but the central claim about optimized thresholds is unproven: no benchmark, and the thresholds are tuned on a different clustering procedure than the one actually deployed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the per-taxon optimized threshold: a genetic-distance cutoff, chosen separately for each ancestor taxon and each taxonomic rank, that makes a single-linkage cut of reference sequences most closely reproduce their labelled descendant taxa as measured by adjusted mutual information. This threshold is what lets clustering adapt to taxa with different rates of within-species variation; thresholds are inherited down the hierarchy and are applied to pseudotaxa when a taxon lacks enough reference sequences to optimize its own.
What would settle it
Construct a mock community with known species spanning several phyla, run OptimOTU on its sequences, and compare the recovered clusters to the true species either by adjusted Rand index or by counting merged and split species; if per-taxon optimized thresholds do not match true species boundaries better than the best single global threshold, the paper's central claim is wrong.
Extended reading notes
Core claim
OptimOTU's central claim is that the genetic-distance threshold at which sequences should be grouped into OTUs is not a single number but a property of each taxon and each rank, and that the right thresholds can be learned from a reference set whose sequences already carry taxonomic labels. The algorithm builds a single-linkage hierarchy of the reference sequences, cuts it at a grid of candidate thresholds, and scores each cut against the reference taxonomy at every rank using adjusted mutual information; the threshold that best recovers the named descendant taxa is stored for that ancestor taxon. Query sequences are first given a preliminary taxonomic identification with a stopping condition, then clustered rank by rank. Within each rank, sequences that share an identification at that rank form cluster cores regardless of distance; unidentified sequences are attached to the nearest core by closed-reference clustering at the learned threshold, with iterations that mimic single linkage; and sequences still unattached are clustered de novo, receiving placeholder 'pseudotaxa' names. The output is a full hierarchical classification in which every OTU is either a named taxon or a pseudotaxon, and the clustering is constrained to be congruent with the preliminary taxonomy.
Load-bearing premise
The method assumes that the reference sequences and their taxonomic labels are accurate and representative enough that the threshold which best matches the reference taxonomy for a taxon will also be the right threshold for query sequences from that taxon.
Editorial extensions
If this is right
- A single metabarcoding dataset spanning several phyla will no longer have to be clustered with one threshold that over-merges some groups and over-splits others; per-taxon thresholds should reduce both errors within the same dataset.
- Because clustering is constrained by the preliminary taxonomy, the resulting OTU table comes with a complete rank-by-rank classification, with pseudotaxa standing in for organisms that have no named close relative in the reference database.
- Optimized thresholds are a reusable resource: computing them once for a marker gene and taxonomic group lets future studies cluster new samples without repeating the expensive optimization step.
- The pipe-based external distance interface removes the need to hold a full distance matrix in memory, which is what allows the pipeline to scale to datasets with millions of reads per sample.
Reading between the lines
- A consequence the paper does not draw is that the threshold optimization could be repurposed for other grouping objectives, such as functional genes or ecological guilds, by swapping the reference labels; nothing in the algorithm itself is species-specific.
- The paper's reliance on reference taxonomy implies that a biased or mislabeled reference database will systematically miscalibrate every threshold inherited from it; this makes the method's real-world accuracy sensitive to database quality in ways a mock-community benchmark could quantify.
- A testable extension would be to measure whether the benefit of per-taxon thresholds grows as the taxonomic breadth of a dataset increases, since the single-threshold approach should degrade most where lineages differ most in their intraspecific variation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents OptimOTU, an OTU clustering algorithm and bioinformatics pipeline for metabarcoding data. The algorithm first uses taxonomically identified reference sequences to estimate per-taxon genetic distance thresholds by comparing single-linkage partitions at various thresholds to reference taxonomy via AMI and other indices. Query sequences are then taxonomically classified, cluster cores are formed from identifications at each rank, unidentified sequences are attached by closed-reference clustering using optimized thresholds, and remaining sequences are clustered de novo into pseudotaxa. The paper also describes a full pipeline for paired-end Illumina data, including denoising, chimera removal, taxonomic assignment, and clustering, with parallelization options and multiple distance calculation methods. The R/C++ package and pipeline code are open source.
Significance. The per-taxon threshold idea is timely and the software is substantial: the implementation is open source, includes a test suite checking consistency of the clustering algorithms with base R hclust(), and the pipeline addresses real bioinformatics needs for large datasets. If the central claim is validated, taxonomically aware threshold optimization could improve OTU delimitation over global thresholds, especially in diverse communities with incomplete reference libraries. However, the paper's central claims are currently unsupported by empirical evidence: there is no benchmark, no mock community analysis, and no demonstration that thresholds generalize from reference sequences to query sequences. Because the clustering is constrained by the same taxonomy used for calibration, the reported congruence with taxonomy is partly by construction. The contribution is therefore a well-described algorithmic framework that needs substantial additional validation before the performance claims can be accepted.
major comments (3)
- [Implementation — threshold optimization] The threshold optimization step is supervised calibration: thresholds are chosen to maximize AMI (or another index) between single-linkage partitions of reference sequences and the reference taxonomy. The same taxonomy is then used at the clustering stage to form cluster cores (all sequences with the same identification at the current rank are grouped regardless of distance). Consequently, any report that the final OTUs 'closely match' the taxonomy is partly by construction. The manuscript provides no independent validation—such as mock communities, leave-one-out cross-validation of thresholds, or a comparison against a global-threshold baseline—that would break this circularity and support the claim that optimized thresholds generalize to query sequences.
- [Implementation — clustering process] The optimization measures how well single-linkage partitions match the reference taxonomy, but the deployed clustering algorithm does not use single-linkage to form the final partition for taxonomically identified sequences. Cluster cores are formed by grouping all sequences with the same taxonomic identification regardless of distance (Figure 3, examples A and B), and species S5 is kept as one cluster despite one sequence lying outside the 'species-level threshold' from the others. The optimized threshold therefore governs only closed-reference attachment of unidentified sequences and de novo clustering of pseudotaxa, not the threshold that splits or merges named taxa. The paper does not demonstrate that thresholds selected by matching single-linkage partitions to the taxonomy are the right radii for those two roles, nor that per-taxon thresholds outperform a single global threshold in those roles. This disconnect needs to be addressed explicitly and tested.
- [Whole manuscript — no empirical validation] The central claims of the paper—that OptimOTU produces OTUs that better reflect species boundaries than a single global threshold and that the pipeline scales to millions of reads per sample across tens of thousands of samples—are not supported by any empirical evaluation. There is no mock community analysis, no comparison with VSEARCH 97% clustering, DADA2 followed by clustering, or dnabarcoder, no cross-validation of the threshold optimization, and no runtime or memory benchmarks to substantiate the scaling statement in the abstract. Since these claims are quantitative and comparative, the current manuscript is a software description rather than a demonstrated method. A benchmark section with at least one public mock community and a comparison to standard tools is necessary to make the claims credible.
minor comments (5)
- [Phase 1 — ASV table construction] The text reads 'In the OptimOTU pipeine'; 'pipeine' should be 'pipeline'.
- [The OptimOTU pipeline] The text reads 'used to process demultipexed paired-end reads'; 'demultipexed' should be 'demultiplexed'.
- [Clustering algorithms] The statement that SLINK 'operates at the theoretically optimal complexities of O(n2) time and O(n) space' should use superscript notation O(n^2), and if the complexity claim is intended literally, a citation or proof should be given; as written it is a typesetting issue.
- [Threshold optimization — Figure 2 caption] The sentence 'Dashed vertical lines indicate thresholds which optimize AMI for each rank, with ties resolved by selecting the median threshold and rounding down' is ambiguous: the median of which set of thresholds, and why rounding down? Please clarify.
- [References] The citation for dnabarcoder (Vu, Nilsson, and Verkley 2022) appears to be a preprint posted on Authorea; if a peer-reviewed version is available, it should be cited instead.
Circularity Check
No circularity: the threshold optimization is an explicitly supervised calibration, and no fitted parameter is renamed as an independent prediction.
full rationale
The paper's central step is the supervised selection of clustering thresholds by maximizing agreement with a reference taxonomy. This is stated openly ('The resulting clusters are then compared to the taxonomic identifications, to determine the threshold that produces clusters that most closely match the taxonomy at each rank'), so the 'optimal' label is definitional rather than a disguised derivation. The use of taxonomy to form cluster cores ('cluster cores are formed by grouping sequences which have the same taxonomic identifications at the current rank') is an explicit design constraint, not a hidden circular step: the paper never claims to predict taxonomy from thresholds, and the thresholds are used for attaching unidentified sequences and de novo clustering rather than for re-deriving the named groups. Self-citations appear only as software or prior-work references (e.g., PROTAX, Ovaskainen et al. 2024, crew.cluster) and are not load-bearing for the algorithm's claimed operation. The absence of an independent mock-community or cross-validated benchmark is a validation gap and a correctness risk, but it is not an instance of a result reducing to its own inputs by construction.
Assumptions & free parameters
free parameters (4)
- Per-taxon genetic distance thresholds =
not reported in paper; selected by maximizing AMI/ARI/FM over a grid, e.g., 0.0 to 0.4
- Minimum taxon size for threshold optimization =
10 sequences and 5 descendant taxa (defaults)
- Taxonomic assignment confidence thresholds =
50% (probable) and 90% (reliable)
- Test threshold grid range and step =
0.0 to 0.4 by 0.001 (example)
assumptions (4)
- domain assumption Reference sequences and their taxonomic identifications are accurate and representative of the target taxa.
- domain assumption Genetic distance computed from pairwise alignment is a biologically meaningful proxy for taxonomic boundaries at all ranks.
- domain assumption Single-linkage clustering is an appropriate model for OTU formation.
- domain assumption Taxonomic identifications supplied by the chosen classifier are reliable enough to constrain clustering (same ID never separated, different IDs never merged).
invented entities (1)
-
Pseudotaxa
Cite this review
Pith. "Pith review of OptimOTU: Taxonomically aware OTU clustering with optimized thresholds and a bioinformatics workflow for metabarcoding data." pith.science (2026). https://pith.science/paper/5SVAOR6X
@misc{pith2026250210350,
author = {Pith},
title = {Pith review of: OptimOTU: Taxonomically aware OTU clustering with optimized thresholds and a bioinformatics workflow for metabarcoding data},
year = {2026},
howpublished = {\url{https://pith.science/paper/5SVAOR6X}},
note = {Machine review of arXiv:2502.10350}
}
read the original abstract
To turn environmentally derived metabarcoding data into community matrices for ecological analysis, sequences must first be clustered into operational taxonomic units (OTUs). This task is particularly complex for data including large numbers of taxa with incomplete reference libraries. OptimOTU offers a taxonomically aware approach to OTU clustering. It uses a set of taxonomically identified reference sequences to choose optimal genetic distance thresholds for grouping each ancestor taxon into clusters which most closely match its descendant taxa. Then, query sequences are clustered according to preliminary taxonomic identifications and the optimized thresholds for their ancestor taxon. The process follows the taxonomic hierarchy, resulting in a full taxonomic classification of all the query sequences into named taxonomic groups as well as placeholder "pseudotaxa" which accommodate the sequences that could not be classified to a named taxon at the corresponding rank. The OptimOTU clustering algorithm is implemented as an R package, with computationally intensive steps implemented in C++ for speed, and incorporating open-source libraries for pairwise sequence alignment. Distances may also be calculated externally, and may be read from a UNIX pipe, allowing clustering of large datasets where the full distance matrix would be inconveniently large to store in memory. The OptimOTU bioinformatics pipeline includes a full workflow for paired-end Illumina sequencing data that incorporates quality filtering, denoising, artifact removal, taxonomic classification, and OTU clustering with OptimOTU. The OptimOTU pipeline is developed for use on high performance computing clusters, and scales to datasets with millions of reads per sample, and tens of thousands of samples.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Henrik Nilsson, Karl-Henrik Larsson, Ian J
Abarenkov, Kessy, R. Henrik Nilsson, Karl-Henrik Larsson, Ian J. Alexander, Ursula Eberhardt, Susanne Erland, Klaus Høiland, et al. 2010. ``The UNITE Database for Molecular Identification of Fungi -- Recent Updates and Future Perspectives.'' New Phytologist 186 (2): 281--85. https://doi.org/10.1111/j.1469-8137.2009.03160.x
arXiv 2010
-
[2]
Abarenkov, Kessy, Panu Somervuo, R. Henrik Nilsson, Paul M. Kirk, Tea Huotari, Nerea Abrego, and Otso Ovaskainen. 2018. ``Protax-Fungi: A Web-Based Tool for Probabilistic Taxonomic Placement of Fungal Internal Transcribed Spacer Sequences.'' New Phytologist 220 (2): 517--25. https://doi.org/10.1111/nph.15301
-
[3]
Navas-Molina, Evguenia Kopylova, James T
Amir, Amnon, Daniel McDonald, Jose A. Navas-Molina, Evguenia Kopylova, James T. Morton, Zhenjiang Zech Xu, Eric P. Kightley, et al. 2017. ``Deblur Rapidly Resolves Single-Nucleotide Community Sequence Patterns .'' mSystems 2 (2): e00191--16. https://doi.org/10.1128/mSystems.00191-16
-
[4]
Anslan, Sten, Mohammad Bahram, Indrek Hiiesalu, and Leho Tedersoo. 2017. `` PipeCraft : Flexible Open-Source Toolkit for Bioinformatics Analysis of Custom High-Throughput Amplicon Sequencing Data.'' Molecular Ecology Resources 17 (6): e234--40. https://doi.org/10.1111/1755-0998.12692
-
[5]
Kaehler, Jai Ram Rideout, Matthew Dillon, Evan Bolyen, Rob Knight, Gavin A
Bokulich, Nicholas A., Benjamin D. Kaehler, Jai Ram Rideout, Matthew Dillon, Evan Bolyen, Rob Knight, Gavin A. Huttley, and J. Gregory Caporaso. 2018. ``Optimizing Taxonomic Classification of Marker-Gene Amplicon Sequences with QIIME 2's Q2-Feature-Classifier Plugin.'' Microbiome 6 (1): 90. https://doi.org/10.1186/s40168-018-0470-z
-
[6]
Callahan, Benjamin J., Paul J. McMurdie, and Susan P. Holmes. 2017. ``Exact Sequence Variants Should Replace Operational Taxonomic Units in Marker-Gene Data Analysis.'' The ISME Journal 11 (12): 2639--43. https://doi.org/10.1038/ismej.2017.119
-
[7]
Callahan, Benjamin J., Paul J. McMurdie, Michael J. Rosen, Andrew W. Han, Amy Jo A. Johnson, and Susan P. Holmes. 2016. `` DADA2 : High-resolution Sample Inference from Illumina Amplicon Data.'' Nature Methods 13 (7): 581--83. https://doi.org/10.1038/nmeth.3869
-
[8]
Ching, Travers. 2024. ``Qs: Quick Serialization of R Objects.'' https://github.com/qsbase/qs
work page 2024
Show all 62 references
-
[9]
conda contributors. n.d. ``Conda: A System-Level, Binary Package and Environment Manager Running on All Major Operating Systems and Platforms.'' https://github.com/conda/conda
-
[10]
Czech, Lucas, Pierre Barbera, and Alexandros Stamatakis. 2020. ``Genesis and Gappa : Processing, Analyzing and Visualizing Phylogenetic (Placement) Data.'' Bioinformatics 36 (10): 3263--65. https://doi.org/10.1093/bioinformatics/btaa070
2020 doi
-
[11]
Dondoshansky, I, and Y Wolf. 2000. `` BLASTCLUST - BLAST Score-Based Single Linkage Clustering.'' National Center for Biotechnology Information. ftp://ftp.ncbi.nih.gov/blast/documents/blastclust.html
2000
-
[12]
Eddy, Sean R. 2011. ``Accelerated Profile HMM Searches .'' PLOS Computational Biology 7 (10): e1002195. https://doi.org/10.1371/journal.pcbi.1002195
2011 doi
-
[13]
Eddy, Sean R., and Richard Durbin. 1994. `` RNA Sequence Analysis Using Covariance Models.'' Nucleic Acids Research 22 (11): 2079--88. https://doi.org/10.1093/nar/22.11.2079
1994 doi
-
[14]
Edgar, Robert C. 2010. ``Search and Clustering Orders of Magnitude Faster Than BLAST .'' Bioinformatics 26 (19): 2460--61. https://doi.org/10.1093/bioinformatics/btq461
2010 doi
-
[15]
---------. 2016a. `` SINTAX : A Simple Non- Bayesian Taxonomy Classifier for 16S and ITS Sequences.'' bioRxiv, September, 074161. https://doi.org/10.1101/074161
-
[16]
---------. 2016b. `` UNOISE2 : Improved Error-Correction for Illumina 16S and ITS Amplicon Sequencing.'' bioRxiv, October, 081257. https://doi.org/10.1101/081257
-
[17]
---------. 2018. `` UNCROSS2 : Identification of Cross-Talk in 16S rRNA OTU Tables.'' August 27, 2018. https://doi.org/10.1101/400762
2018 doi
-
[18]
Haas, Jose C
Edgar, Robert C., Brian J. Haas, Jose C. Clemente, Christopher Quince, and Rob Knight. 2011. `` UCHIME Improves Sensitivity and Speed of Chimera Detection.'' Bioinformatics 27 (16): 2194--2200. https://doi.org/10.1093/bioinformatics/btr381
2011 doi
-
[19]
B., and C
Fowlkes, E. B., and C. L. Mallows. 1983. ``A Method for Comparing Two Hierarchical Clusterings .'' Journal of the American Statistical Association 78 (383): 553--69. https://doi.org/10.1080/01621459.1983.10478008
1983
-
[20]
Guillou, Laure, Dipankar Bachar, Stéphane Audic, David Bass, Cédric Berney, Lucie Bittner, Christophe Boutte, et al. 2013. ``The Protist Ribosomal Reference Database ( PR2 ): A Catalog of Unicellular Eukaryote Small Sub-Unit rRNA Sequences with Curated Taxonomy.'' Nucleic Acid...
2013 doi
-
[21]
Hubert, Lawrence, and Phipps Arabie. 1985. ``Comparing Partitions.'' Journal of Classification 2 (1): 193--218. https://doi.org/10.1007/BF01908075
1985 doi
-
[22]
Jamy, Mahwash, Rachel Foster, Pierre Barbera, Lucas Czech, Alexey Kozlov, Alexandros Stamatakis, Gary Bending, Sally Hilton, David Bass, and Fabien Burki. 2020. ``Long-Read Metabarcoding of the Eukaryotic rDNA Operon to Phylogenetically and Taxonomically Resolve Environmental ...
2020
-
[23]
Kauserud, Håvard. 2023. `` ITS Alchemy: On the Use of ITS as a DNA Marker in Fungal Ecology.'' Fungal Ecology, July, 101274. https://doi.org/10.1016/j.funeco.2023.101274
2023
-
[24]
Klik, Mark. 2022. ``Fst: Lightning Fast Serialization of Data Frames.'' http://www.fstpackage.org
2022
-
[25]
Saira Mian, Kimmen Sjölander, and David Haussler
Krogh, Anders, Michael Brown, I. Saira Mian, Kimmen Sjölander, and David Haussler. 1994. ``Hidden Markov Models in Computational Biology : Applications to Protein Modeling .'' Journal of Molecular Biology 235 (5): 1501--31. https://doi.org/10.1006/jmbi.1994.1104
1994
-
[26]
Landau, William Michael. 2021. ``The Targets R Package: A Dynamic Make-like Function-Oriented Pipeline Toolkit for Reproducibility and High-Performance Computing.'' Journal of Open Source Software 6 (57): 2959. https://doi.org/10.21105/joss.02959
2021 doi
-
[27]
Landau, William Michael, Michael Gilbert Levin, and Brendan Furneaux. 2024. ``Crew.cluster: Crew Launcher Plugins for Traditional High-Performance Computing Clusters.'' https://wlandau.github.io/crew.cluster/
2024
-
[28]
Jørgensen, Daniel H
Lanzén, Anders, Steffen L. Jørgensen, Daniel H. Huson, Markus Gorfer, Svenn Helge Grindhaug, Inge Jonassen, Lise Øvreås, and Tim Urich. 2012. `` CREST -- Classification Resources for Environmental Sequence Tags .'' PLOS ONE 7 (11): e49334. https://doi.org/10.1371/journal.pone.0049334
2012 doi
-
[29]
Li, Roy, Sujeevan Ratnasingham, Iuliia Zarubiieva, Panu Somervuo, and Graham W. Taylor. 2024. `` PROTAX-GPU : A Scalable Probabilistic Taxonomic Classification System for DNA Barcodes.'' Philosophical Transactions of the Royal Society B: Biological Sciences 379 (1904): 2023012...
1904
-
[30]
Li, Weizhong, and Adam Godzik. 2006. ``Cd-Hit: A Fast Program for Clustering and Comparing Large Sets of Protein or Nucleotide Sequences.'' Bioinformatics (Oxford, England) 22 (13): 1658--59. https://doi.org/10.1093/bioinformatics/btl158
2006 doi
-
[31]
Mahé, Frédéric, Torbjørn Rognes, Christopher Quince, Colomban de Vargas, and Micah Dunthorn. 2015. ``Swarm V2: Highly-Scalable and High-Resolution Amplicon Clustering.'' PeerJ 3: e1420. https://doi.org/10.7717/peerj.1420
2015 doi
-
[32]
Marco-Sola, Santiago, Juan Carlos Moure, Miquel Moreto, and Antonio Espinosa. 2021. ``Fast Gap-Affine Pairwise Alignment Using the Wavefront Algorithm.'' Bioinformatics 37 (4): 456--63. https://doi.org/10.1093/bioinformatics/btaa777
2021 doi
-
[33]
Martin, Marcel. 2011. ``Cutadapt Removes Adapter Sequences from High-Throughput Sequencing Reads.'' EMBnet.journal 17 (1): 10--12. https://doi.org/10.14806/ej.17.1.200
2011 doi
-
[34]
Matthews, B. W. 1975. ``Comparison of the Predicted and Observed Secondary Structure of T4 Phage Lysozyme.'' Biochimica Et Biophysica Acta (BBA) - Protein Structure 405 (2): 442--51. https://doi.org/10.1016/0005-2795(75)90109-9
1975 doi
-
[35]
Murali, Adithya, Aniruddha Bhargava, and Erik S. Wright. 2018. `` IDTAXA : A Novel Approach for Accurate Taxonomic Classification of Microbiome Sequences.'' Microbiome 6 (1): 140. https://doi.org/10.1186/s40168-018-0521-5
2018 doi
-
[36]
Nawrocki, Eric P. 2014. ``Annotating Functional RNAs in Genomes Using Infernal .'' In RNA Sequence , Structure , and Function : Computational and Bioinformatic Methods , edited by Jan Gorodkin and Walter L. Ruzzo, 163--97. Methods in Molecular Biology . Totowa, NJ: Humana Pres...
2014 doi
-
[37]
Henrik, Erik Kristiansson, Martin Ryberg, Nils Hallenberg, and Karl-Henrik Larsson
Nilsson, R. Henrik, Erik Kristiansson, Martin Ryberg, Nils Hallenberg, and Karl-Henrik Larsson. 2008. ``Intraspecific ITS Variability in the Kingdom Fungi as Expressed in the International Sequence Databases and Its Implications for Molecular Species Identification .'' Evoluti...
2008 doi
-
[38]
Nolet, Corey J., Divye Gala, Alex Fender, Mahesh Doijade, Joe Eaton, Edward Raff, John Zedlewski, Brad Rees, and Tim Oates. 2023. `` cuSLINK : Single-Linkage Agglomerative Clustering on the GPU .'' In Machine Learning and Knowledge Discovery in Databases : Research Track , edi...
2023 doi
-
[39]
Andrew, et al
Ovaskainen, Otso, Nerea Abrego, Brendan Furneaux, Bess Hardwick, Panu Somervuo, Isabella Palorinne, Nigel R. Andrew, et al. 2024. ``Global Spore Sampling Project : A Global, Standardized Dataset of Airborne Fungal DNA .'' Scientific Data 11 (1): 561. https://doi.org/10.1038/s4...
2024 doi
-
[40]
Andrew, et al
Ovaskainen, Otso, Nerea Abrego, Panu Somervuo, Isabella Palorinne, Bess Hardwick, Juha-Matti Pitkänen, Nigel R. Andrew, et al. 2020. ``Monitoring Fungal Communities With the Global Spore Sampling Project .'' Frontiers in Ecology and Evolution 7. https://doi.org/10.3389/fevo.2019.00511
2020
-
[41]
Pentinsaari, Mikko, Heli Salmela, Marko Mutanen, and Tomas Roslin. 2016. ``Molecular Evolution of a Widely-Adopted Taxonomic Marker ( COI ) Across the Animal Tree of Life.'' Scientific Reports 6 (1, 1): 35275. https://doi.org/10.1038/srep35275
2016 doi
-
[42]
M., and M
Porter, T. M., and M. Hajibabaei. 2021. ``Profile Hidden Markov Model Sequence Analysis Can Help Remove Putative Pseudogenes from DNA Barcoding and Metabarcoding Datasets.'' BMC Bioinformatics 22 (1): 256. https://doi.org/10.1186/s12859-021-04180-x
2021 doi
-
[43]
Rand, William M. 1971. ``Objective Criteria for the Evaluation of Clustering Methods .'' Journal of the American Statistical Association 66 (336): 846--50. https://doi.org/10.1080/01621459.1971.10482356
1971
-
[44]
Ratnasingham, Sujeevan, and Paul D. N. Hebert. 2013. ``A DNA-Based Registry for All Animal Species : The Barcode Index Number ( BIN ) System .'' PLOS ONE 8 (7): e66213. https://doi.org/10.1371/journal.pone.0066213
2013 doi
-
[45]
Rognes, Torbjørn, Tomáš Flouri, Ben Nichols, Christopher Quince, and Frédéric Mahé. 2016. `` VSEARCH : A Versatile Open Source Tool for Metagenomics.'' PeerJ 4 (October): e2584. https://doi.org/10.7717/peerj.2584
2016 doi
-
[46]
Romeijn, Luuk, Andrius Bernatavicius, and Duong Vu. 2024. `` MycoAI : Fast and Accurate Taxonomic Classification for Fungal ITS Sequences.'' Molecular Ecology Resources n/a (n/a): e14006. https://doi.org/10.1111/1755-0998.14006
2024
-
[47]
Roslin, Tomas, Panu Somervuo, Mikko Pentinsaari, Paul D. N. Hebert, Jireh Agda, Petri Ahlroth, Perttu Anttonen, et al. 2022. ``A Molecular-Based Identification Resource for the Arthropods of Finland .'' Molecular Ecology Resources 22 (2): 803--22. https://doi.org/10.1111/1755-...
2022
-
[48]
Westcott, Thomas Ryabin, Justine R
Schloss, Patrick D., Sarah L. Westcott, Thomas Ryabin, Justine R. Hall, Martin Hartmann, Emily B. Hollister, Ryan A. Lesniewski, et al. 2009. ``Introducing Mothur: Open-Source , Platform-Independent , Community-Supported Software for Describing and Comparing Microbial Communit...
2009 doi
-
[49]
Sibson, R. 1973. `` SLINK : An Optimally Efficient Algorithm for the Single-Link Cluster Method.'' The Computer Journal 16 (1): 30--34. https://doi.org/10.1093/comjnl/16.1.30
1973 doi
-
[50]
Henrik Nilsson, and Otso Ovaskainen
Somervuo, Panu, Sonja Koskela, Juho Pennanen, R. Henrik Nilsson, and Otso Ovaskainen. 2016. ``Unbiased Probabilistic Taxonomic Classification for DNA Barcoding.'' Bioinformatics 32 (19): 2920--27. https://doi.org/10.1093/bioinformatics/btw346
2016 doi
-
[51]
Yu, Charles C
Somervuo, Panu, Douglas W. Yu, Charles C. Y. Xu, Yinqiu Ji, Jenni Hultman, Helena Wirta, and Otso Ovaskainen. 2017. ``Quantifying Uncertainty of Taxonomic Placement in DNA Barcoding and Metabarcoding.'' Methods in Ecology and Evolution 8 (4): 398--407. https://doi.org/10.1111/...
2017 doi
-
[52]
Buhay, Michael F
Song, Hojun, Jennifer E. Buhay, Michael F. Whiting, and Keith A. Crandall. 2008. ``Many Species in One: DNA Barcoding Overestimates the Number of Species When Nuclear Mitochondrial Pseudogenes Are Coamplified.'' Proceedings of the National Academy of Sciences 105 (36): 13486--...
2008 doi
-
[53]
Šošić, Martin, and Mile Šikić. 2017. ``Edlib: A C / C ++ Library for Fast, Exact Sequence Alignment Using Edit Distance.'' Bioinformatics 33 (9): 1394--95. https://doi.org/10.1093/bioinformatics/btw753
2017 doi
-
[54]
Steinbach, Michael. 2000. ``A Comparison of Document Clustering Techniques.'' Technical Report\# 00\_034/University of Minnesota
2000
-
[55]
Strehl, Alexander, and Joydeep Ghosh. 2002. ``Cluster Ensembles -- A Knowledge Reuse Framework for Combining Multiple Partitions .'' Journal of Machine Learning Research 3: 583--617
2002
-
[56]
Tedersoo, Leho, Mahdieh S Hosseyni Moghaddam, Vladimir Mikryukov, Ali Hakimzadeh, Mohammad Bahram, R Henrik Nilsson, Iryna Yatsiuk, et al. 2024. `` EUKARYOME : The rRNA Gene Reference Database for Identification of All Eukaryotes.'' Database 2024 (February): baae043. https://d...
2024 doi
-
[57]
Tedersoo, Leho, Vladimir Mikryukov, Sten Anslan, Mohammad Bahram, Abdul Nasir Khalid, Adriana Corrales, Ahto Agan, et al. 2021. ``The Global Soil Mycobiome Consortium Dataset for Boosting Fungal Diversity Research.'' Fungal Diversity 111 (1): 573--88. https://doi.org/10.1007/s...
2021 doi
-
[58]
Vinh, Nguyen Xuan, Julien Epps, and James Bailey. 2010. ``Information Theoretic Measures for Clusterings Comparison : Variants , Properties , Normalization and Correction for Chance .'' Journal of Machine Learning Research 11: 2837--54
2010
-
[59]
Vu, Thuy, Rolf Henrik Nilsson, and Gerard Verkley. 2022. Dnabarcoder: An Open-Source Software Package for Analyzing and Predicting DNA Sequence Similarity Cut-Offs for Fungal Sequence Identification . https://doi.org/10.22541/au.164201896.67817672/v1
2022
-
[60]
Garrity, James M
Wang, Qiong, George M. Garrity, James M. Tiedje, and James R. Cole. 2007. ``Naïve Bayesian Classifier for Rapid Assignment of rRNA Sequences into the New Bacterial Taxonomy .'' Appl. Environ. Microbiol. 73 (16): 5261--67. https://doi.org/10.1128/AEM.00062-07
2007 doi
-
[61]
Watts, Corinne, Andrew Dopheide, Robert Holdaway, Carina Davis, Jamie Wood, Danny Thornburrow, and Ian A Dickie. 2019. `` DNA Metabarcoding as a Tool for Invertebrate Community Monitoring: A Case Study Comparison with Conventional Techniques.'' Austral Entomology 58 (3): 675--...
2019 doi
-
[62]
MVWߺ7ȈȈ/p[ox .5ۏ _ w;CK Ve1mKn+߷ > _||^g pÏ 3<Зn iYk<[b??O֗Qk ߭k Z Wq =Elsq]W?_^3 led
Zito, Alessandro, Tommaso Rigon, and David B. Dunson. 2023. ``Inferring Taxonomic Placement from DNA Barcoding Aiding in Discovery of New Taxa.'' Methods in Ecology and Evolution 14 (2): 529--42. https://doi.org/10.1111/2041-210X.14009. CSLReferences document pipeline-key-1.pd...
2023 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.