Pith. sign in

REVIEW 3 major objections 2 minor 204 references

Learning ON Large Datasets Using Bit-String Trees

T0 review · 3 major / 2 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The thesis claims that ComBI, a compressed BST of inverted hash tables, delivers fast approximate nearest-neighbor search at billion-sample scale with 0.90 precision and 4–296× speed-ups over Multi-Index Hashing, along with companion…

desk verdict Not reviewable as submitted: the full text is an entirely different paper, so none of the abstract's claims can be checked. read the letter →

arxiv 2508.17083 v1 pith:ZWRX7VFT submitted 2025-08-23 cs.LG

classification cs.LG
keywords approximatenearestneighborsearchbit-stringtreescompressedBSTofinvertedhashtablessimilarity-preservinghashingguidedrandomforestsingle-cellRNA-seqcancergenomicscodonswitchrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The thesis abstract proposes ComBI, a compressed BST of inverted hash tables, as a replacement for standard space-partitioning hashing in approximate nearest-neighbor search. The claimed payoff is fast, memory-reduced search on datasets up to one billion samples, with 0.90 precision and 4–296× speed-ups over Multi-Index Hashing and 2–13× gains over Cellfishing.jl on single-cell RNA-seq searches. The same hashing ideas are extended into GRAF, a guided random forest, and uGRAF, its unsupervised variant, which together with ComBI are said to estimate per-sample classifiability for scalable cancer-patient survival prediction. A third component, CRCS, embeds codon switches into numerical vectors for somatic mutation identification, driver-gene discovery, and tumor mutation scoring. The submitted full text is a different manuscript, so these are abstract-level claims that the supplied text does not document.

What carries the argument

The load-bearing objects are three. ComBI is a compressed BST of inverted hash tables, meaning each internal node stores an inverted index rather than a plain split; it is the mechanism claimed to cut memory and search time while keeping precision high. GRAF (and uGRAF) is a tree-ensemble classifier that combines global and local partitioning, bridging decision trees and boosting. CRCS is a deep embedding that maps codon switches to continuous vectors, letting mutations be scored without matched normal samples. Together they are claimed to form a single pipeline from hashing to classification to genomic prediction.

What would settle it

Re-run ComBI against Multi-Index Hashing and Cellfishing.jl on a publicly documented billion-vector dataset at matched recall levels; if the speed-ups fall below the claimed ranges or precision drops well below 0.90 at the standard operating point, the central claim fails.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that a compressed binary search tree made of inverted hash tables (ComBI) can preserve the indexing benefits of space-partitioning hashing while avoiding the exponential growth and sparsity that make ordinary BST-based hashes inefficient on large data. The abstract further claims that this structure yields 0.90 precision at up to one billion samples, outrunning Multi-Index Hashing by 4–296× and Cellfishing.jl by 2–13× on single-cell RNA-seq searches. The paper also claims that a guided random forest (GRAF), an unsupervised variant (uGRAF), and a continuous representation of codon switches (CRCS) extend the same hashing-and-partitioning ideas to competitive classification across 115 datasets and to cancer genomics, including survival prediction in bladder, liver, and brain cancers.

Load-bearing premise

The abstract's speed and precision numbers stand on the assumptions that the billion-sample benchmarks are representative, the baselines are configured competitively, precision is reported at a standard operating point, and the submitted full text actually contains these experiments; none of this can be checked from the supplied manuscript.

Editorial extensions

If this is right

  • If ComBI's speed and precision hold, billion-sample similarity search could move from cluster-scale hashing to a single machine without much accuracy loss.
  • The 2–13× gain over Cellfishing.jl would make single-cell RNA-seq marker searches interactive on large cell atlases.
  • GRAF's reported accuracy on 115 datasets would make guided tree ensembles a competitive default for tabular classification.
  • Per-sample classifiability from ComBI/GRAF could let survival models be trained and evaluated on very large cancer cohorts without matched normal tissue.
  • CRCS, if valid, would expand somatic-mutation discovery to tumor samples where matched normals are unavailable.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The abstract does not state the recall level at which 0.90 precision is measured; a natural test is to re-run the comparison at matched recall values, since speed-ups can change sharply with operating point.
  • If the compression in ComBI is what drives the gains, the same inverted-hash compressed tree should transfer to metric nearest-neighbor search beyond bit-string spaces; that is an extension the abstract does not claim.
  • The submitted full text is a different paper, so the empirical numbers should be treated as unverified until the experiments appear; this is a reader-side caution, not a verdict on the methods.
  • GRAF and ComBI's claimed ability to estimate per-sample classifiability could be tested directly by comparing its ranking of patients against standard survival-risk scores on the same cohorts.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The submission arXiv:2508.17083 consists of an abstract announcing four methods (ComBI, GRAF, uGRAF, CRCS) for similarity-preserving hashing, classification, and cancer genomics, with quantitative claims including 0.90 precision and 4X-296X speed-ups over Multi-Index Hashing on datasets of up to one billion samples. The supplied full text, however, is arXiv:2508.17090v4, a paper on neural stochastic differential equations on compact state spaces applied to suicide-risk modeling. That full text contains no mention of ComBI, GRAF, uGRAF, CRCS, bit-string trees, Multi-Index Hashing, Cellfishing.jl, or any of the abstract's experiments. The stress-test concern is confirmed: as submitted, the manuscript contains only an abstract with no matching technical content that can be checked.

Significance. If the abstract's claims were substantiated, the work would be significant: an order-of-magnitude faster approximate nearest-neighbor search at billion-sample scale with high precision, plus competitive classifiers and cancer-genomics tools, would be useful to the machine learning and biomedical communities. However, the submitted material provides no derivations, algorithms, datasets, experimental protocols, code, or proofs for any of these methods. The supplied full text is an unrelated paper, so the abstract's claims cannot be verified, reproduced, or placed in context. No strength of the claimed contributions can be assessed from the submission as it stands.

major comments (3)
  1. [Full Text (supplied manuscript, arXiv:2508.17090v4)] The submitted full text is a different paper. It is titled 'Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling' and its abstract, theorems, experiments, and appendices concern viability of SDEs on compact polyhedra and EMA suicide-risk data. It contains no occurrence of ComBI, GRAF, uGRAF, CRCS, bit-string trees, Multi-Index Hashing, Cellfishing.jl, or any of the benchmarks described in the submission's abstract. Because the central claims of the manuscript reside entirely in the abstract, and because the supplied full text does not address them, there is no object to review.
  2. [Abstract] The abstract's empirical claims are unsupported by any experimental detail. The sentence reporting '0.90 precision with 4X-296X speed-ups over Multi-Index Hashing' and '2X-13X gains' over Cellfishing.jl gives no precision-recall operating point, no dataset construction or size verification, no baseline configuration, no runtime measurement methodology, and no error bars or statistical tests. Even if a matching full text were supplied, these details would be required to evaluate whether the reported numbers are meaningful; in the present submission they are entirely absent.
  3. [Full Text / Appendices] The only code link and the only experimental tables in the supplied full text belong to the SDE paper, not to the abstract's methods. For example, Table 1 and Appendix G.3 describe EMA forecasting experiments for WSP-based latent neural SDEs, and the GitHub repository linked in the introduction is the WSP demo. These artifacts cannot provide support for the abstract's hashing, classification, or cancer-genomics claims, and no corresponding artifacts for ComBI, GRAF, uGRAF, or CRCS are present.
minor comments (2)
  1. [Title / Abstract] The title uses 'Bit-String Trees' while the abstract defines the approach in terms of 'Binary Search Trees' (BSTs); the relationship between these terms should be clarified or made consistent.
  2. [Abstract] The abstract combines four distinct contributions (ComBI, GRAF, uGRAF, CRCS) in a single submission without indicating how they relate methodologically beyond a shared hashing or tree-based theme; the intended narrative connection should be stated explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identifiable: the submitted full text is an unrelated paper, so the abstract's claimed derivation chain for ComBI is absent and cannot be audited.

full rationale

The manuscript under review consists of an abstract describing ComBI, GRAF, uGRAF, and CRCS, but the supplied full text is a different paper, 'Neural Stochastic Differential Equations on Compact State Spaces: Theory, Methods, and Application to Suicide Risk Modeling.' The full text contains no mention of ComBI, GRAF, uGRAF, CRCS, bit-string trees, Multi-Index Hashing, Cellfishing.jl, or any of the abstract's experiments. Consequently, there is no derivation chain, no fitted parameter, no self-citation chain, and no equation in the submitted text that could reduce the abstract's claims to their inputs by construction. The abstract's empirical assertions—'0.90 precision with 4X-296X speed-ups' and '2X-13X gains'—are unsupported by the supplied text, but missing evidence is not circularity. Under the hard rule to claim circularity only when the specific reduction can be quoted and exhibited, no circular step can be identified. The correct finding is therefore a non-finding on circularity: the artifact's substance is absent, and its correctness and reproducibility cannot be assessed. This is a completeness and provenance problem, not a circularity problem.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The abstract relies on background assertions about the inefficiency of BST-based hashing and about the transferability of classifier confidence to survival prediction. Because the submitted full text is a different paper, none of the underlying analyses, proofs, or experimental controls that would substantiate these premises are available. No invented physical entities are introduced; the contributions are algorithmic.

assumptions (2)
  • domain assumption Standard space partitioning-based hashing relies on Binary Search Trees, whose exponential growth and sparsity hinder efficiency.
    Stated in the abstract as the motivation for ComBI; no supporting analysis is visible in the abstract.
  • domain assumption Per-sample classifiability estimated by GRAF and uGRAF enables scalable prediction of cancer patient survival.
    Asserted in the abstract as a bridge between the ML methods and the clinical application; no evidence is available in the submission.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning ON Large Datasets Using Bit-String Trees." pith.science (2026). https://pith.science/paper/ZWRX7VFT

@misc{pith2026250817083,
  author       = {Pith},
  title        = {Pith review of: Learning ON Large Datasets Using Bit-String Trees},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZWRX7VFT}},
  note         = {Machine review of arXiv:2508.17083}
}
read the original abstract

This thesis develops computational methods in similarity-preserving hashing, classification, and cancer genomics. Standard space partitioning-based hashing relies on Binary Search Trees (BSTs), but their exponential growth and sparsity hinder efficiency. To overcome this, we introduce Compressed BST of Inverted hash tables (ComBI), which enables fast approximate nearest-neighbor search with reduced memory. On datasets of up to one billion samples, ComBI achieves 0.90 precision with 4X-296X speed-ups over Multi-Index Hashing, and also outperforms Cellfishing.jl on single-cell RNA-seq searches with 2X-13X gains. Building on hashing structures, we propose Guided Random Forest (GRAF), a tree-based ensemble classifier that integrates global and local partitioning, bridging decision trees and boosting while reducing generalization error. Across 115 datasets, GRAF delivers competitive or superior accuracy, and its unsupervised variant (uGRAF) supports guided hashing and importance sampling. We show that GRAF and ComBI can be used to estimate per-sample classifiability, which enables scalable prediction of cancer patient survival. To address challenges in interpreting mutations, we introduce Continuous Representation of Codon Switches (CRCS), a deep learning framework that embeds genetic changes into numerical vectors. CRCS allows identification of somatic mutations without matched normals, discovery of driver genes, and scoring of tumor mutations, with survival prediction validated in bladder, liver, and brain cancers. Together, these methods provide efficient, scalable, and interpretable tools for large-scale data analysis and biomedical applications.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

204 extracted references · 69 canonical work pages

  1. [1]

    Supervised hashing with kernels

    Wei Liu, Jun Wang, Rongrong Ji, Yu-Gang Jiang, and Shih-Fu Chang. Supervised hashing with kernels. In 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 2074--2081. IEEE, 2012

  2. [2]

    Similarity search in high dimensions via hashing

    Aristides Gionis, Piotr Indyk, Rajeev Motwani, et al. Similarity search in high dimensions via hashing. In Vldb , volume 99, pages 518--529, 1999

  3. [3]

    Similarity estimation techniques from rounding algorithms

    Moses S Charikar. Similarity estimation techniques from rounding algorithms. In Proceedings of the thiry-fourth annual ACM symposium on Theory of computing , pages 380--388, 2002

  4. [4]

    Semantic hashing

    Ruslan Salakhutdinov and Geoffrey Hinton. Semantic hashing. International Journal of Approximate Reasoning , 50(7):969--978, 2009

  5. [5]

    Spectral hashing

    Yair Weiss, Antonio Torralba, and Rob Fergus. Spectral hashing. In Advances in neural information processing systems , pages 1753--1760, 2009

  6. [6]

    Spherical hashing

    Jae-Pil Heo, Youngwoon Lee, Junfeng He, Shih-Fu Chang, and Sung-Eui Yoon. Spherical hashing. In 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 2957--2964. IEEE, 2012

  7. [7]

    Semi-supervised hashing for large-scale search

    Jun Wang, Sanjiv Kumar, and Shih-Fu Chang. Semi-supervised hashing for large-scale search. IEEE Transactions on Pattern Analysis and Machine Intelligence , 34(12):2393--2406, 2012

  8. [8]

    Hash function learning via codewords

    Yinjie Huang, Michael Georgiopoulos, and Georgios C Anagnostopoulos. Hash function learning via codewords. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases , pages 659--674. Springer, 2015

Show all 204 references
  1. [9]

    Visual saliency guided complex image retrieval

    Haoxiang Wang, Zhihui Li, Yang Li, BB Gupta, and Chang Choi. Visual saliency guided complex image retrieval. Pattern Recognition Letters , 130:64--72, 2020

  2. [10]

    Feature extraction and a database strategy for video fingerprinting

    Job Oostveen, Ton Kalker, and Jaap Haitsma. Feature extraction and a database strategy for video fingerprinting. In International Conference on Advances in Visual Information Systems , pages 117--128. Springer, 2002

  3. [11]

    A robust and fast video copy detection system using content-based fingerprinting

    Mani Malek Esmaeili, Mehrdad Fatourechi, and Rabab Kreidieh Ward. A robust and fast video copy detection system using content-based fingerprinting. IEEE Transactions on information forensics and security , 6(1):213--226, 2010

  4. [12]

    Fast matching for video/audio fingerprinting algorithms

    Mani Malek Esmaeili, Rabab K Ward, and Mehrdad Fatourechi. Fast matching for video/audio fingerprinting algorithms. In 2011 IEEE International Workshop on Information Forensics and Security , pages 1--6. IEEE, 2011

  5. [13]

    Audio fingerprinting: nearest neighbor search in high dimensional binary spaces

    Matthew L Miller, Manuel Acevedo Rodriguez, and Ingemar J Cox. Audio fingerprinting: nearest neighbor search in high dimensional binary spaces. Journal of VLSI signal processing systems for signal, image and video technology , 41(3):285--291, 2005

  6. [14]

    Cellfishing

    Kenta Sato, Koki Tsuyuzaki, Kentaro Shimizu, and Itoshi Nikaido. Cellfishing. jl: an ultrafast and scalable cell search method for single-cell rna sequencing. Genome biology , 20(1):31, 2019

  7. [15]

    Efficient iot-based sensor big data collection--processing and analysis in smart buildings

    Andreas P Plageras, Kostas E Psannis, Christos Stergiou, Haoxiang Wang, and Brij B Gupta. Efficient iot-based sensor big data collection--processing and analysis in smart buildings. Future Generation Computer Systems , 82:349--357, 2018

  8. [16]

    Security, privacy & efficiency of sustainable cloud computing for big data & iot

    Christos Stergiou, Kostas E Psannis, Brij B Gupta, and Yutaka Ishibashi. Security, privacy & efficiency of sustainable cloud computing for big data & iot. Sustainable Computing: Informatics and Systems , 19:174--184, 2018

  9. [17]

    Four-image encryption scheme based on quaternion fresnel transform, chaos and computer generated hologram

    Chuying Yu, Jianzhong Li, Xuan Li, Xuechang Ren, and Brij B Gupta. Four-image encryption scheme based on quaternion fresnel transform, chaos and computer generated hologram. Multimedia Tools and Applications , 77(4):4585--4608, 2018

  10. [18]

    Cellatlassearch: a scalable search engine for single cells

    Divyanshu Srivastava, Arvind Iyer, Vibhor Kumar, and Debarka Sengupta. Cellatlassearch: a scalable search engine for single cells. Nucleic acids research , 46(W1):W141--W147, 2018

  11. [19]

    Approximate nearest neighbors: towards removing the curse of dimensionality

    Piotr Indyk and Rajeev Motwani. Approximate nearest neighbors: towards removing the curse of dimensionality. In Proceedings of the thirtieth annual ACM symposium on Theory of computing , pages 604--613, 1998

  12. [20]

    Fast exact search in hamming space with multi-index hashing

    Mohammad Norouzi, Ali Punjani, and David J Fleet. Fast exact search in hamming space with multi-index hashing. IEEE transactions on pattern analysis and machine intelligence , 36(6):1107--1119, 2013

  13. [21]

    Fast nearest neighbor search in the hamming space

    Zhansheng Jiang, Lingxi Xie, Xiaotie Deng, Weiwei Xu, and Jingdong Wang. Fast nearest neighbor search in the hamming space. In International Conference on Multimedia Modeling , pages 325--336. Springer, 2016

  14. [22]

    A fast approximate nearest neighbor search algorithm in the hamming space

    Mani Malek Esmaeili, Rabab Kreidieh Ward, and Mehrdad Fatourechi. A fast approximate nearest neighbor search algorithm in the hamming space. IEEE transactions on pattern analysis and machine intelligence , 34(12):2481--2488, 2012

  15. [23]

    Lsh forest: self-tuning indexes for similarity search

    Mayank Bawa, Tyson Condie, and Prasanna Ganesan. Lsh forest: self-tuning indexes for similarity search. In Proceedings of the 14th international conference on World Wide Web , pages 651--660, 2005

  16. [24]

    Random projection trees and low dimensional manifolds

    Sanjoy Dasgupta and Yoav Freund. Random projection trees and low dimensional manifolds. In Proceedings of the fortieth annual ACM symposium on Theory of computing , pages 537--546, 2008

  17. [25]

    Randomized partition trees for exact nearest neighbor search

    Sanjoy Dasgupta and Kaushik Sinha. Randomized partition trees for exact nearest neighbor search. In Conference on Learning Theory , pages 317--337, 2013

  18. [26]

    Multidimensional binary search trees used for associative searching

    Jon Louis Bentley. Multidimensional binary search trees used for associative searching. Communications of the ACM , 18(9):509--517, 1975

  19. [27]

    Olshen and oj

    L Breiman and RA lH Friedman. Olshen and oj. Stone, Classification and Regression Trees, Wadsworth and Brooks , 1984

  20. [28]

    Random forests

    Leo Breiman. Random forests. Machine learning , 45(1):5--32, 2001

  21. [29]

    Extremely randomized trees

    Pierre Geurts, Damien Ernst, and Louis Wehenkel. Extremely randomized trees. Machine learning , 63(1):3--42, 2006

  22. [30]

    Five balltree construction algorithms

    Stephen M Omohundro. Five balltree construction algorithms . International Computer Science Institute Berkeley, 1989

  23. [31]

    An optimal algorithm for approximate nearest neighbor searching fixed dimensions

    Sunil Arya, David M Mount, Nathan S Netanyahu, Ruth Silverman, and Angela Y Wu. An optimal algorithm for approximate nearest neighbor searching fixed dimensions. Journal of the ACM (JACM) , 45(6):891--923, 1998

  24. [32]

    Birch: an efficient data clustering method for very large databases

    Tian Zhang, Raghu Ramakrishnan, and Miron Livny. Birch: an efficient data clustering method for very large databases. ACM sigmod record , 25(2):103--114, 1996

  25. [33]

    hdbscan: Hierarchical density based clustering

    Leland McInnes, John Healy, and Steve Astels. hdbscan: Hierarchical density based clustering. Journal of Open Source Software , 2(11):205, 2017

  26. [34]

    Agglomerative clustering via maximum incremental path integral

    Wei Zhang, Deli Zhao, and Xiaogang Wang. Agglomerative clustering via maximum incremental path integral. Pattern Recognition , 46(11):3056--3065, 2013

  27. [35]

    Isolation forest

    Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou. Isolation forest. In 2008 eighth ieee international conference on data mining , pages 413--422. IEEE, 2008

  28. [36]

    Adaptive random forests for evolving data stream classification

    Heitor M Gomes, Albert Bifet, Jesse Read, Jean Paul Barddal, Fabr \' cio Enembreck, Bernhard Pfharinger, Geoff Holmes, and Talel Abdessalem. Adaptive random forests for evolving data stream classification. Machine Learning , 106(9):1469--1495, 2017

  29. [37]

    Discovery of rare cells from voluminous single cell expression data

    Aashi Jindal, Prashant Gupta, Debarka Sengupta, et al. Discovery of rare cells from voluminous single cell expression data. Nature communications , 9(1):1--9, 2018

  30. [38]

    Classifying many-class high-dimensional fingerprint datasets using random forest of oblique decision trees

    Thanh-Nghi Do, Philippe Lenca, and St \'e phane Lallich. Classifying many-class high-dimensional fingerprint datasets using random forest of oblique decision trees. Vietnam journal of computer science , 2(1):3--12, 2015

  31. [39]

    Random forest in remote sensing: A review of applications and future directions

    Mariana Belgiu and Lucian Dr a gu t . Random forest in remote sensing: A review of applications and future directions. ISPRS journal of photogrammetry and remote sensing , 114:24--31, 2016

  32. [40]

    Oblique random forest based on partial least squares applied to pedestrian detection

    Artur Jordao Lima Correia and William Robson Schwartz. Oblique random forest based on partial least squares applied to pedestrian detection. In 2016 IEEE International Conference on Image Processing (ICIP) , pages 2931--2935. IEEE, 2016

  33. [41]

    Oblique random forest ensemble via least square estimation for time series forecasting

    Xueheng Qiu, Le Zhang, Ponnuthurai Nagaratnam Suganthan, and Gehan AJ Amaratunga. Oblique random forest ensemble via least square estimation for time series forecasting. Information Sciences , 420:249--262, 2017

  34. [42]

    Robust visual tracking using oblique random forests

    Le Zhang, Jagannadan Varadarajan, Ponnuthurai Nagaratnam Suganthan, Narendra Ahuja, and Pierre Moulin. Robust visual tracking using oblique random forests. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 5589--5598, 2017

  35. [43]

    Random forests for genomic data analysis

    Xi Chen and Hemant Ishwaran. Random forests for genomic data analysis. Genomics , 99(6):323--329, 2012

  36. [44]

    Resource-efficient machine learning in 2 kb ram for the internet of things

    Ashish Kumar, Saurabh Goyal, and Manik Varma. Resource-efficient machine learning in 2 kb ram for the internet of things. In International Conference on Machine Learning , pages 1935--1944. PMLR, 2017

  37. [45]

    Robust head pose estimation using dirichlet-tree distribution enhanced random forests

    Yuanyuan Liu, Jingying Chen, Zhiming Su, Zhenzhen Luo, Nan Luo, Leyuan Liu, and Kun Zhang. Robust head pose estimation using dirichlet-tree distribution enhanced random forests. Neurocomputing , 173:42--53, 2016

  38. [46]

    Ferret: a toolkit for content-based similarity search of feature-rich data

    Qin Lv, William Josephson, Zhe Wang, Moses Charikar, and Kai Li. Ferret: a toolkit for content-based similarity search of feature-rich data. In Proceedings of the 1st ACM SIGOPS/EuroSys European Conference on Computer Systems 2006 , pages 317--330, 2006

  39. [47]

    Sizing sketches: a rank-based analysis for similarity search

    Zhe Wang, Wei Dong, William Josephson, Qin Lv, Moses Charikar, and Kai Li. Sizing sketches: a rank-based analysis for similarity search. In Proceedings of the 2007 ACM SIGMETRICS international conference on Measurement and modeling of computer systems , pages 157--168, 2007

  40. [48]

    A method and server for predicting damaging missense mutations

    Ivan A Adzhubei, Steffen Schmidt, Leonid Peshkin, Vasily E Ramensky, Anna Gerasimova, Peer Bork, Alexey S Kondrashov, and Shamil R Sunyaev. A method and server for predicting damaging missense mutations. Nature methods , 7(4):248--249, 2010

  41. [49]

    Sift missense predictions for genomes

    Robert Vaser, Swarnaseetha Adusumalli, Sim Ngak Leng, Mile Sikic, and Pauline C Ng. Sift missense predictions for genomes. Nature protocols , 11(1):1, 2016

  42. [50]

    Disease variant prediction with deep generative models of evolutionary data

    Jonathan Frazer, Pascal Notin, Mafalda Dias, Aidan Gomez, Joseph K Min, Kelly Brock, Yarin Gal, and Debora S Marks. Disease variant prediction with deep generative models of evolutionary data. Nature , 599(7883):91--95, 2021

  43. [51]

    Efficient estimation of word representations in vector space

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 , 2013

  44. [52]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 , 2018

  45. [53]

    Do we need hundreds of classifiers to solve real world classification problems? The Journal of Machine Learning Research , 15(1):3133--3181, 2014

    Manuel Fern \'a ndez-Delgado, Eva Cernadas, Sen \'e n Barro, and Dinani Amorim. Do we need hundreds of classifiers to solve real world classification problems? The Journal of Machine Learning Research , 15(1):3133--3181, 2014

  46. [54]

    Hashing for similarity search: A survey

    Jingdong Wang, Heng Tao Shen, Jingkuan Song, and Jianqiu Ji. Hashing for similarity search: A survey. arXiv preprint arXiv:1408.2927 , 2014

  47. [55]

    Approximate nearest neighbor search on high dimensional data-experiments, analyses, and improvement

    Wen Li, Ying Zhang, Yifang Sun, Wei Wang, Mingjie Li, Wenjie Zhang, and Xuemin Lin. Approximate nearest neighbor search on high dimensional data-experiments, analyses, and improvement. IEEE Transactions on Knowledge and Data Engineering , 2019

  48. [56]

    Hbst: A hamming distance embedding binary search tree for feature-based visual place recognition

    Dominik Schlegel and Giorgio Grisetti. Hbst: A hamming distance embedding binary search tree for feature-based visual place recognition. IEEE Robotics and Automation Letters , 3(4):3741--3748, 2018

  49. [57]

    Locality-sensitive hashing without false negatives

    Rasmus Pagh. Locality-sensitive hashing without false negatives. In Proceedings of the twenty-seventh annual ACM-SIAM symposium on Discrete algorithms , pages 1--9. SIAM, 2016

  50. [58]

    Scalability and total recall with fast coveringlsh

    Ninh Pham and Rasmus Pagh. Scalability and total recall with fast coveringlsh. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management , pages 1109--1118, 2016

  51. [59]

    Online nearest neighbor search in binary space

    Sepehr Eghbali, Hassan Ashtiani, and Ladan Tahvildari. Online nearest neighbor search in binary space. In 2017 IEEE International Conference on Data Mining (ICDM) , pages 853--858. IEEE, 2017

  52. [60]

    Online nearest neighbor search using hamming weight trees

    Sepehr Eghbali, Hassan Ashtiani, and Ladan Tahvildari. Online nearest neighbor search using hamming weight trees. IEEE transactions on pattern analysis and machine intelligence , 2019

  53. [61]

    Fast and compact hamming distance index

    Simon Gog and Rossano Venturini. Fast and compact hamming distance index. In Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval , pages 285--294, 2016

  54. [62]

    o nen, Teemu Pitk \

    Ville Hyv \"o nen, Teemu Pitk \"a nen, Sotiris Tasoulis, Elias J \"a \"a saari, Risto Tuomainen, Liang Wang, Jukka Corander, and Teemu Roos. Fast nearest neighbor search through sparse random projections and voting. In 2016 IEEE International Conference on Big Data (Big Data) ...

  55. [63]

    dropclust: efficient clustering of ultra-large scrna-seq data

    Debajyoti Sinha, Akhilesh Kumar, Himanshu Kumar, Sanghamitra Bandyopadhyay, and Debarka Sengupta. dropclust: efficient clustering of ultra-large scrna-seq data. Nucleic acids research , 46(6):e36--e36, 2018

  56. [64]

    dropclust2: An r package for resource efficient analysis of large scale single cell rna-seq data

    Debajyoti Sinha, Pradyumn Sinha, Ritwik Saha, Sanghamitra Bandyopadhyay, and Debarka Sengupta. dropclust2: An r package for resource efficient analysis of large scale single cell rna-seq data. bioRxiv , page 596924, 2019

  57. [65]

    Cellfishing

    Kenta Sato, Koki Tsuyuzaki, Kentaro Shimizu, and Itoshi Nikaido. Cellfishing. jl: an ultrafast and scalable cell search method for single-cell rna sequencing. Genome biology , 20(1):1--23, 2019

  58. [66]

    The evolving concept of cell identity in the single cell era

    Samantha A Morris. The evolving concept of cell identity in the single cell era. Development , 146(12):dev169748, 2019

  59. [67]

    Defining cell types and states with single-cell genomics

    Cole Trapnell. Defining cell types and states with single-cell genomics. Genome research , 25(10):1491--1498, 2015

  60. [68]

    Cell state transitions: definitions and challenges

    Carla Mulas, Agathe Chaigne, Austin Smith, and Kevin J Chalut. Cell state transitions: definitions and challenges. Development , 148(20):dev199950, 2021

  61. [69]

    Scalable nearest neighbor algorithms for high dimensional data

    Marius Muja and David G Lowe. Scalable nearest neighbor algorithms for high dimensional data. IEEE transactions on pattern analysis and machine intelligence , 36(11):2227--2240, 2014

  62. [70]

    A single-cell transcriptomic map of the human and mouse pancreas reveals inter-and intra-cell population structure

    Maayan Baron, Adrian Veres, Samuel L Wolock, Aubrey L Faust, Renaud Gaujoux, Amedeo Vetere, Jennifer Hyoje Ryu, Bridget K Wagner, Shai S Shen-Orr, Allon M Klein, et al. A single-cell transcriptomic map of the human and mouse pancreas reveals inter-and intra-cell population str...

  63. [71]

    Cell type atlas and lineage tree of a whole complex animal by single-cell transcriptomics

    Mireya Plass, Jordi Solana, F Alexander Wolf, Salah Ayoub, Aristotelis Misios, Petar Gla z ar, Benedikt Obermayer, Fabian J Theis, Christine Kocks, and Nikolaus Rajewsky. Cell type atlas and lineage tree of a whole complex animal by single-cell transcriptomics. Science , 360(6...

  64. [72]

    Comprehensive classification of retinal bipolar neurons by single-cell transcriptomics

    Karthik Shekhar, Sylvain W Lapan, Irene E Whitney, Nicholas M Tran, Evan Z Macosko, Monika Kowalczyk, Xian Adiconis, Joshua Z Levin, James Nemesh, Melissa Goldman, et al. Comprehensive classification of retinal bipolar neurons by single-cell transcriptomics. Cell , 166(5):1308...

  65. [73]

    Deep discrete supervised hashing

    Qing-Yuan Jiang, Xue Cui, and Wu-Jun Li. Deep discrete supervised hashing. IEEE Transactions on Image Processing , 27(12):5996--6009, 2018

  66. [74]

    Supervised hashing with latent factor models

    Peichao Zhang, Wei Zhang, Wu-Jun Li, and Minyi Guo. Supervised hashing with latent factor models. In Proceedings of the 37th international ACM SIGIR conference on Research & development in information retrieval , pages 173--182, 2014

  67. [75]

    Asymmetric deep supervised hashing

    Qing-Yuan Jiang and Wu-Jun Li. Asymmetric deep supervised hashing. In Proceedings of the AAAI conference on artificial intelligence , volume 32, 2018

  68. [76]

    Combi: Compressed binary search tree for approximate k-nn searches in hamming space

    Prashant Gupta, Aashi Jindal, Debarka Sengupta, et al. Combi: Compressed binary search tree for approximate k-nn searches in hamming space. Big Data Research , 25:100223, 2021

  69. [77]

    Ensemble methods in machine learning

    Thomas G Dietterich. Ensemble methods in machine learning. In International workshop on multiple classifier systems , pages 1--15. Springer, 2000

  70. [78]

    Greedy function approximation: a gradient boosting machine

    Jerome H Friedman. Greedy function approximation: a gradient boosting machine. Annals of statistics , pages 1189--1232, 2001

  71. [79]

    Xgboost: A scalable tree boosting system

    Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining , pages 785--794. ACM, 2016

  72. [80]

    Oc1: A randomized algorithm for building oblique decision trees

    Sreerama K Murthy, Simon Kasif, Steven Salzberg, and Richard Beigel. Oc1: A randomized algorithm for building oblique decision trees. In Proceedings of AAAI , volume 93, pages 322--327. Citeseer, 1993

  73. [81]

    A system for induction of oblique decision trees

    Sreerama K Murthy, Simon Kasif, and Steven Salzberg. A system for induction of oblique decision trees. Journal of artificial intelligence research , 2:1--32, 1994

  74. [82]

    On oblique random forests

    Bjoern H Menze, B Michael Kelm, Daniel N Splitthoff, Ullrich Koethe, and Fred A Hamprecht. On oblique random forests. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases , pages 453--469. Springer, 2011

  75. [83]

    Hhcart: An oblique decision tree

    DC Wickramarachchi, BL Robertson, Marco Reale, Christopher John Price, and J Brown. Hhcart: An oblique decision tree. Computational Statistics & Data Analysis , 96:12--23, 2016

  76. [84]

    Guided random forest and its application to data approximation

    Prashant Gupta, Aashi Jindal, Debarka Sengupta, et al. Guided random forest and its application to data approximation. arXiv preprint arXiv:1909.00659 , 2019

  77. [85]

    Classification and regression trees

    Leo Breiman, Jerome Friedman, Charles J Stone, and Richard A Olshen. Classification and regression trees . CRC press, 1984

  78. [86]

    A support vector machine approach to decision trees

    Kristin P Bennett and JA Blue. A support vector machine approach to decision trees. In 1998 IEEE International Joint Conference on Neural Networks Proceedings. IEEE World Congress on Computational Intelligence (Cat. No. 98CH36227) , volume 3, pages 2396--2401. IEEE, 1998

  79. [87]

    A pyramidal delayed perceptron

    G Martinelli, L Prina Ricotti, S Ragazzini, and FM Mascioli. A pyramidal delayed perceptron. IEEE transactions on circuits and systems , 37(9):1176--1181, 1990

  80. [88]

    Binary classification by svm based tree type neural networks

    AK Deb, Suresh Chandra, et al. Binary classification by svm based tree type neural networks. In Proceedings of the 2002 International Joint Conference on Neural Networks. IJCNN'02 (Cat. No. 02CH37290) , volume 3, pages 2773--2778. IEEE, 2002

  81. [89]

    Mml inference of oblique decision trees

    Peter J Tan and David L Dowe. Mml inference of oblique decision trees. In Australasian Joint Conference on Artificial Intelligence , pages 1082--1088. Springer, 2004

  82. [90]

    Decision forests with oblique decision trees

    Peter J Tan and David L Dowe. Decision forests with oblique decision trees. In Mexican International Conference on Artificial Intelligence , pages 593--603. Springer, 2006

  83. [91]

    An information measure for classification

    Chris S Wallace and David M Boulton. An information measure for classification. The Computer Journal , 11(2):185--194, 1968

  84. [92]

    Hybrid extreme point tabu search

    Jennifer A Blue and Kristin P Bennett. Hybrid extreme point tabu search. European journal of operational research , 106(2-3):676--688, 1998

  85. [93]

    Decision-tree-based multiclass support vector machines

    Fumitake Takahashi and Shigeo Abe. Decision-tree-based multiclass support vector machines. In Proceedings of the 9th International Conference on Neural Information Processing, 2002. ICONIP'02. , volume 3, pages 1418--1422. IEEE, 2002

  86. [94]

    An improved algorithm for decision-tree-based svm

    Xiaodan Wang, Zhaohui Shi, Chongming Wu, and Wei Wang. An improved algorithm for decision-tree-based svm. In 2006 6th World Congress on Intelligent Control and Automation , volume 1, pages 4234--4238. IEEE, 2006

  87. [95]

    Geometric decision tree

    Naresh Manwani and PS Sastry. Geometric decision tree. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) , 42(1):181--192, 2011

  88. [96]

    Multisurface proximal support vector machine classification via generalized eigenvalues

    Olvi L Mangasarian and Edward W Wild. Multisurface proximal support vector machine classification via generalized eigenvalues. IEEE transactions on pattern analysis and machine intelligence , 28(1):69--74, 2005

  89. [97]

    Oblique decision tree ensemble via multisurface proximal support vector machine

    Le Zhang and Ponnuthurai N Suganthan. Oblique decision tree ensemble via multisurface proximal support vector machine. IEEE transactions on cybernetics , 45(10):2165--2176, 2014

  90. [98]

    Rotation forest: A new classifier ensemble method

    Juan Jos \'e Rodriguez, Ludmila I Kuncheva, and Carlos J Alonso. Rotation forest: A new classifier ensemble method. IEEE transactions on pattern analysis and machine intelligence , 28(10):1619--1630, 2006

  91. [99]

    An experimental study on rotation forest ensembles

    Ludmila I Kuncheva and Juan J Rodr \' guez. An experimental study on rotation forest ensembles. In International workshop on multiple classifier systems , pages 459--468. Springer, 2007

  92. [100]

    Co2 forest: Improved random forest by continuous optimization of oblique splits

    Mohammad Norouzi, Maxwell D Collins, David J Fleet, and Pushmeet Kohli. Co2 forest: Improved random forest by continuous optimization of oblique splits. arXiv preprint arXiv:1506.06155 , 2015

  93. [101]

    Efficient non-greedy optimization of decision trees

    Mohammad Norouzi, Maxwell Collins, Matthew A Johnson, David J Fleet, and Pushmeet Kohli. Efficient non-greedy optimization of decision trees. In Advances in neural information processing systems , pages 1729--1737, 2015

  94. [102]

    Learning structural svms with latent variables

    Chun-Nam John Yu and Thorsten Joachims. Learning structural svms with latent variables. In Proceedings of the 26th annual international conference on machine learning , pages 1169--1176, 2009

  95. [103]

    Heterogeneous oblique random forest

    Rakesh Katuwal, PN Suganthan, and Le Zhang. Heterogeneous oblique random forest. Pattern Recognition , 99:107078, 2020

  96. [104]

    New algorithms for efficient high-dimensional nonparametric classification

    Ting Liu, Andrew W Moore, and Alexander Gray. New algorithms for efficient high-dimensional nonparametric classification. Journal of Machine Learning Research , 7(Jun):1135--1158, 2006

  97. [105]

    Missforest—non-parametric missing value imputation for mixed-type data

    Daniel J Stekhoven and Peter B \"u hlmann. Missforest—non-parametric missing value imputation for mixed-type data. Bioinformatics , 28(1):112--118, 2012

  98. [106]

    Deep neural decision forests

    Peter Kontschieder, Madalina Fiterau, Antonio Criminisi, and Samuel Rota Bulo. Deep neural decision forests. In Proceedings of the IEEE international conference on computer vision , pages 1467--1475, 2015

  99. [107]

    Neural oblivious decision ensembles for deep learning on tabular data

    Sergei Popov, Stanislav Morozov, and Artem Babenko. Neural oblivious decision ensembles for deep learning on tabular data. arXiv preprint arXiv:1909.06312 , 2019

  100. [108]

    An ensemble of decision trees with random vector functional link networks for multi-class classification

    Rakesh Katuwal, Ponnuthurai N Suganthan, and Le Zhang. An ensemble of decision trees with random vector functional link networks for multi-class classification. Applied Soft Computing , 70:1146--1153, 2018

  101. [109]

    Enhancing multi-class classification of random forest using random vector functional neural network and oblique decision surfaces

    Rakesh Katuwal and Ponnuthurai N Suganthan. Enhancing multi-class classification of random forest using random vector functional neural network and oblique decision surfaces. In 2018 International Joint Conference on Neural Networks (IJCNN) , pages 1--8. IEEE, 2018

  102. [110]

    Boosting the margin: A new explanation for the effectiveness of voting methods

    Robert E Schapire, Yoav Freund, Peter Bartlett, Wee Sun Lee, et al. Boosting the margin: A new explanation for the effectiveness of voting methods. The annals of statistics , 26(5):1651--1686, 1998

  103. [111]

    A decision-theoretic generalization of on-line learning and an application to boosting

    Yoav Freund and Robert E Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of computer and system sciences , 55(1):119--139, 1997

  104. [112]

    Least squares quantization in pcm

    Stuart Lloyd. Least squares quantization in pcm. IEEE transactions on information theory , 28(2):129--137, 1982

  105. [113]

    Some methods for classification and analysis of multivariate observations

    J MacQueen. Some methods for classification and analysis of multivariate observations. In Proc. 5th Berkeley Symposium on Math., Stat., and Prob , page 281, 1965

  106. [114]

    The self-organizing map

    Teuvo Kohonen. The self-organizing map. Proceedings of the IEEE , 78(9):1464--1480, 1990

  107. [115]

    A density-based algorithm for discovering clusters in large spatial databases with noise

    Martin Ester, Hans-Peter Kriegel, J \"o rg Sander, Xiaowei Xu, et al. A density-based algorithm for discovering clusters in large spatial databases with noise. In kdd , volume 96, pages 226--231, 1996

  108. [116]

    Accelerated hierarchical density based clustering

    Leland McInnes and John Healy. Accelerated hierarchical density based clustering. In 2017 IEEE International Conference on Data Mining Workshops (ICDMW) , pages 33--42. IEEE, 2017

  109. [117]

    Learning despite concept variation by finding structure in attribute-based data

    Eduardo Perez and Larry A Rendell. Learning despite concept variation by finding structure in attribute-based data. In In Proceedings of the Thirteenth International Conference on Machine Learning . Citeseer, 1996

  110. [118]

    Decision trees do not generalize to new variations

    Yoshua Bengio, Olivier Delalleau, and Clarence Simard. Decision trees do not generalize to new variations. Computational Intelligence , 26(4):449--467, 2010

  111. [119]

    Weka: Practical machine learning tools and techniques with java implementations

    Rossen Dimov, Michael Feld, Dr Michael Kipp, Dr Alassane Ndiaye, and Dr Dominik Heckmann. Weka: Practical machine learning tools and techniques with java implementations. AI Tools SeminarUniversity of Saarland, WS , 6(07), 2007

  112. [120]

    Bias, variance, and arcing classifiers

    Leo Breiman. Bias, variance, and arcing classifiers. Technical report, Tech. Rep. 460, Statistics Department, University of California, Berkeley …, 1996

  113. [121]

    Bias plus variance decomposition for zero-one loss functions

    Ron Kohavi, David H Wolpert, et al. Bias plus variance decomposition for zero-one loss functions. In ICML , volume 96, pages 275--83, 1996

  114. [122]

    A unified bias-variance decomposition

    Pedro Domingos. A unified bias-variance decomposition. In Proceedings of 17th International Conference on Machine Learning , pages 231--238, 2000

  115. [123]

    Variance and bias for general loss functions

    Gareth M James. Variance and bias for general loss functions. Machine Learning , 51(2):115--135, 2003

  116. [124]

    The random subspace method for constructing decision forests

    Tin Kam Ho. The random subspace method for constructing decision forests. IEEE transactions on pattern analysis and machine intelligence , 20(8):832--844, 1998

  117. [125]

    UCI machine learning repository, 2017

    Dua Dheeru and Efi Karra Taniskidou. UCI machine learning repository, 2017

  118. [126]

    A general introduction to adjustment for multiple comparisons

    Shi-Yi Chen, Zhe Feng, and Xiaolian Yi. A general introduction to adjustment for multiple comparisons. Journal of thoracic disease , 9(6):1725, 2017

  119. [127]

    Practical coreset constructions for machine learning

    Olivier Bachem, Mario Lucic, and Andreas Krause. Practical coreset constructions for machine learning. arXiv preprint arXiv:1703.06476 , 2017

  120. [128]

    Selecting training sets for support vector machines: a review

    Jakub Nalepa and Michal Kawulok. Selecting training sets for support vector machines: a review. Artificial Intelligence Review , 52(2):857--900, 2019

  121. [129]

    Neighborhood property--based pattern selection for support vector machines

    Hyunjung Shin and Sungzoon Cho. Neighborhood property--based pattern selection for support vector machines. Neural Computation , 19(3):816--855, 2007

  122. [130]

    Fast data selection for svm training using ensemble margin

    Li Guo and Samia Boukir. Fast data selection for svm training using ensemble margin. Pattern Recognition Letters , 51:112--119, 2015

  123. [131]

    Feature subset selection using a new definition of classifiability

    Ming Dong and Ravi Kothari. Feature subset selection using a new definition of classifiability. Pattern Recognition Letters , 24(9-10):1215--1225, 2003

  124. [132]

    Classifiability-based omnivariate decision trees

    Yuanhong Li, Ming Dong, and Ravi Kothari. Classifiability-based omnivariate decision trees. IEEE Transactions on Neural Networks , 16(6):1547--1560, 2005

  125. [133]

    Stem cell divisions, somatic mutations, cancer etiology, and cancer prevention

    Cristian Tomasetti, Lu Li, and Bert Vogelstein. Stem cell divisions, somatic mutations, cancer etiology, and cancer prevention. Science , 355(6331):1330--1334, 2017

  126. [134]

    The damaging effect of passenger mutations on cancer progression

    Christopher D McFarland, Julia A Yaglom, Jonathan W Wojtkowiak, Jacob G Scott, David L Morse, Michael Y Sherman, and Leonid A Mirny. The damaging effect of passenger mutations on cancer progression. Cancer research , 77(18):4763--4772, 2017

  127. [135]

    A mutational signature reveals alterations underlying deficient homologous recombination repair in breast cancer

    Paz Polak, Jaegil Kim, Lior Z Braunstein, Rosa Karlic, Nicholas J Haradhavala, Grace Tiao, Daniel Rosebrock, Dimitri Livitz, Kirsten K \"u bler, Kent W Mouw, et al. A mutational signature reveals alterations underlying deficient homologous recombination repair in breast cancer...

  128. [136]

    The repertoire of mutational signatures in human cancer

    Ludmil B Alexandrov, Jaegil Kim, Nicholas J Haradhvala, Mi Ni Huang, Alvin Wei Tian Ng, Yang Wu, Arnoud Boot, Kyle R Covington, Dmitry A Gordenin, Erik N Bergstrom, et al. The repertoire of mutational signatures in human cancer. Nature , 578(7793):94--101, 2020

  129. [137]

    Clustered mutations in yeast and in human cancers can arise from damaged long single-strand dna regions

    Steven A Roberts, Joan Sterling, Cole Thompson, Shawn Harris, Deepak Mav, Ruchir Shah, Leszek J Klimczak, Gregory V Kryukov, Ewa Malc, Piotr A Mieczkowski, et al. Clustered mutations in yeast and in human cancers can arise from damaged long single-strand dna regions. Molecular...

  130. [138]

    An apobec cytidine deaminase mutagenesis pattern is widespread in human cancers

    Steven A Roberts, Michael S Lawrence, Leszek J Klimczak, Sara A Grimm, David Fargo, Petar Stojanov, Adam Kiezun, Gregory V Kryukov, Scott L Carter, Gordon Saksena, et al. An apobec cytidine deaminase mutagenesis pattern is widespread in human cancers. Nature genetics , 45(9):9...

  131. [139]

    Danq: a hybrid convolutional and recurrent deep neural network for quantifying the function of dna sequences

    Daniel Quang and Xiaohui Xie. Danq: a hybrid convolutional and recurrent deep neural network for quantifying the function of dna sequences. Nucleic acids research , 44(11):e107--e107, 2016

  132. [140]

    Predicting effects of noncoding variants with deep learning--based sequence model

    Jian Zhou and Olga G Troyanskaya. Predicting effects of noncoding variants with deep learning--based sequence model. Nature methods , 12(10):931--934, 2015

  133. [141]

    Genomic analyses implicate noncoding de novo variants in congenital heart disease

    Felix Richter, Sarah U Morton, Seong Won Kim, Alexander Kitaygorodsky, Lauren K Wasson, Kathleen M Chen, Jian Zhou, Hongjian Qi, Nihir Patel, Steven R DePalma, et al. Genomic analyses implicate noncoding de novo variants in congenital heart disease. Nature genetics , 52(8):769...

  134. [142]

    Analysis of protein-coding genetic variation in 60,706 humans

    Monkol Lek, Konrad J Karczewski, Eric V Minikel, Kaitlin E Samocha, Eric Banks, Timothy Fennell, Anne H O’Donnell-Luria, James S Ware, Andrew J Hill, Beryl B Cummings, et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature , 536(7616):285--291, 2016

  135. [143]

    Deep learning discerns cancer mutation exclusivity

    Prashant Gupta, Aashi Jindal, Debarka Sengupta, et al. Deep learning discerns cancer mutation exclusivity. bioRxiv , 2020

  136. [144]

    A new deep learning technique reveals the exclusive functional contributions of individual cancer mutations

    Prashant Gupta, Aashi Jindal, Gaurav Ahuja, Debarka Sengupta, et al. A new deep learning technique reveals the exclusive functional contributions of individual cancer mutations. Journal of Biological Chemistry , 298(8), 2022

  137. [145]

    A computational approach to distinguish somatic vs

    James X Sun, Yuting He, Eric Sanford, Meagan Montesion, Garrett M Frampton, St \'e phane Vignot, Jean-Charles Soria, Jeffrey S Ross, Vincent A Miller, Phil J Stephens, et al. A computational approach to distinguish somatic vs. germline origin of genomic alterations from deep s...

  138. [146]

    Variation across 141,456 human exomes and genomes reveals the spectrum of loss-of-function intolerance across human protein-coding genes

    Konrad J Karczewski, Laurent C Francioli, Grace Tiao, Beryl B Cummings, Jessica Alf \"o ldi, Qingbo Wang, Ryan L Collins, Kristen M Laricchia, Andrea Ganna, Daniel P Birnbaum, et al. Variation across 141,456 human exomes and genomes reveals the spectrum of loss-of-function int...

  139. [147]

    dbsnp: the ncbi database of genetic variation

    Stephen T Sherry, M-H Ward, M Kholodov, J Baker, Lon Phan, Elizabeth M Smigielski, and Karl Sirotkin. dbsnp: the ncbi database of genetic variation. Nucleic acids research , 29(1):308--311, 2001

  140. [148]

    Cosmic: the catalogue of somatic mutations in cancer

    John G Tate, Sally Bamford, Harry C Jubb, Zbyslaw Sondka, David M Beare, Nidhi Bindal, Harry Boutselakis, Charlotte G Cole, Celestino Creatore, Elisabeth Dawson, et al. Cosmic: the catalogue of somatic mutations in cancer. Nucleic acids research , 47(D1):D941--D947, 2019

  141. [149]

    The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data, 2012

    Ethan Cerami, Jianjiong Gao, Ugur Dogrusoz, Benjamin E Gross, Selcuk Onur Sumer, B \"u lent Arman Aksoy, Anders Jacobsen, Caitlin J Byrne, Michael L Heuer, Erik Larsson, et al. The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data, 2012

  142. [150]

    Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal

    Jianjiong Gao, B \"u lent Arman Aksoy, Ugur Dogrusoz, Gideon Dresdner, Benjamin Gross, S Onur Sumer, Yichao Sun, Anders Jacobsen, Rileen Sinha, Erik Larsson, et al. Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal. Science signaling , ...

  143. [151]

    Oncokb: a precision oncology knowledge base

    Debyani Chakravarty, Jianjiong Gao, Sarah Phillips, Ritika Kundra, Hongxin Zhang, Jiaojiao Wang, Julia E Rudolph, Rona Yaeger, Tara Soumerai, Moriah H Nissan, et al. Oncokb: a precision oncology knowledge base. JCO precision oncology , 1:1--16, 2017

  144. [152]

    Intogen-mutations identifies cancer drivers across tumor types

    Abel Gonzalez-Perez, Christian Perez-Llamas, Jordi Deu-Pons, David Tamborero, Michael P Schroeder, Alba Jene-Sanz, Alberto Santos, and Nuria Lopez-Bigas. Intogen-mutations identifies cancer drivers across tumor types. Nature methods , 10(11):1081--1082, 2013

  145. [153]

    Cancer genome interpreter annotates the biological and clinical relevance of tumor alterations

    David Tamborero, Carlota Rubio-Perez, Jordi Deu-Pons, Michael P Schroeder, Ana Vivancos, Ana Rovira, Ignasi Tusquets, Joan Albanell, Jordi Rodon, Josep Tabernero, et al. Cancer genome interpreter annotates the biological and clinical relevance of tumor alterations. Genome medi...

  146. [154]

    The human genome browser at ucsc

    W James Kent, Charles W Sugnet, Terrence S Furey, Krishna M Roskin, Tom H Pringle, Alan M Zahler, and David Haussler. The human genome browser at ucsc. Genome research , 12(6):996--1006, 2002

  147. [155]

    The ucsc table browser data retrieval tool

    Donna Karolchik, Angela S Hinrichs, Terrence S Furey, Krishna M Roskin, Charles W Sugnet, David Haussler, and W James Kent. The ucsc table browser data retrieval tool. Nucleic acids research , 32(suppl\_1):D493--D496, 2004

  148. [156]

    The somatic mutation landscape of the human body

    Pablo E Garc \' a-Nieto, Ashby J Morrison, and Hunter B Fraser. The somatic mutation landscape of the human body. Genome Biology , 20(1):1--20, 2019

  149. [157]

    Pan-cancer whole-genome analyses of metastatic solid tumours

    Peter Priestley, Jonathan Baber, Martijn P Lolkema, Neeltje Steeghs, Ewart de Bruijn, Charles Shale, Korneel Duyvesteyn, Susan Haidari, Arne van Hoeck, Wendy Onstenk, et al. Pan-cancer whole-genome analyses of metastatic solid tumours. Nature , 575(7781):210--216, 2019

  150. [158]

    Adam: A method for stochastic optimization

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 , 2014

  151. [159]

    Long short-term memory

    Sepp Hochreiter and J \"u rgen Schmidhuber. Long short-term memory. Neural computation , 9(8):1735--1780, 1997

  152. [160]

    Bidirectional recurrent neural networks

    Mike Schuster and Kuldip K Paliwal. Bidirectional recurrent neural networks. IEEE transactions on Signal Processing , 45(11):2673--2681, 1997

  153. [161]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 , 2015

  154. [162]

    Hierarchical attention networks for document classification

    Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. Hierarchical attention networks for document classification. In Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language techn...

  155. [163]

    dna2vec: Consistent vector representations of variable-length k-mers

    Patrick Ng. dna2vec: Consistent vector representations of variable-length k-mers. arXiv preprint arXiv:1701.06279 , 2017

  156. [164]

    Metascape provides a biologist-oriented resource for the analysis of systems-level datasets

    Yingyao Zhou, Bin Zhou, Lars Pache, Max Chang, Alireza Hadj Khodabakhshi, Olga Tanaseichuk, Christopher Benner, and Sumit K Chanda. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nature communications , 10(1):1--10, 2019

  157. [165]

    Oncotree: a cancer classification system for precision oncology

    Ritika Kundra, Hongxin Zhang, Robert Sheridan, Sahussapont Joseph Sirintrapun, Avery Wang, Angelica Ochoa, Manda Wilson, Benjamin Gross, Yichao Sun, Ramyasree Madupuri, et al. Oncotree: a cancer classification system for precision oncology. JCO Clinical Cancer Informatics , 5:...

  158. [166]

    Determining the optimal number and location of cutoff points with application to data of cervical cancer

    Chung Chang, Meng-Ke Hsieh, Wen-Yi Chang, An Jen Chiang, and Jiabin Chen. Determining the optimal number and location of cutoff points with application to data of cervical cancer. PloS one , 12(4):e0176231, 2017

  159. [167]

    Continuous distributed representation of biological sequences for deep proteomics and genomics

    Ehsaneddin Asgari and Mohammad RK Mofrad. Continuous distributed representation of biological sequences for deep proteomics and genomics. PloS one , 10(11):e0141287, 2015

  160. [168]

    Archetypal organization of the amphioxus hox gene cluster

    Jordi Garcia-Fern \`a ndez and Peter WH Holland. Archetypal organization of the amphioxus hox gene cluster. Nature , 370(6490):563--566, 1994

  161. [169]

    The human olfactory receptor gene family

    Bettina Malnic, Paul A Godfrey, and Linda B Buck. The human olfactory receptor gene family. Proceedings of the National Academy of Sciences , 101(8):2584--2589, 2004

  162. [170]

    Extensive local gene duplication and functional divergence among paralogs in atlantic salmon

    Ian A Warren, Kate L Ciborowski, Elisa Casadei, David G Hazlerigg, Sam Martin, William C Jordan, and Seirian Sumner. Extensive local gene duplication and functional divergence among paralogs in atlantic salmon. Genome biology and evolution , 6(7):1790--1805, 2014

  163. [171]

    Tumor mutational burden quantification from targeted gene panels: major advancements and challenges

    Laura Fancello, Sara Gandini, Pier Giuseppe Pelicci, and Luca Mazzarella. Tumor mutational burden quantification from targeted gene panels: major advancements and challenges. Journal for immunotherapy of cancer , 7(1):1--13, 2019

  164. [172]

    The duchenne muscular dystrophy gene and cancer

    Leanne Jones, Michael Naidoo, Lee R Machado, and Karen Anthony. The duchenne muscular dystrophy gene and cancer. Cellular Oncology , 44(1):19--32, 2021

  165. [173]

    Ribosomal s6 protein kinase 4 promotes radioresistance in esophageal squamous cell carcinoma

    Ming-Yang Li, Lin-Ni Fan, Dong-Hui Han, Zhou Yu, Jing Ma, Yi-Xiong Liu, Pei-Feng Li, Dan-Hui Zhao, Jia Chai, Lei Jiang, et al. Ribosomal s6 protein kinase 4 promotes radioresistance in esophageal squamous cell carcinoma. The Journal of clinical investigation , 130(8):4301--4319, 2020

  166. [174]

    Autophagy promotes primary ciliogenesis by removing ofd1 from centriolar satellites

    Zaiming Tang, Mary Grace Lin, Timothy Richard Stowe, She Chen, Muyuan Zhu, Tim Stearns, Brunella Franco, and Qing Zhong. Autophagy promotes primary ciliogenesis by removing ofd1 from centriolar satellites. Nature , 502(7470):254--257, 2013

  167. [175]

    Primary cilium in cancer hallmarks

    Lucilla Fabbri, Fr \'e d \'e ric Bost, and Nathalie M Mazure. Primary cilium in cancer hallmarks. International journal of molecular sciences , 20(6):1336, 2019

  168. [176]

    Akt regulates a rab11-effector switch required for ciliogenesis

    Vijay Walia, Adrian Cuenca, Melanie Vetter, Christine Insinna, Sumeth Perera, Quanlong Lu, Daniel A Ritt, Elizabeth Semler, Suzanne Specht, Jimmy Stauffer, et al. Akt regulates a rab11-effector switch required for ciliogenesis. Developmental cell , 50(2):229--246, 2019

  169. [177]

    deepdriver: predicting cancer driver genes based on somatic mutations using deep convolutional neural networks

    Ping Luo, Yulian Ding, Xiujuan Lei, and Fang-Xiang Wu. deepdriver: predicting cancer driver genes based on somatic mutations using deep convolutional neural networks. Frontiers in genetics , page 13, 2019

  170. [178]

    Crucial role of the pentose phosphate pathway in malignant tumors

    Lin Jin and Yanhong Zhou. Crucial role of the pentose phosphate pathway in malignant tumors. Oncology letters , 17(5):4213--4221, 2019

  171. [179]

    The pentose phosphate pathway dynamics in cancer and its dependency on intracellular ph

    Khalid O Alfarouk, Samrein Ahmed, Robert L Elliott, Amanda Benoit, Saad S Alqahtani, Muntaser E Ibrahim, Adil HH Bashir, Sari TS Alhoufie, Gamal O Elhassan, Christian C Wales, et al. The pentose phosphate pathway dynamics in cancer and its dependency on intracellular ph. Metab...

  172. [180]

    Lipid metabolism gene-wide profile and survival signature of lung adenocarcinoma

    Jinyou Li, Qiang Li, Zhenyu Su, Qi Sun, Yong Zhao, Tienan Feng, Jiayuan Jiang, Feng Zhang, and Haitao Ma. Lipid metabolism gene-wide profile and survival signature of lung adenocarcinoma. Lipids in health and disease , 19(1):1--9, 2020

  173. [181]

    Revisiting glycogen in cancer: a conspicuous and targetable enabler of malignant transformation

    Tashbib Khan, Mitchell A Sullivan, Jennifer H Gunter, Thomas Kryza, Nicholas Lyons, Yaowu He, and John D Hooper. Revisiting glycogen in cancer: a conspicuous and targetable enabler of malignant transformation. Frontiers in Oncology , 10:2161, 2020

  174. [182]

    Polymorphisms in one-carbon metabolism and trans-sulfuration pathway genes and susceptibility to bladder cancer

    Lee E Moore, N \'u ria Malats, Nathaniel Rothman, Francisco X Real, Manolis Kogevinas, Sara Karami, Reina Garc \' a-Closas, Debra Silverman, Stephen Chanock, Robert Welch, et al. Polymorphisms in one-carbon metabolism and trans-sulfuration pathway genes and susceptibility to b...

  175. [183]

    Metabolic phenotype of bladder cancer

    Francesco Massari, Chiara Ciccarese, Matteo Santoni, Roberto Iacovelli, Roberta Mazzucchelli, Francesco Piva, Marina Scarpelli, Rossana Berardi, Giampaolo Tortora, Antonio Lopez-Beltran, et al. Metabolic phenotype of bladder cancer. Cancer treatment reviews , 45:46--57, 2016

  176. [184]

    One-carbon metabolism in cancer

    Alice C Newman and Oliver DK Maddocks. One-carbon metabolism in cancer. British journal of cancer , 116(12):1499--1504, 2017

  177. [185]

    Histone modifications as markers of cancer prognosis: a cellular view

    SK Kurdistani. Histone modifications as markers of cancer prognosis: a cellular view. British journal of cancer , 97(1):1--5, 2007

  178. [186]

    Telomere maintenance mechanisms in cancer

    Tiago Bordeira Gaspar, Ana S \'a , Jos \'e Manuel Lopes, Manuel Sobrinho-Sim \ o es, Paula Soares, and Jo \ a o Vinagre. Telomere maintenance mechanisms in cancer. Genes , 9(5):241, 2018

  179. [187]

    Network-based stratification of tumor mutations

    Matan Hofree, John P Shen, Hannah Carter, Andrew Gross, and Trey Ideker. Network-based stratification of tumor mutations. Nature methods , 10(11):1108--1115, 2013

  180. [188]

    Etumormetastasis: A network-based algorithm predicts clinical outcomes using whole-exome sequencing data of cancer patients

    Jean-S \'e bastien Milanese, Chabane Tibiche, Naif Zaman, Jinfeng Zou, Pengyong Han, Zhigang Meng, Andre Nantel, Arnaud Droit, and Edwin Wang. Etumormetastasis: A network-based algorithm predicts clinical outcomes using whole-exome sequencing data of cancer patients. Genomics,...

  181. [189]

    Detection of the aberration of x chromosome in hepatocellular carcinoma cell lin e hcc-9903 by fluorescence in situ hybridization (fish)

    Jun Liu, Zhanmin Wang, and Shufang Zheng. Detection of the aberration of x chromosome in hepatocellular carcinoma cell lin e hcc-9903 by fluorescence in situ hybridization (fish). Chinese Journal of General Surgery , 1993

  182. [190]

    Aberration of x chromosome in liver neoplasm detected by fluorescence in situ hybridization

    Jun Liu, Zhan-Min Wang, Shu-Fang Zhen, Xiao-Peng Wu, Dao-Xin Ma, Zhao-Hui Li, Bo Liu, Zhi-Lun Zhao, and Yang Ke. Aberration of x chromosome in liver neoplasm detected by fluorescence in situ hybridization. Hepatobiliary & Pancreatic Diseases International: HBPD INT , 3(1):110-...

  183. [191]

    Significance of kdm6a mutation in bladder cancer immune escape

    Xingxing Chen, Xuehua Lin, Guofu Pang, Jian Deng, Qun Xie, and Zhengrong Zhang. Significance of kdm6a mutation in bladder cancer immune escape. BMC cancer , 21(1):1--6, 2021

  184. [192]

    Mutually exclusive mutation profiles define functionally related genes in muscle invasive bladder cancer

    Ami G Sangster, Robert J Gooding, Andrew Garven, Hamid Ghaedi, David M Berman, and Scott K Davey. Mutually exclusive mutation profiles define functionally related genes in muscle invasive bladder cancer. PloS one , 17(1):e0259992, 2022

  185. [193]

    X chromosome protects against bladder cancer in females via a kdm6a-dependent epigenetic mechanism

    Satoshi Kaneko and Xue Li. X chromosome protects against bladder cancer in females via a kdm6a-dependent epigenetic mechanism. Science advances , 4(6):eaar5598, 2018

  186. [194]

    Clonal origin of bladder cancer

    David Sidransky, Philip Frost, Andy Von Eschenbach, Ryoichi Oyasu, Antonette C Preisinger, and Bert Vogelstein. Clonal origin of bladder cancer. New England Journal of Medicine , 326(11):737--740, 1992

  187. [195]

    Tumor mutation burden in chinese cancer patients and the underlying driving pathways of high tumor mutation burden across different cancer types

    Xiao-Dong Jiao, Xiao-Chun Zhang, Bao-Dong Qin, Dong Liu, Liang Liu, Jian-Jiao Ni, Zhou-Yu Ning, Ling-Xiang Chen, Liang-Jun Zhu, Song-Bing Qin, et al. Tumor mutation burden in chinese cancer patients and the underlying driving pathways of high tumor mutation burden across diffe...

  188. [196]

    Predicting the functional impact of protein mutations: application to cancer genomics

    Boris Reva, Yevgeniy Antipin, and Chris Sander. Predicting the functional impact of protein mutations: application to cancer genomics. Nucleic acids research , 39(17):e118--e118, 2011

  189. [197]

    Evolutionary conservation and somatic mutation hotspot maps of p53: correlation with p53 protein structural and functional features

    D Roland Walker, Jeffrey P Bond, Robert E Tarone, Curtis C Harris, Wojciech Makalowski, Mark S Boguski, and Marc S Greenblatt. Evolutionary conservation and somatic mutation hotspot maps of p53: correlation with p53 protein structural and functional features. Oncogene , 18(1):...

  190. [198]

    Mutation signatures of carcinogen exposure: genome-wide detection and new opportunities for cancer prevention

    Song Ling Poon, John R McPherson, Patrick Tan, Bin Tean Teh, and Steven G Rozen. Mutation signatures of carcinogen exposure: genome-wide detection and new opportunities for cancer prevention. Genome medicine , 6:1--14, 2014

  191. [199]

    Mut2vec: distributed representation of cancerous mutations

    Sunkyu Kim, Heewon Lee, Keonwoo Kim, and Jaewoo Kang. Mut2vec: distributed representation of cancerous mutations. BMC medical genomics , 11(2):33, 2018

  192. [200]

    Block-recurrent transformers

    DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur. Block-recurrent transformers. arXiv preprint arXiv:2203.07852 , 2022

  193. [201]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems , pages 5998--6008, 2017

  194. [202]

    Dnabert: pre-trained bidirectional encoder representations from transformers model for dna-language in genome

    Yanrong Ji, Zhihan Zhou, Han Liu, and Ramana V Davuluri. Dnabert: pre-trained bidirectional encoder representations from transformers model for dna-language in genome. Bioinformatics , 37(15):2112--2120, 2021

  195. [203]

    Protein misfolding and degenerative diseases

    Enrique Reynaud et al. Protein misfolding and degenerative diseases. Nature Education , 3(9):28, 2010

  196. [204]

    Splicing mutations in human genetic disorders: examples, detection, and confirmation

    Abramowicz Anna and Gos Monika. Splicing mutations in human genetic disorders: examples, detection, and confirmation. Journal of applied genetics , 59(3):253--268, 2018

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.