Pith. sign in

REVIEW 3 major objections 4 minor 121 references

Deep Learning, Machine Learning, Advancing Big Data Analytics and Management

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A systematic survey positions the entire big-data analytics workflow as one pipeline running from data collection to deployment, with Python examples at each stage.

desk verdict Survey, not research: broad coverage, no new results, and the tutorial examples have enough internal inconsistencies that I'd desk reject it. read the letter →

arxiv 2412.02187 v1 pith:PUD4NFR7 submitted 2024-12-03 cs.LG

classification cs.LG
keywords bigdataanalyticsmachinelearningdeeppreprocessingclassificationclusteringwarehousingpythonimplementations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a monograph-style survey that tries to establish that the whole of big-data analytics—from raw data collection and warehousing through preprocessing, modeling, evaluation, and real-world applications—can be presented as one coherent, practice-oriented pipeline. The authors aim to bridge theory and practice by pairing standard algorithms and formulas with runnable Python snippets. If successful, the work would serve as an introductory reference that lets researchers, practitioners, and students move from raw data to predictive models without consulting multiple textbooks. The claimed payoff is a single vocabulary and workflow that transfers across healthcare, finance, marketing, and policy domains.

What carries the argument

The organizing device is the analytics pipeline: data collection (surveys, sensors, web scraping, logs, open data), data warehousing with ETL and schema models, preprocessing (cleaning, integration, reduction, sampling), modeling (classification, clustering, regression, anomaly detection, text analytics), evaluation and validation, and finally case studies across healthcare, finance, marketing, and policy. The 5Vs framework defines what counts as big data, the pipeline orders the methods, and the Python snippets act as the claimed executable bridge to practice.

What would settle it

Re-run the code blocks as printed: for example, the date-standardization example in Section 4.1.5 should output exactly the table shown, the bootstrap example in Section 4.5.7 should produce the stated 95% confidence interval, and the pandas missing-data examples should match their displayed frames; a mismatch would falsify the paper's implicit claim that its Python implementations are reliable bridges to practice. A reader can also check whether the unresolved citation markers ("[ ?]") and the truncated sentence in Section 4.5.7 are corrected, since the claim to be a self-contained reference depends on completeness.

Watch

Extended reading notes

Core claim

Stated on the paper's own terms, the central claim is that a systematic overview of AI, machine learning, and deep learning for big data analytics can meaningfully bridge the gap between theoretical foundations and practical implementation. The book's contribution is organizational and pedagogical: it assembles established techniques—5Vs characterization, ETL and warehousing, data cleaning, normalization, dimensionality reduction, sampling, classification, clustering, regression, anomaly detection, text analytics, evaluation, forecasting, recommender systems, and distributed computing—into one sequential narrative. The tone is that of a reference or textbook, not a research monograph, and the intended reader is someone seeking a guided path through the analytics stack rather than a new algorithmic result.

Load-bearing premise

The book's usefulness as a theory-to-practice bridge rests on the assumption that the standard algorithms and formulas it describes are stated correctly and that every printed Python example produces exactly the output shown.

Editorial extensions

If this is right

  • If the overview is accurate, a newcomer can follow the chapters in order to build a working data-analysis workflow from scratch.
  • The book provides a shared vocabulary for teams working across data engineering, statistics, and machine learning, easing coordination on big-data projects.
  • The Python implementations become reusable templates for standard tasks such as handling missing data, selecting features, and training classifiers.
  • The framing of cloud and edge computing as part of the pipeline implies that scalability concerns are integral to analytics, not add-ons.
  • The case-study chapters suggest the same pipeline transfers across domains, from patient-care prediction to fraud detection and customer segmentation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The pipeline structure would make a natural syllabus skeleton for an introductory data-science course, with each chapter paired to its code blocks as lab exercises.
  • A testable extension of the paper's 'bridge to practice' promise is to assemble every Python snippet into a single runnable notebook and verify that each printed output is reproduced exactly.
  • Because the survey's value is pedagogical, its accuracy hinges on the code examples; a companion repository with versioned, tested code would be the concrete artifact the introduction implies.
  • The same pipeline template could be reused for discipline-specific big-data texts by swapping the fish and iris illustrations for domain datasets from healthcare, finance, or public policy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript is an expository survey and tutorial that moves from the 5Vs of big data through data collection, data warehousing, preprocessing, sampling, classification, clustering, frequent-pattern mining, regression, anomaly detection, text analytics, model evaluation, time series, recommender systems, and deep-learning/cloud-based analytics. Its abstract claims a systematic overview that bridges theory and practice through Python-based implementations. The paper presents no new algorithms, datasets, experiments, or falsifiable predictions; its contribution is pedagogical and archival rather than research-oriented.

Significance. If the manuscript were cleaned and completed, it could serve as a broad introductory reference for big-data analytics: the prose is generally readable, the organization is logical, and the algorithmic descriptions are consistent with standard textbook treatments. The many self-contained Python snippets are a useful feature, and the chapter structure would help newcomers navigate the field. However, the current value is undercut by the absence of a bibliography, unresolved placeholder citations, truncated passages, and at least one example whose printed output cannot be produced by the surrounding code. The claim of a 'systematic overview' is therefore not yet supported; the paper is a promising draft rather than a reliable reference.

major comments (3)
  1. [References; §1.2; §2.3; §1.5] The manuscript as submitted has no bibliography, and the text still contains unresolved placeholder citations: '[ ?]' in §1.2 and §2.3, and '[?, 6]' in §1.5. Because the paper's value as a 'systematic overview' depends on being able to verify which sources support each claim, this is not a cosmetic issue. The authors must supply a complete, correctly numbered reference list and remove all placeholders before the manuscript can be considered publishable.
  2. [§4.5.1] The code `df_stratified_sample, _ = train_test_split(df, test_size=0.87, stratify=df['species'], random_state=42)` on the 150-row Iris dataset returns a training DataFrame of 19 rows (150 − ⌈0.87×150⌉ = 19), but the printed output shows 10 rows. The examples are advertised in the abstract as Python-based implementations; a printed output that cannot be produced by the surrounding code means the practical part of the paper is currently unverified. Re-run the example and either correct the output or change the split parameter so that the code, output, and prose are consistent.
  3. [§4.1.5; §4.5.7] The manuscript is incomplete in two visible places: §4.1.5 ends with 'Depending on the type of inconsistency, different techniques such as date parsing' and §4.5.7 ends with 'and the 95' before the printed text cuts off. A survey that claims to equip readers with tools cannot be assessed as a complete manuscript while sentences are truncated; these passages need to be finished.
minor comments (4)
  1. [§3.3.4] The section heading and the table of contents use 'OL TP' with an internal space; this should be 'OLTP'.
  2. [§4.5.3] The systematic sampling code uses `np.random.randint` without setting a random seed, so the printed systematic sample is not reproducible; add a seed or a `random_state` parameter.
  3. [§4.5.4] The cluster sampling examples also draw random clusters without a fixed seed, so the printed 'Selected Cluster' values will vary between runs; either seed the random number generator or note that the output is illustrative.
  4. [Abstract; §3.7.2] The abstract promises attention to 'global standards' for data privacy and compliance, but the text's compliance discussion is limited to GDPR; either expand the coverage or soften the wording.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the manuscript is an expository survey with no derivation chain, fitted parameters, or self-citation load-bearing claims; its internal code/output inconsistencies are reproducibility defects, not circular reasoning.

full rationale

This paper is a textbook-style compilation covering big data concepts, preprocessing, classification, clustering, and Python examples. It makes no quantitative predictions, fits no parameters to data, and tests no hypotheses, so there is no derivation chain in which a claimed output reduces to its own input by construction. The abstract's promise to 'bridge the gap between theory and practice' is a descriptive editorial claim rather than a derived result. The manuscript does cite prior work with bracketed references, but the bibliography is omitted from the supplied text and no load-bearing argument rests on a self-citation or on a uniqueness theorem imported from the authors' own prior work. The visible inconsistencies, such as the stratified random sampling output in Section 4.5.1 showing ten rows where the code with test_size=0.87 on the 150-row Iris dataset would produce a different-sized sample, and the truncated bootstrap output in Section 4.5.7, are internal correctness and completeness issues relevant to the paper's utility as a tutorial, but they are not circular reasoning. Similarly, the missing bibliography and '[?]' citation placeholders in Sections 1.2 and 2.3 are completeness defects rather than evidence that the paper's claims reduce to their own premises. Under the stated rules, absence of derivation warrants a score of 0 rather than a penalty for lack of originality or polish.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

This is a review paper; it introduces no free parameters, postulates no new entities, and depends only on standard background knowledge of machine learning and statistics.

assumptions (2)
  • domain assumption The paper assumes the correctness of the standard descriptions of algorithms (e.g., SVM margin maximization, PCA eigendecomposition) as presented in the surveyed literature.
    Section 5.4 defines SVM as finding the optimal hyperplane and Section 4.3.1 describes PCA steps; these are conventional textbook facts the paper does not prove.
  • domain assumption The paper assumes the Python code snippets, as printed, faithfully reflect the behavior of the pandas and scikit-learn libraries.
    All code examples in Sections 4 and 5 rely on this assumption, but at least one printed output does not exactly match the code (Section 4.1.5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Learning, Machine Learning, Advancing Big Data Analytics and Management." pith.science (2026). https://pith.science/paper/PUD4NFR7

@misc{pith2026241202187,
  author       = {Pith},
  title        = {Pith review of: Deep Learning, Machine Learning, Advancing Big Data Analytics and Management},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PUD4NFR7}},
  note         = {Machine review of arXiv:2412.02187}
}
read the original abstract

Advancements in artificial intelligence, machine learning, and deep learning have catalyzed the transformation of big data analytics and management into pivotal domains for research and application. This work explores the theoretical foundations, methodological advancements, and practical implementations of these technologies, emphasizing their role in uncovering actionable insights from massive, high-dimensional datasets. The study presents a systematic overview of data preprocessing techniques, including data cleaning, normalization, integration, and dimensionality reduction, to prepare raw data for analysis. Core analytics methodologies such as classification, clustering, regression, and anomaly detection are examined, with a focus on algorithmic innovation and scalability. Furthermore, the text delves into state-of-the-art frameworks for data mining and predictive modeling, highlighting the role of neural networks, support vector machines, and ensemble methods in tackling complex analytical challenges. Special emphasis is placed on the convergence of big data with distributed computing paradigms, including cloud and edge computing, to address challenges in storage, computation, and real-time analytics. The integration of ethical considerations, including data privacy and compliance with global standards, ensures a holistic perspective on data management. Practical applications across healthcare, finance, marketing, and policy-making illustrate the real-world impact of these technologies. Through comprehensive case studies and Python-based implementations, this work equips researchers, practitioners, and data enthusiasts with the tools to navigate the complexities of modern data analytics. It bridges the gap between theory and practice, fostering the development of innovative solutions for managing and leveraging data in the era of artificial intelligence.

Figures

Figures reproduced from arXiv: 2412.02187 by the authors.

Figure 1.1
Figure 1.1. The 5Vs of Big Data huge amounts of user data daily, and e-commerce platforms process thousands of transactions every second. These are classic examples of big data [2]. 13 [PITH_FULL_IMAGE:figures/full_fig_p013_1_1.png] view at source ↗
Figure 5.1
Figure 5.1. The Decision surface of decision trees trained [PITH_FULL_IMAGE:figures/full_fig_p092_5_1.png] view at source ↗
Figure 5.2
Figure 5.2. Decision tree trained on all the iris features [PITH_FULL_IMAGE:figures/full_fig_p093_5_2.png] view at source ↗
Figures from the paper (2 more)
Figure 5.3
Figure 5.3. Figure 5.3: Decision tree pruned on all the features [PITH_FULL_IMAGE:figures/full_fig_p094_5_3.png]
Figure 5.4
Figure 5.4. Figure 5.4: Linear SVM Decision Boundary In this example, we train a linear SVM using the first two features of the Iris dataset and visualize [PITH_FULL_IMAGE:figures/full_fig_p100_5_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 75 canonical work pages

  1. [1]

    Mayer-Schönberger and K

    V . Mayer-Schönberger and K. Cukier, Big data: A revolution that will transform how we live, work, and think. Houghton Mifflin Harcourt, 2013

  2. [2]

    Provost, Data Science for Business: What you need to know about data mining and data-analytic thinking, vol

    F . Provost, Data Science for Business: What you need to know about data mining and data-analytic thinking, vol. 355. O’Reilly Media, Inc, 2013

  3. [3]

    Advanced customer analytics: Strategic value through integration of relationship-oriented big data,

    B. Kitchens, D. Dobolyi, J. Li, and A. Abbasi, “Advanced customer analytics: Strategic value through integration of relationship-oriented big data,” Journal of Management Information Sys- tems, vol. 35, no. 2, pp. 540–574, 2018

  4. [4]

    Big data in healthcare: management, analysis and future prospects,

    S. Dash, S. K. Shakyawar, M. Sharma, and S. Kaushik, “Big data in healthcare: management, analysis and future prospects,” Journal of big data , vol. 6, no. 1, pp. 1–25, 2019

  5. [5]

    Designing a smart transportation system: An internet of things and big data approach,

    B. Jan, H. Farman, M. Khan, M. Talha, and I. U. Din, “Designing a smart transportation system: An internet of things and big data approach,” IEEE Wireless Communications, vol. 26, no. 4, pp. 73– 79, 2019

  6. [6]

    Privacy in the age of medical big data,

    W. N. Price and I. G. Cohen, “Privacy in the age of medical big data,” Nature medicine, vol. 25, no. 1, pp. 37–43, 2019

  7. [7]

    Does the gdpr enhance consumersâĂŹ control over personal data? an analysis from a behavioural perspective,

    I. Van Ooijen and H. U. Vrabec, “Does the gdpr enhance consumersâĂŹ control over personal data? an analysis from a behavioural perspective,” Journal of consumer policy , vol. 42, pp. 91– 107, 2019

  8. [8]

    L. M. Rea and R. A. Parker, Designing and conducting survey research: A comprehensive guide . John Wiley & Sons, 2014

Show all 121 references
  1. [9]

    R. M. Groves, F . J. Fowler Jr, M. P . Couper, J. M. Lepkowski, E. Singer, and R. Tourangeau, Survey methodology. John Wiley & Sons, 2011

  2. [10]

    Internet of things: A hands-on approach,

    A. Bahga, “Internet of things: A hands-on approach,” Bahga & Madissetti, 2014

  3. [11]

    Review of agricultural iot technology,

    J. Xu, B. Gu, and G. Tian, “Review of agricultural iot technology,” Artificial Intelligence in Agricul- ture, vol. 6, pp. 10–22, 2022

  4. [12]

    A transaction data study of weekly and intradaily patterns in stock returns,

    L. Harris, “A transaction data study of weekly and intradaily patterns in stock returns,” Journal of financial economics, vol. 16, no. 1, pp. 99–117, 1986

  5. [13]

    Vegetables procurement by asian supermarkets: a transaction cost approach,

    R. Ruben, D. Boselie, and H. Lu, “Vegetables procurement by asian supermarkets: a transaction cost approach,” Supply Chain Management: an international journal , vol. 12, no. 1, pp. 60–68, 2007. 167 168 BIBLIOGRAPHY

  6. [14]

    Social media analytics: a survey of techniques, tools and plat- forms,

    B. Batrinca and P . C. Treleaven, “Social media analytics: a survey of techniques, tools and plat- forms,” Ai & Society, vol. 30, pp. 89–116, 2015

  7. [15]

    The digital architectures of social media: Comparing political campaigning on facebook, twitter, instagram, and snapchat in the 2016 us election,

    M. Bossetta, “The digital architectures of social media: Comparing political campaigning on facebook, twitter, instagram, and snapchat in the 2016 us election,” Journalism & mass commu- nication quarterly, vol. 95, no. 2, pp. 471–496, 2018

  8. [16]

    Schmidt, C

    K. Schmidt, C. Phillips, and A. Chuvakin, Logging and log management: the authoritative guide to understanding the concepts surrounding logging and log management . Newnes, 2012

  9. [17]

    System log clustering approaches for cyber security applications: A survey,

    M. Landauer, F . Skopik, M. Wurzenberger, and A. Rauber, “System log clustering approaches for cyber security applications: A survey,” Computers & Security, vol. 92, p. 101739, 2020

  10. [18]

    Open data now: the secret to hot startups, smart investing, savvy marketing, and fast innovation,

    J. Gurin, “Open data now: the secret to hot startups, smart investing, savvy marketing, and fast innovation,” (No Title), 2014

  11. [19]

    A review of data quality assessment methods for public health information systems,

    H. Chen, D. Hailey, N. Wang, and P . Yu, “A review of data quality assessment methods for public health information systems,” International journal of environmental research and public health , vol. 11, no. 5, pp. 5170–5207, 2014

  12. [20]

    W. H. Inmon, Building the data warehouse . John wiley & sons, 2005

  13. [21]

    Comprehensive survey on data warehousing research,

    P . Chandra and M. K. Gupta, “Comprehensive survey on data warehousing research,” Interna- tional Journal of Information Technology , vol. 10, pp. 217–224, 2018

  14. [22]

    Adamson and M

    C. Adamson and M. Venerable, Data warehouse design solutions. John Wiley & Sons, Inc., 1998

  15. [23]

    Data governance model to enhance data quality in financial institutions,

    S. Karkošková, “Data governance model to enhance data quality in financial institutions,” Infor- mation Systems Management, vol. 40, no. 1, pp. 90–110, 2023

  16. [24]

    Kavis, Architecting the cloud: design decisions for cloud computing service models (SaaS, PaaS, and IaaS)

    M. Kavis, Architecting the cloud: design decisions for cloud computing service models (SaaS, PaaS, and IaaS). Wiley Online Library, 2014

  17. [25]

    Imhoff, N

    C. Imhoff, N. Galemmo, and J. G. Geiger, Mastering data warehouse design: relational and dimen- sional techniques. John Wiley & Sons, 2003

  18. [26]

    Kimball and M

    R. Kimball and M. Ross, The data warehouse toolkit: The definitive guide to dimensional modeling. John Wiley & Sons, 2013

  19. [27]

    An analysis of many-to-many relationships be- tween fact and dimension tables in dimensional modeling,

    W. Rowen, I.-Y . Song, C. Medsker, and E. Ewen, “An analysis of many-to-many relationships be- tween fact and dimension tables in dimensional modeling,” in International Workshop on Design and Management of Data Warehouses (DMDW 2001), Interlaken Switzerland , pp. 1–13, 2001

  20. [28]

    A common database approach for oltp and olap using an in-memory column database,

    H. Plattner, “A common database approach for oltp and olap using an in-memory column database,” in Proceedings of the 2009 ACM SIGMOD International Conference on Management of data, pp. 1–2, 2009

  21. [29]

    Caserta and R

    J. Caserta and R. Kimball, The Data Warehouseetl Toolkit: Practical Techniques for Extracting, Cleaning, Conforming, and Delivering Data . Wiley, 2013

  22. [30]

    The challenges of extract, transform and loading (etl) system implementation for near real- time environment,

    A. Sabtu, N. F . M. Azmi, N. N. A. Sjarif, S. A. Ismail, O. M. Yusop, H. Sarkan, and S. Chuprat, “The challenges of extract, transform and loading (etl) system implementation for near real- time environment,” in 2017 International Conference on Research and Innovation in Infor...

  23. [31]

    O’Reilly Media, Inc

    T . White, Hadoop: The definitive guide . " O’Reilly Media, Inc.", 2012

  24. [32]

    Ellis, Real-time analytics: Techniques to analyze and visualize streaming data

    B. Ellis, Real-time analytics: Techniques to analyze and visualize streaming data . John Wiley & Sons, 2014

  25. [33]

    Sql performance explained,

    M. Winand, “Sql performance explained,” Self-published, Vienna, 2012

  26. [34]

    Database tuning principles, experiments, and trou- bleshooting techniques,

    D. Shasha, P . Bonnet, and N. H. Bercich, “Database tuning principles, experiments, and trou- bleshooting techniques,” ACM SIGMOD Record, vol. 33, no. 2, pp. 115–116, 2004

  27. [35]

    Knowledge warehouse: an architec- tural integration of knowledge management, decision support, artificial intelligence and data warehousing,

    H. R. Nemati, D. M. Steiger, L. S. Iyer, and R. T . Herschel, “Knowledge warehouse: an architec- tural integration of knowledge management, decision support, artificial intelligence and data warehousing,” Decision Support Systems , vol. 33, no. 2, pp. 143–161, 2002

  28. [36]

    Edge computing in industrial inter- net of things: Architecture, advances and challenges,

    T . Qiu, J. Chi, X. Zhou, Z. Ning, M. Atiquzzaman, and D. O. Wu, “Edge computing in industrial inter- net of things: Architecture, advances and challenges,” IEEE Communications Surveys & Tutorials, vol. 22, no. 4, pp. 2462–2488, 2020

  29. [37]

    Big data preprocessing: methods and prospects,

    S. García, S. Ramírez-Gallego, J. Luengo, J. M. Benítez, and F . Herrera, “Big data preprocessing: methods and prospects,” Big data analytics, vol. 1, pp. 1–22, 2016

  30. [38]

    Data cleaning: Problems and current approaches,

    E. Rahm, H. H. Do, et al., “Data cleaning: Problems and current approaches,” IEEE Data Eng. Bull., vol. 23, no. 4, pp. 3–13, 2000

  31. [39]

    Handling missing data in survey research,

    J. M. Brick and G. Kalton, “Handling missing data in survey research,” Statistical methods in medical research, vol. 5, no. 3, pp. 215–238, 1996

  32. [40]

    Correcting noisy data.,

    C. M. Teng, “Correcting noisy data.,” in ICML, vol. 99, pp. 239–248, Citeseer, 1999

  33. [41]

    Handling duplicate data in data warehouse for data mining,

    J. J. Tamilselvi and C. B. Gifta, “Handling duplicate data in data warehouse for data mining,” International Journal of Computer Applications , vol. 15, no. 4, pp. 7–15, 2011

  34. [42]

    Normalization: A preprocessing stage,

    S. Patro, “Normalization: A preprocessing stage,” arXiv preprint arXiv:1503.06462, 2015

  35. [43]

    Big data reduction methods: a survey,

    M. H. ur Rehman, C. S. Liew, A. Abbas, P . P . Jayaraman, T . Y . Wah, and S. U. Khan, “Big data reduction methods: a survey,” Data Science and Engineering , vol. 1, pp. 265–284, 2016

  36. [44]

    Feature selection: A data perspective,

    J. Li, K. Cheng, S. Wang, F . Morstatter, R. P . Trevino, J. Tang, and H. Liu, “Feature selection: A data perspective,” ACM computing surveys (CSUR) , vol. 50, no. 6, pp. 1–45, 2017

  37. [45]

    Dong and H

    G. Dong and H. Liu, Feature engineering for machine learning and data analytics. CRC press, 2018

  38. [46]

    S. L. Lohr, Sampling: design and analysis . Chapman and Hall/CRC, 2021

  39. [47]

    Cluster sampling,

    P . Sedgwick, “Cluster sampling,” Bmj, vol. 348, 2014

  40. [48]

    Convenience sampling,

    P . Sedgwick, “Convenience sampling,” Bmj, vol. 347, 2013

  41. [49]

    Snowball sampling,

    L. A. Goodman, “Snowball sampling,” The annals of mathematical statistics , pp. 148–170, 1961

  42. [50]

    Bootstrap,

    T . Hesterberg, “Bootstrap,” Wiley Interdisciplinary Reviews: Computational Statistics , vol. 3, no. 6, pp. 497–526, 2011

  43. [51]

    Pattern recognition and machine learning,

    C. M. Bishop, “Pattern recognition and machine learning,” Springer google schola , vol. 2, pp. 1122–1128, 2006. 170 BIBLIOGRAPHY

  44. [52]

    Classification assessment methods,

    A. Tharwat, “Classification assessment methods,” Applied computing and informatics , vol. 17, no. 1, pp. 168–192, 2021

  45. [53]

    Decision tree methods: applications for classification and prediction,

    Y .-Y . Song and L. Ying, “Decision tree methods: applications for classification and prediction,” Shanghai archives of psychiatry , vol. 27, no. 2, p. 130, 2015

  46. [54]

    K-nearest neighbor,

    L. E. Peterson, “K-nearest neighbor,” Scholarpedia, vol. 4, no. 2, p. 1883, 2009

  47. [55]

    Support vector machines,

    M. A. Hearst, S. T . Dumais, E. Osuna, J. Platt, and B. Scholkopf, “Support vector machines,” IEEE Intelligent Systems and their applications , vol. 13, no. 4, pp. 18–28, 1998

  48. [56]

    Deep learning and machine learning, advancing big data analytics and management: Handy appetizer,

    B. Peng, X. Pan, Y . Wen, Z. Bi, K. Chen, M. Li, M. Liu, Q. Niu, J. Liu, J. Wang, S. Zhang, J. Xu, and P . Feng, “Deep learning and machine learning, advancing big data analytics and management: Handy appetizer,” 2024

  49. [57]

    Deep learning, machine learning – digital signal and image processing: From theory to application,

    W. Hsieh, Z. Bi, J. Liu, B. Peng, S. Zhang, X. Pan, J. Xu, J. Wang, K. Chen, C. H. Yin, P . Feng, Y . Wen, T . Wang, M. Li, J. Ren, Q. Niu, S. Chen, and M. Liu, “Deep learning, machine learning – digital signal and image processing: From theory to application,” 2024

  50. [58]

    Bayesian classification (autoclass): theory and results.,

    P . C. Cheeseman, J. C. Stutz, et al. , “Bayesian classification (autoclass): theory and results.,” Advances in knowledge discovery and data mining , vol. 180, pp. 153–180, 1996

  51. [59]

    Lazy learning of bayesian rules,

    Z. Zheng and G. I. Webb, “Lazy learning of bayesian rules,” Machine learning, vol. 41, pp. 53–84, 2000

  52. [60]

    Rule-based classification.,

    X. Li and B. Liu, “Rule-based classification.,” 2014

  53. [61]

    Reducing bias through directed acyclic graphs,

    I. Shrier and R. W. Platt, “Reducing bias through directed acyclic graphs,” BMC medical research methodology, vol. 8, pp. 1–15, 2008

  54. [62]

    A tutorial on learning with bayesian networks,

    D. Heckerman, “A tutorial on learning with bayesian networks,” Learning in graphical models , pp. 301–354, 1998

  55. [63]

    Deep learning and machine learning, advancing big data analytics and management: Tensor- flow pretrained models,

    K. Chen, Z. Bi, Q. Niu, J. Liu, B. Peng, S. Zhang, M. Liu, M. Li, X. Pan, J. Xu, J. Wang, and P . Feng, “Deep learning and machine learning, advancing big data analytics and management: Tensor- flow pretrained models,” 2024

  56. [64]

    Backpropagation neural networks: a tutorial,

    B. J. Wythoff, “Backpropagation neural networks: a tutorial,” Chemometrics and Intelligent Lab- oratory Systems, vol. 18, no. 2, pp. 115–155, 1993

  57. [65]

    Sequential covering rule induction algorithm for variable consistency rough set approaches,

    J. Błaszczyński, R. Słowiński, and M. Szeląg, “Sequential covering rule induction algorithm for variable consistency rough set approaches,” Information Sciences, vol. 181, no. 5, pp. 987–1002, 2011

  58. [66]

    Fast effective rule induction,

    W. W. Cohen, “Fast effective rule induction,” in Machine learning proceedings 1995 , pp. 115–123, Elsevier, 1995

  59. [67]

    Data mining: Concepts and techniques,

    W. I. D. Mining, “Data mining: Concepts and techniques,” Morgan Kaufinann, vol. 10, no. 559-569, p. 4, 2006

  60. [68]

    Analysis of k-means and k-medoids algorithm for big data,

    P . Arora, S. Varshney,et al., “Analysis of k-means and k-medoids algorithm for big data,”Procedia Computer Science, vol. 78, pp. 507–512, 2016. BIBLIOGRAPHY 171

  61. [69]

    Density-based clustering,

    H.-P . Kriegel, P . Kröger, J. Sander, and A. Zimek, “Density-based clustering,”Wiley interdisciplinary reviews: data mining and knowledge discovery , vol. 1, no. 3, pp. 231–240, 2011

  62. [70]

    Density-based spatial clustering of applications with noise,

    M. Ester, H.-P . Kriegel, J. Sander, and X. Xu, “Density-based spatial clustering of applications with noise,” in Int. Conf. knowledge discovery and data mining , 1996

  63. [71]

    Optics: Ordering points to identify the clustering structure,

    M. Ankerst, M. M. Breunig, H.-P . Kriegel, and J. Sander, “Optics: Ordering points to identify the clustering structure,” ACM Sigmod record, vol. 28, no. 2, pp. 49–60, 1999

  64. [72]

    Grid-based clustering,

    W. Cheng, W. Wang, and S. Batista, “Grid-based clustering,” in Data clustering, pp. 128–148, Chap- man and Hall/CRC, 2018

  65. [73]

    Clique: Clustering based on density on web usage data: Ex- periments and test results,

    K. Santhisree and A. Damodaram, “Clique: Clustering based on density on web usage data: Ex- periments and test results,” in 2011 3rd International Conference on Electronics Computer Tech- nology, vol. 4, pp. 233–236, IEEE, 2011

  66. [74]

    G. J. McLachlan and T . Krishnan, The EM algorithm and extensions . John Wiley & Sons, 2008

  67. [75]

    Self-organizing maps.,

    M. M. Van Hulle, “Self-organizing maps.,” Handbook of natural computing , vol. 1, pp. 585–622, 2012

  68. [76]

    Subspace clustering for high dimensional data: a review,

    L. Parsons, E. Haque, and H. Liu, “Subspace clustering for high dimensional data: a review,” Acm sigkdd explorations newsletter, vol. 6, no. 1, pp. 90–105, 2004

  69. [77]

    Frequent pattern mining: current status and future direc- tions,

    J. Han, H. Cheng, D. Xin, and X. Yan, “Frequent pattern mining: current status and future direc- tions,” Data mining and knowledge discovery , vol. 15, no. 1, pp. 55–86, 2007

  70. [78]

    The apriori algorithm–a tutorial,

    M. Hegland, “The apriori algorithm–a tutorial,” Mathematics and computation in imaging science and information processing, pp. 209–262, 2007

  71. [79]

    Mining frequent patterns without candidate generation,

    J. Han, J. Pei, and Y . Yin, “Mining frequent patterns without candidate generation,” ACM sigmod record, vol. 29, no. 2, pp. 1–12, 2000

  72. [80]

    Closet: An efficient algorithm for mining frequent closed itemsets.,

    J. Pei, J. Han, R. Mao, et al., “Closet: An efficient algorithm for mining frequent closed itemsets.,” in ACM SIGMOD workshop on research issues in data mining and knowledge discovery , vol. 4, pp. 21–30, 2000

  73. [81]

    Constraint-based pattern mining,

    S. Nijssen and A. Zimmermann, “Constraint-based pattern mining,” Frequent pattern mining , pp. 147–163, 2014

  74. [82]

    Selecting the right interestingness measure for associa- tion patterns,

    P .-N. Tan, V . Kumar, and J. Srivastava, “Selecting the right interestingness measure for associa- tion patterns,” in Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining , pp. 32–41, 2002

  75. [83]

    Hastie, R

    T . Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman, The elements of statistical learning: data mining, inference, and prediction , vol. 2. Springer, 2009

  76. [84]

    Non-linear regression models,

    T . Amemiya, “Non-linear regression models,” Handbook of econometrics , vol. 1, pp. 333–389, 1983

  77. [85]

    Locally weighted regression: an approach to regression analy- sis by local fitting,

    W. S. Cleveland and S. J. Devlin, “Locally weighted regression: an approach to regression analy- sis by local fitting,” Journal of the American statistical association , vol. 83, no. 403, pp. 596–610, 1988. 172 BIBLIOGRAPHY

  78. [86]

    C. C. Aggarwal and C. C. Aggarwal, An introduction to outlier analysis . Springer, 2017

  79. [87]

    A review of local outlier factor algorithms for outlier detection in big data streams,

    O. Alghushairy, R. Alsini, T . Soule, and X. Ma, “A review of local outlier factor algorithms for outlier detection in big data streams,” Big Data and Cognitive Computing , vol. 5, no. 1, p. 1, 2020

  80. [88]

    Statistical fraud detection: A review,

    R. J. Bolton and D. J. Hand, “Statistical fraud detection: A review,” Statistical science , vol. 17, no. 3, pp. 235–255, 2002

  81. [89]

    Network intrusion detection,

    B. Mukherjee, L. T . Heberlein, and K. N. Levitt, “Network intrusion detection,” IEEE network, vol. 8, no. 3, pp. 26–41, 1994

  82. [90]

    Distributed denial of service attacks,

    F . Lau, S. H. Rubin, M. H. Smith, and L. Trajkovic, “Distributed denial of service attacks,” in Smc 2000 conference proceedings. 2000 ieee international conference on systems, man and cybernet- ics.’cybernetics evolving to systems, humans, organizations, and their complex int...

  83. [91]

    Speech and language processing,

    D. Jurafsky, “Speech and language processing,” 2000

  84. [92]

    Understanding bag-of-words model: a statistical framework,

    Y . Zhang, R. Jin, and Z.-H. Zhou, “Understanding bag-of-words model: a statistical framework,” International journal of machine learning and cybernetics , vol. 1, pp. 43–52, 2010

  85. [93]

    Stemming algorithms: A case study for detailed evaluation,

    D. A. Hull, “Stemming algorithms: A case study for detailed evaluation,” Journal of the American Society for Information Science , vol. 47, no. 1, pp. 70–84, 1996

  86. [94]

    A systematic review on stopword removal algorithms,

    J. Kaur and P . K. Buttar, “A systematic review on stopword removal algorithms,” International Journal on Future Revolution in Computer Science & Communication Engineering , vol. 4, no. 4, pp. 207–210, 2018

  87. [95]

    Stemming and lemmatization in the clus- tering of finnish text documents,

    T . Korenius, J. Laurikkala, K. Järvelin, and M. Juhola, “Stemming and lemmatization in the clus- tering of finnish text documents,” in Proceedings of the thirteenth ACM international conference on Information and knowledge management , pp. 625–633, 2004

  88. [96]

    A comparative study of query and document translation for cross-language informa- tion retrieval,

    D. W. Oard, “A comparative study of query and document translation for cross-language informa- tion retrieval,” in Conference of the Association for Machine Translation in the Americas, pp. 472– 483, Springer, 1998

  89. [97]

    Interpreting tf-idf term weights as making relevance decisions,

    H. C. Wu, R. W. P . Luk, K. F . Wong, and K. L. Kwok, “Interpreting tf-idf term weights as making relevance decisions,” ACM Transactions on Information Systems (TOIS) , vol. 26, no. 3, pp. 1–37, 2008

  90. [98]

    Semantic cosine similarity,

    F . Rahutomo, T . Kitasuka, M. Aritsugi, et al., “Semantic cosine similarity,” in The 7th international student conference on advanced science and technology ICAST , vol. 4, p. 1, University of Seoul South Korea, 2012

  91. [99]

    On the evaluation of boolean operators in the extended boolean retrieval framework,

    J. H. Lee, W. Y . Kin, M. H. Kim, and Y . J. Lee, “On the evaluation of boolean operators in the extended boolean retrieval framework,” inProceedings of the 16th annual international ACM SIGIR conference on Research and development in information retrieval , pp. 291–297, 1993

  92. [100]

    Sentiment analysis algorithms and applications: A sur- vey,

    W. Medhat, A. Hassan, and H. Korashy, “Sentiment analysis algorithms and applications: A sur- vey,”Ain Shams engineering journal , vol. 5, no. 4, pp. 1093–1113, 2014

  93. [101]

    Lexicon-based methods for sentiment analysis,

    M. Taboada, J. Brooke, M. Tofiloski, K. Voll, and M. Stede, “Lexicon-based methods for sentiment analysis,” Computational linguistics, vol. 37, no. 2, pp. 267–307, 2011. BIBLIOGRAPHY 173

  94. [102]

    Deep learning and machine learning – object detection and semantic segmentation: From theory to applications,

    J. Ren, Z. Bi, Q. Niu, J. Liu, B. Peng, S. Zhang, X. Pan, J. Wang, K. Chen, C. H. Yin, P . Feng, Y . Wen, T . Wang, S. Chen, M. Li, J. Xu, and M. Liu, “Deep learning and machine learning – object detection and semantic segmentation: From theory to applications,” 2024

  95. [103]

    G. E. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time series analysis: forecasting and control. John Wiley & Sons, 2015

  96. [104]

    P . J. Brockwell and R. A. Davis, Introduction to time series and forecasting . Springer, 2002

  97. [105]

    Jannach, Recommender Systems: An Introduction

    D. Jannach, Recommender Systems: An Introduction . Cambridge University Press, 2010

  98. [106]

    Advances in collaborative filtering,

    Y . Koren, S. Rendle, and R. Bell, “Advances in collaborative filtering,”Recommender systems hand- book, pp. 91–142, 2021

  99. [107]

    Empirical analysis of predictive algorithms for collab- orative filtering,

    J. S. Breese, D. Heckerman, and C. Kadie, “Empirical analysis of predictive algorithms for collab- orative filtering,” arXiv preprint arXiv:1301.7363, 2013

  100. [108]

    Content-based recommender systems: State of the art and trends,

    P . Lops, M. De Gemmis, and G. Semeraro, “Content-based recommender systems: State of the art and trends,” Recommender systems handbook, pp. 73–105, 2011

  101. [109]

    Hybrid recommender systems: Survey and experiments,

    R. Burke, “Hybrid recommender systems: Survey and experiments,” User modeling and user- adapted interaction, vol. 12, pp. 331–370, 2002

  102. [110]

    Deep learning and machine learning, advancing big data analytics and management: Unveiling ai’s potential through tools, techniques, and applications,

    P . Feng, Z. Bi, Y . Wen, X. Pan, B. Peng, M. Liu, J. Xu, K. Chen, J. Liu, C. H. Yin, S. Zhang, J. Wang, Q. Niu, M. Li, and T . Wang, “Deep learning and machine learning, advancing big data analytics and management: Unveiling ai’s potential through tools, techniques, and appli...

  103. [111]

    Deep learning and machine learning with gpgpu and cuda: Unlocking the power of parallel computing,

    M. Li, Z. Bi, T . Wang, Y . Wen, Q. Niu, J. Liu, B. Peng, S. Zhang, X. Pan, J. Xu, J. Wang, K. Chen, C. H. Yin, P . Feng, and M. Liu, “Deep learning and machine learning with gpgpu and cuda: Unlocking the power of parallel computing,” 2024

  104. [112]

    Imagenet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, 2012

  105. [113]

    Recurrent neural networks,

    L. R. Medsker, L. Jain, et al., “Recurrent neural networks,” Design and Applications, vol. 5, no. 64- 67, p. 2, 2001

  106. [114]

    Surveying the mllm landscape: A meta-review of current surveys,

    M. Li, K. Chen, Z. Bi, M. Liu, B. Peng, Q. Niu, J. Liu, J. Wang, S. Zhang, X. Pan, J. Xu, and P . Feng, “Surveying the mllm landscape: A meta-review of current surveys,” 2024

  107. [115]

    Mapreduce: simplified data processing on large clusters,

    J. Dean and S. Ghemawat, “Mapreduce: simplified data processing on large clusters,” Commu- nications of the ACM , vol. 51, no. 1, pp. 107–113, 2008

  108. [116]

    From text to multimodality: Exploring the evolution and impact of large language models in medical practice,

    Q. Niu, K. Chen, M. Li, P . Feng, Z. Bi, L. K. Yan, Y . Zhang, C. H. Yin, C. Fei, J. Liu, and B. Peng, “From text to multimodality: Exploring the evolution and impact of large language models in medical practice,” 2024

  109. [117]

    Personalized medicine.,

    K. K. Jain, “Personalized medicine.,” Current opinion in molecular therapeutics , vol. 4, no. 6, pp. 548–558, 2002

  110. [118]

    Big data in finance: Evidence and challenges,

    A. Subrahmanyam, “Big data in finance: Evidence and challenges,” Borsa Istanbul Review, vol. 19, no. 4, pp. 283–287, 2019. 174 BIBLIOGRAPHY

  111. [119]

    Algorithmic trading,

    G. Nuti, M. Mirghaemi, P . Treleaven, and C. Yingsaeree, “Algorithmic trading,” Computer, vol. 44, no. 11, pp. 61–69, 2011

  112. [120]

    Customer segmentation and strategy develop- ment based on customer lifetime value: A case study,

    S.-Y . Kim, T .-S. Jung, E.-H. Suh, and H.-S. Hwang, “Customer segmentation and strategy develop- ment based on customer lifetime value: A case study,” Expert systems with applications , vol. 31, no. 1, pp. 101–107, 2006

  113. [121]

    Real-time smart traffic management system for smart cities by using internet of things and big data,

    P . Rizwan, K. Suresh, and M. R. Babu, “Real-time smart traffic management system for smart cities by using internet of things and big data,” in 2016 international conference on emerging technological trends (ICETT), pp. 1–7, IEEE, 2016

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.