REVIEW 2 major objections 5 minor 38 references
Performance Reporting of Mathematical Library Installations with LAAB - An Overview
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The LAAB framework ties every performance report of a mathematical-library installation to the exact build, toolchain, and run settings that produced it, and makes the benchmark code retrievable for rerunning.
desk verdict Useful integration of existing HPC benchmarking pieces with a clear objective list, but one of the four stated objectives (compatibility) is not actually backed by any LAAB component described. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a pipeline of components. Build recipes identify the exact installation under test, including version, toolchain, compiler flags, and dependencies. The job-parameterisation tool creates a self-contained run directory for every job, so the executed code and configuration remain retrievable. An inspector component parses the output logs, constructs performance profiles, and applies a partial-ranking method that allows ties when performance distributions overlap. A dashboard renders the profiles as reports, with reusable user-interface components for tables, box plots, and scaling plots, and a transfer component moves the profiles and benchmark source to a serving repository. The partial-ranking-with-ties step is what carries the reliability objective, because it keeps noisy measurements from producing a false strict ordering.
What would settle it
Run a LAAB benchmark whose build recipe and run directory are preserved, then reproduce the same report on a different login node or after a system update; if the median times or the partial ranking of versions change materially, the claim that reports are traceable and reliable fails.
Extended reading notes
Core claim
LAAB's central claim is that a single workflow can produce performance reports that are traceable to the exact library installation and runtime configuration, compatible with evaluation of high-level frameworks built on core linear-algebra libraries, robust to run-to-run noise through ranking with ties, and accessible because benchmark source and metadata can be transferred and rerun. The paper demonstrates the workflow in two experiments: a single-core comparison of double-precision matrix multiplication across several library versions surfaces a recently introduced build that underperforms, and a distributed ScaLAPACK comparison shows that a 512-by-512 block size scales less well than 256 by 256 while the variability in the measurements is large enough that neither block size has a consistent advantage. In both cases the report preserves the ambiguity instead of forcing a clean ranking.
Load-bearing premise
Everything hangs on the assumption that the build recipe and the job run directory capture every setting that affects performance; if the environment adds hidden state, the report cannot be reproduced.
Editorial extensions
If this is right
- A site adopting LAAB can compare a newly deployed library or toolchain version against previous ones and see regressions before users are affected.
- Users can estimate the compute-time contribution of a step in a production workload by multiplying the median execution time from the report by the number of times that step is expected to run.
- Reviewers and allocation committees receive an auditable artifact: report, benchmark source, and run directory are all retrievable, so a reported number can be reproduced or adapted to a different problem size.
- The dashboard's comparison view shows when two configurations are statistically indistinguishable, as in the 256 vs 512 block-size experiment, so decisions are not based on noise.
Reading between the lines
- A natural extension the paper does not pursue is to use the four objectives as a checklist for performance reporting of any managed software stack, not just linear-algebra libraries; the same recipe-to-report pipeline would apply to compilers, MPI implementations, and I/O libraries.
- The compatibility objective connects LAAB to the Linear Algebra Mapping Problem: if a high-level framework expresses a symmetric operation through general matrix multiplication, LAAB-style reports would quantify the lost FLOP/s, which could be used to rank frameworks by how well they map high-level operations to kernels.
- A testable extension would be to turn the reliability objective into an automated gate: declare a regression only when the partial ranks shift beyond the observed noise band, effectively making regression detection a continuous integration check.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents LAAB, a framework for benchmarking and reporting the performance of mathematical library installations on HPC systems. It defines four objectives (traceability, compatibility, reliability, accessibility) and describes a design that combines EasyBuild recipe identification, JUBE-managed benchmark execution, a laab-inspector that computes metrics and applies partial-ranking, dashboard visualizations, and Pigeon for archive transfer. A demonstration on ScaLAPACK distributed DGEMM compares two block sizes and shows that the report captures run-to-run variability. The paper is explicitly an overview and announces future contributions for detailed development guidance and additional use cases.
Significance. If the framework works as intended, it addresses a practical gap: HPC sites rarely report mathematical library performance in a structured, reproducible way. The paper brings together mature components (EasyBuild, JUBE) with the authors' own partial-ranking and dashboard tools, and the open-source availability is a concrete strength. The demo honestly presents variability, which supports the reliability objective. However, the evaluation is limited to a single use case, and the compatibility objective is not yet demonstrated; these limitations affect the strength of the central claim that LAAB systematically addresses all four objectives.
major comments (2)
- [Secs. III.2 and V.2] The compatibility objective is not substantiated by any LAAB-specific component described in the paper. The only mechanisms discussed are the separate maintenance of laab-core and laab-domain-* benchmarks (Sec. V.2) and a reference to prior work [19]; Fig. 2 includes laab-domain-* and dashboard-domain-* boxes without explaining how they implement a compatibility metric or map high-level operations to core-library operations. Consequently, the Abstract's claim that LAAB "systematically addresses" four objectives is only supported for three of them. Please either describe the LAAB-side implementation of compatibility or explicitly scope this objective as future work.
- [Sec. VI] The experimental demonstration covers only traceability (the specification panel in Fig. 3) and reliability (the distribution plot showing run-to-run variation); it does not exercise compatibility, nor does it apply the partial-ranking method that is central to the reliability objective (that method appears only in Fig. 1, which is based on prior work). If the paper's goal is to show how LAAB addresses all four objectives, the evaluation should include at least a minimal example for compatibility, and the reliability demonstration should show the partial-ranking output generated by laab-inspector.
minor comments (5)
- [Sec. V.1] Typo: "definied" should be "defined" in the description of the EasyBuild recipe file name.
- [Sec. VI] The phrase "For demo, let us consider" is awkward; consider "As a demonstration, consider" or similar.
- [Fig. 2] The figure would benefit from a short caption or legend that explicitly identifies the yellow boxes as LAAB components and explains the role of the laab-domain-* and dashboard-domain-* boxes, which are not described in the text.
- [Sec. IV, last paragraph] The statement that "no single existing tool encompasses the proposed workflow in its entirety" is asserted rather than demonstrated; only a handful of tools are mentioned. A brief capability matrix or a more systematic comparison would strengthen this claim, or the wording could be relaxed.
- [Sec. V.3] The description of continuous benchmarking does not specify how many repetitions per job and how many repeated job executions are recommended; adding a suggested minimum would improve reproducibility guidance.
Circularity Check
No significant circularity: LAAB is a framework-integration paper; its self-cited components are used as building blocks, not derived from the claimed objectives.
full rationale
The paper's central claim is that LAAB provides a systematic workflow for performance reporting by integrating EasyBuild recipes, JUBE-managed benchmark execution, partial-ranking-based inspectors, and dashboard visualizations. This is a framework-design claim, not a derived quantitative prediction, and the experimental section is an illustrative self-demonstration rather than a fitted model whose outputs are asserted as predictions. Self-citations appear in the reliability component (partial-ranking method [25]), the compatibility example ([19]), and the software components Tvastar [35] and Pigeon [36]), but in each case the cited work supplies an external component or example; the paper does not derive the validity of those components from the present objectives, nor does it rename a fitted parameter as a prediction. The compatibility objective is admittedly less substantiated within the paper than the other three, but that is a completeness or evidence gap, not a circular reduction: no equation, definition, or result in the paper is equivalent by construction to its own input. The paper is self-contained as an overview of a system design, and no load-bearing step reduces to a self-citation chain or to the framework's own definitions.
Assumptions & free parameters
assumptions (4)
- domain assumption Mathematical libraries form the computational building blocks of scientific applications.
- domain assumption EasyBuild recipes provide a practical and sufficient basis for selecting representative software builds.
- domain assumption Measurement variability is inherent and must be handled statistically.
- domain assumption Users and reviewers have access to the HPC system to re-run benchmarks.
invented entities (1)
-
LAAB framework
independent evidence
Cite this review
Pith. "Pith review of Performance Reporting of Mathematical Library Installations with LAAB - An Overview." pith.science (2026). https://pith.science/paper/WWNQDVST
@misc{pith2026260813512,
author = {Pith},
title = {Pith review of: Performance Reporting of Mathematical Library Installations with LAAB - An Overview},
year = {2026},
howpublished = {\url{https://pith.science/paper/WWNQDVST}},
note = {Machine review of arXiv:2608.13512}
}
read the original abstract
We present the Linear Algebra Aware Benchmarks (LAAB) framework for systematically assessing and reporting the performance of mathematical library installations on HPC systems. Mathematical libraries provide interfaces for operations that form the computational building blocks of scientific applications. Reporting their performance is important for assessing application efficiency, estimating compute-time requirements, and preparing resource-allocation requests. In this paper, we define four objectives for performance reporting: 1) traceability, linking each report to the exact library installation and execution settings; 2) compatibility, relating library-operation performance to higher-level scientific applications that use them; 3) reliability, supporting interpretation in the presence of measurement variability; and 4) accessibility, ensuring that reports, benchmark definitions, and relevant metadata are available for inspection and reproduction. We then present the design of LAAB and show how it addresses the challenges associated with these objectives.
Figures
Reference graph
Works this paper leans on
-
[19]
Benchmarking the Linear Algebra Awareness of TensorFlow and PyTorch,
A. Sankaran, N. A. Alashti, C. Psarras, and P. Bientinesi, “Benchmarking the Linear Algebra Awareness of TensorFlow and PyTorch,” in2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). Lyon, France: IEEE, May 2022, pp. 924–933. [Online]. Available: https://ieeexplore.ieee.org/document/9835658/
-
[25]
Ranking with ties based on noisy performance data,
A. Sankaran, L. Karlsson, and P. Bientinesi, “Ranking with ties based on noisy performance data,”International Journal of Data Science and Analytics, vol. 20, no. 5, pp. 4363–4384, Oct. 2025. [Online]. Available: https://link.springer.com/10.1007/s41060-025-00722-1
-
[1]
Basic Linear Algebra Subprograms for Fortran Usage,
C. L. Lawson, R. J. Hanson, D. R. Kincaid, and F. T. Krogh, “Basic Linear Algebra Subprograms for Fortran Usage,”ACM Transactions on Mathematical Software, vol. 5, no. 3, pp. 308–323, Sep. 1979. [Online]. Available: https://dl.acm.org/doi/10.1145/355841.355847
arXiv 1979
-
[2]
LAPACK: A portable linear algebra library for high- performance computers,
E. Angerson, D. Sorensen, Z. Bai, J. Dongarra, A. Greenbaum, A. McKenney, J. Du Croz, S. Hammarling, J. Demmel, and C. Bischof, “LAPACK: A portable linear algebra library for high- performance computers,” inProceedings SUPERCOMPUTING ’90. New York, NY , USA: IEEE, 1990, pp. 2–11. [Online]. Available: https://ieeexplore.ieee.org/document/129995/
work page 1990
-
[3]
Modern Scientific Software Management Using EasyBuild and Lmod,
M. Geimer, K. Hoste, and R. McLay, “Modern Scientific Software Management Using EasyBuild and Lmod,” in2014 First International Workshop on HPC User Support Tools. New Orleans, LA: IEEE, Nov. 2014, pp. 41–51. [Online]. Available: https://ieeexplore.ieee.org/document/7081225/
-
[5]
EESSI: A cross-platform ready-to-use optimised scientific software stack,
B. Dr ¨oge, V . Holanda Rusu, K. Hoste, C. Van Leeuwen, A. O’Cais, and T. R ¨oblitz, “EESSI: A cross-platform ready-to-use optimised scientific software stack,”Software: Practice and Experience, vol. 53, no. 1, pp. 176–210, Jan. 2023. [Online]. Available: https://onlinelibrary.wiley.com/doi/10.1002/spe.3075
-
[6]
QUANTUM ESPRESSO: a modular and open-source software project for quantum simulations of materials,
P. Giannozzi, S. Baroni, N. Bonini, M. Calandra, R. Car, C. Cavazzoni, D. Ceresoli, G. L. Chiarotti, M. Cococcioni, I. Dabo, A. Dal Corso, S. De Gironcoli, S. Fabris, G. Fratesi, R. Gebauer, U. Gerstmann, C. Gougoussis, A. Kokalj, M. Lazzeri, L. Martin-Samos, N. Marzari, F. Mauri, R. Mazzarello, S. Paolini, A. Pasquarello, L. Paulatto, C. Sbraccia, S. Sca...
2009
-
[7]
PyTorch: an imperative style, high- performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: an imperative style, high- performance deep learning library,” inProceedings of the 33rd Inter- national C...
work page 2019
Show all 38 references
-
[8]
OpenMP: an industry standard API for shared-memory programming,
L. Dagum and R. Menon, “OpenMP: an industry standard API for shared-memory programming,”IEEE Computational Science and Engineering, vol. 5, no. 1, pp. 46–55, Mar. 1998. [Online]. Available: http://ieeexplore.ieee.org/document/660313/
1998
-
[9]
V oss, R
M. V oss, R. Asenjo, and J. Reinders,Pro TBB: C++ Parallel Programming with Threading Building Blocks. Berkeley, CA: Apress,
-
[10]
Open MPI: Goals, Concept, and Design of a Next Generation MPI Implementation,
E. Gabriel, G. E. Fagg, G. Bosilca, T. Angskun, J. J. Dongarra, J. M. Squyres, V . Sahay, P. Kambadur, B. Barrett, A. Lumsdaine, R. H. Castain, D. J. Daniel, R. L. Graham, and T. S. Woodall, “Open MPI: Goals, Concept, and Design of a Next Generation MPI Implementation,” inRece...
2004
-
[11]
Modular Supercomputing Architecture,
E. Suarez, N. Eicker, T. Moschny, S. Pickartz, C. Clauss, V . Plugaru, A. Herten, K. Michielsen, and T. Lippert, “Modular Supercomputing Architecture,” Zenodo, Tech. Rep., May 2022. [Online]. Available: https://zenodo.org/record/6508394
2022
-
[12]
Anatomy of high-performance matrix multiplication,
K. Goto and R. A. V . D. Geijn, “Anatomy of high-performance matrix multiplication,”ACM Transactions on Mathematical Software, vol. 34, no. 3, pp. 1–25, May 2008. [Online]. Available: https://dl.acm.org/doi/10.1145/1356052.1356053
2008
-
[13]
BLIS: A Framework for Rapidly Instantiating BLAS Functionality,
F. G. Van Zee and R. A. Van De Geijn, “BLIS: A Framework for Rapidly Instantiating BLAS Functionality,”ACM Transactions on Mathematical Software, vol. 41, no. 3, pp. 1–33, Jun. 2015. [Online]. Available: https://dl.acm.org/doi/10.1145/2764454
2015 doi
-
[14]
L. S. Blackford, J. Choi, A. Cleary, E. D’Azevedo, J. Demmel, I. Dhillon, J. Dongarra, S. Hammarling, G. Henry, A. Petitet, K. Stanley, D. Walker, and R. C. Whaley,ScaLAPACK Users’ Guide. Society for Industrial and Applied Mathematics, Jan. 1997. [Online]. Available: http://ep...
1997 doi
-
[15]
SLATE: design of a modern distributed and accelerated linear algebra library,
M. Gates, J. Kurzak, A. Charara, A. YarKhan, and J. Dongarra, “SLATE: design of a modern distributed and accelerated linear algebra library,” inProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Denver Colorado: ACM, N...
2019
-
[16]
The Design and Implementation of FFTW3,
M. Frigo and S. Johnson, “The Design and Implementation of FFTW3,” Proceedings of the IEEE, vol. 93, no. 2, pp. 216–231, Feb. 2005. [Online]. Available: http://ieeexplore.ieee.org/document/1386650/
2005
-
[17]
PETSc/TAO Users Manual Revision 3.24,
USDOE Office of Science (SC), Advanced Scientific Computing Research (ASCR), Argonne National Laboratory (ANL), Argonne, IL (United States), S. Balay, S. Abhyankar, M. Adams, J. Brown, P. Brune, K. Buschelman, E. Constantinescu, L. Dalcin, S. Benson, A. Dener, V . Eijkhout, J....
2025
-
[18]
The Linear Algebra Mapping Problem. Current State of Linear Algebra Languages and Libraries,
C. Psarras, H. Barthels, and P. Bientinesi, “The Linear Algebra Mapping Problem. Current State of Linear Algebra Languages and Libraries,” ACM Transactions on Mathematical Software, vol. 48, no. 3, pp. 1–30, Sep. 2022. [Online]. Available: https://dl.acm.org/doi/10.1145/3549935
2022 doi
-
[20]
Reproducible MPI Benchmarking is Still Not as Easy as You Think,
S. Hunold and A. Carpen-Amarie, “Reproducible MPI Benchmarking is Still Not as Easy as You Think,”IEEE Transactions on Parallel and Distributed Systems, vol. 27, no. 12, pp. 3617–3630, Dec. 2016. [Online]. Available: http://ieeexplore.ieee.org/document/7426807/
2016
-
[21]
Scientific benchmarking of parallel computing systems: twelve ways to tell the masses when reporting performance results,
T. Hoefler and R. Belli, “Scientific benchmarking of parallel computing systems: twelve ways to tell the masses when reporting performance results,” inProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Austin Texas: AC...
2015
-
[22]
Influence of Noisy Environments on Behavior of HPC Applications,
D. A. Nikitenko, F. Wolf, B. Mohr, T. Hoefler, K. S. Stefanov, V . V . V oevodin, A. S. Antonov, and A. Calotoiu, “Influence of Noisy Environments on Behavior of HPC Applications,”Lobachevskii Journal of Mathematics, vol. 42, no. 7, pp. 1560–1570, Jul. 2021. [Online]. Availabl...
2021 doi
-
[23]
Statistical Performance Comparisons of Computers,
T. Chen, Q. Guo, O. Temam, Y . Wu, Y . Bao, Z. Xu, and Y . Chen, “Statistical Performance Comparisons of Computers,”IEEE Transactions on Computers, vol. 64, no. 5, pp. 1442–1455, May 2015. [Online]. Available: https://ieeexplore.ieee.org/document/6783811/
2015
-
[24]
Improving high- performance computations on clouds through resource underutilization,
R. Iakymchuk, J. Napper, and P. Bientinesi, “Improving high- performance computations on clouds through resource underutilization,” inProceedings of the 2011 ACM Symposium on Applied Computing. TaiChung Taiwan: ACM, Mar. 2011, pp. 119–126. [Online]. Available: https://dl.acm.o...
2011
-
[26]
Flexible and Generic Workflow Management,
Luhres Sebastian, Rohe Daniel, Schnurpfeil Alexander, Thust Kay, and Frings Wolfgang, “Flexible and Generic Workflow Management,” in Advances in Parallel Computing. IOS Press, 2016
2016
-
[27]
Breuer, S
T. Breuer, S. L ¨uhrs, A. Klasen, and J. Wellmann, “JUBE,” Aug. 2022. [Online]. Available: https://zenodo.org/record/7534373
2022
-
[28]
Enabling Continuous Testing of HPC Systems Using ReFrame,
V . Karakasis, T. Manitaras, V . H. Rusu, R. Sarmiento-P ´erez, C. Big- namini, M. Kraushaar, A. Jocksch, S. Omlin, G. Peretti-Pezzi, J. P. S. C. Augusto, B. Friesen, Y . He, L. Gerhardt, B. Cook, Z.-Q. You, S. Khuvis, and K. Tomko, “Enabling Continuous Testing of HPC Systems ...
2020
-
[29]
Supporting HPC Users with LLview,
F. S. M. Guimar ˜aes, A. Sankaran, and W. Frings, “Supporting HPC Users with LLview,” inHigh Performance Computing, S. Neuwirth, A. K. Paul, T. Weinzierl, and E. C. Carson, Eds. Cham: Springer Nature Switzerland, 2026, pp. 40–51
2026
-
[30]
The Vampir Performance Analysis Tool-Set,
A. Kn ¨upfer, H. Brunst, J. Doleschal, M. Jurenz, M. Lieber, H. Mickler, M. S. M ¨uller, and W. E. Nagel, “The Vampir Performance Analysis Tool-Set,” inTools for High Performance Computing, M. Resch, R. Keller, V . Himmler, B. Krammer, and A. Schulz, Eds. Berlin, Heidelberg: S...
2008 doi
-
[31]
The Scalasca performance toolset architecture,
M. Geimer, F. Wolf, B. J. N. Wylie, E. ´Abrah´am, D. Becker, and B. Mohr, “The Scalasca performance toolset architecture,”Concurrency and Computation: Practice and Experience, vol. 22, no. 6, pp. 702–719, Apr. 2010. [Online]. Available: https://onlinelibrary.wiley.com/doi/10.1...
2010 doi
-
[32]
The Tau Parallel Performance System,
S. S. Shende and A. D. Malony, “The Tau Parallel Performance System,”The International Journal of High Performance Computing Applications, vol. 20, no. 2, pp. 287–311, May 2006. [Online]. Available: https://journals.sagepub.com/doi/10.1177/1094342006064482
2006 doi
-
[33]
SLURM: Simple Linux Utility for Resource Management,
A. B. Yoo, M. A. Jette, and M. Grondona, “SLURM: Simple Linux Utility for Resource Management,” inJob Scheduling Strategies for Parallel Processing, D. Feitelson, L. Rudolph, and U. Schwiegelshohn, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2003, pp. 44–60
2003
-
[34]
Flux: Overcoming Scheduling Challenges for Exascale Workflows,
D. H. Ahn, N. Bass, A. Chu, J. Garlick, M. Grondona, S. Herbein, J. Koning, T. Patki, T. R. W. Scogland, B. Springmeyer, and M. Taufer, “Flux: Overcoming Scheduling Challenges for Exascale Workflows,” in2018 IEEE/ACM Workflows in Support of Large-Scale Science (WORKS). Dallas,...
2018
-
[35]
laab-tvastar,
A. Sankaran, “laab-tvastar,” Aug. 2026. [Online]. Available: https://zenodo.org/doi/10.5281/zenodo.21920713
2026 doi
-
[36]
laab-pigeon,
——, “laab-pigeon,” Aug. 2026. [Online]. Available: https://zenodo.org/doi/10.5281/zenodo.21920065
2026 doi
-
[37]
Pumma: Parallel universal matrix multiplication algorithms on distributed memory concurrent computers,
J. Choi, D. W. Walker, and J. J. Dongarra, “Pumma: Parallel universal matrix multiplication algorithms on distributed memory concurrent computers,”Concurrency: Practice and Experience, vol. 6, no. 7, pp. 543–570, Oct. 1994. [Online]. Available: https://onlinelibrary.wiley.com/...
1994 doi
-
[38]
LAAB-HPC,
A. Sankaran, “LAAB-HPC,” Aug. 2026, language: en. [Online]. Available: https://zenodo.org/doi/10.5281/zenodo.21921182
2026 doi
-
[2019]
Available: https://link.springer.com/10.1007/978-1- 4842-4398-5
[Online]. Available: https://link.springer.com/10.1007/978-1- 4842-4398-5
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.