REVIEW 3 major objections 4 minor 1 cited by
Comparing Parallel Functional Array Languages: Programming and Performance
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Five functional array languages, compiled from a single high-level source to both multicore and GPU, reach 80% or more of hand-optimized OpenMP, Fortran, and CUDA performance on 25 of 36 baseline instances, with codebases at least 2x…
desk verdict A serious, useful five-language benchmark study whose headline performance claims are inflated by counting single-core baseline wins and by a self-written FlashAttention baseline; the underlying Futhark/DaCe results are credible and worth publishing after the claims are softened. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the benchmark comparison protocol: each of the four problems is implemented once per language from a correct starting point, then repeatedly refactored for performance, while the same source is compiled to both multicore and GPU. The underlying mechanism is the data-parallel compilation pipeline, in which second-order array combinators such as map, reduce, scan, and scatter are fused, flattened, and tiled by the compiler; Futhark's incremental flattening, a multi-version compilation technique that maps nested application parallelism onto GPU grid and block levels, and DaCe's graph-based SDFG transformations are the most developed examples. The performance ratios against hand-optimized baselines, together with source-line-of-code counts, are the quantities that carry the argument.
What would settle it
Run the same four benchmarks with baselines that are independently maintained and unmodified public releases, and have the functional implementations written by programmers who did not design the languages under a fixed, realistic time budget; if the fraction of baseline instances where at least one functional language reaches 80% of baseline performance falls below a majority, the central claim would be undermined.
Extended reading notes
Core claim
The central claim is that mature functional array languages can deliver performance competitive with the best available conventional techniques. The evidence is a systematic comparison on 39 instances of four benchmarks: N-body simulation, MultiGrid, Quickhull, and Flash Attention, run on a 32-core AMD EPYC 7313 multicore system and an NVIDIA A30 GPU. Across the 36 instances with hand-optimized baselines, at least one functional language matched or outperformed the baseline in 30% of cases, exceeded 80% of baseline performance in 70% of cases, and only in 2 instances — Quickhull on the multicore CPU — did no language exceed 50% of baseline performance. This is achieved from a single high-level source per language, compiled to both CPU and GPU, with functional codebases at least 10 times smaller than the combined CPU-plus-GPU baseline codebase. The authors argue that the primary reason some languages underperform is incomplete engineering of their backends, not a fundamental limitation of the functional array approach, and they note that only DaCe and Futhark consistently achieve good performance on both architectures.
Load-bearing premise
The comparison assumes the functional implementations and the hand-optimized baselines received comparable and equally expert tuning effort, but the benchmark teams largely wrote their own languages and some baselines were produced by the same group, so unequal effort could explain the performance results.
Editorial extensions
If this is right
- A single functional array source can replace separate CPU and GPU implementations, cutting code by at least 2x versus the CPU baseline and at least 8x versus the GPU baseline.
- Porting a functional array program to new hardware becomes the task of writing a good compiler backend, not rewriting the application.
- The two languages with the most developed backends on both targets, DaCe and Futhark, are the ones that deliver portable performance, suggesting backend quality is the current bottleneck.
- On Quickhull, Futhark's GPU version ran the 100-million-point datasets 2.3x to 4.3x faster than the multicore baseline, showing that at least some irregular divide-and-conquer workloads can be won by flat data-parallel code.
- Because every functional language implemented the simpler Custom Attention rather than the full Flash Attention algorithm, the Flash Attention shortfall is partly an algorithmic choice and adopting the tiled algorithm should close much of the gap.
Reading between the lines
- Beyond the paper: if backend engineering is the main constraint, a follow-up study of the same languages five years later should show CPU-side gaps closing without source changes, which would confirm the portability claim.
- Beyond the paper: the benchmark set implies a minimal primitive checklist for a performant array language — scan or prefix-sum, reduce-by-index, recursion or its unfolding, and tiled matrix multiplication — since languages missing any of these were unable to compete on the corresponding benchmark.
- Beyond the paper: the source-line-of-code evidence suggests that even at half the baseline speed, functional array code may win on total development cost; a direct test would measure end-to-end programmer time, including debugging and autotuning, to first correct and then fast versions.
- Beyond the paper: both test architectures share the same programming model family, so a stronger portability test would compile the same sources to a non-NVIDIA GPU or an FPGA, where the paper's backend-engineering explanation predicts DaCe and Futhark would still lead.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a systematic comparison of five parallel functional array languages (Accelerate, APL, DaCe, Futhark, and SaC) on four benchmarks (N-body, MultiGrid, Quickhull, and Flash Attention), targeting a 32-core AMD EPYC 7313 and an NVIDIA A30 GPU. For each language it describes design and implementation choices, reports SLOC as a proxy for programming effort, and gives performance tables for multicore and GPU execution. The central claim is that mature functional array languages have the potential to deliver performance competitive with the best available conventional techniques, supported mainly by aggregate statistics over 36 benchmark instances and by the observation that a single high-level source replaces separate CPU and GPU baselines.
Significance. If the performance and code-size claims survive scrutiny, this would be a valuable reference for the functional array programming community: it covers five mature implementations, uses two external baseline suites (NAS MG and PBBS Quickhull), publishes an open-source benchmark repository, and honestly reports many negative results (e.g., SaC's missing GPU results for MG and FlashAttention, APL's generally low performance, and the universal FlashAttention shortfall). The design and implementation comparison in Sections 3 and 4 is informative and largely independent of the performance claims. However, the headline statistics in Section 10 are not reproducible from the tables, and the FlashAttention baseline is not the best available conventional technique, so the central conclusion is currently overstated and needs correction before the paper can be accepted.
major comments (3)
- [Section 10, Tables 5-8] The aggregate statistics in Section 10 do not match the data in Tables 5-8. An independent recount of all 36 baseline instances gives 24/36 (67%) with at least one language at or above 80% of baseline, 12/36 (33%) at or above 100%, 10/36 (28%) strictly between 50% and 80%, and 2/36 (6%) at or below 50%. The paper reports 25 (70%), 11 (30%), 9 (25%), and 2 (6%). One concrete discrepancy is the FlashAttention dataset (d=64, N=32768) on 1 CPU core, where Accelerate reaches 150 Gflops against a baseline of 77 Gflops (ratio 1.95): this instance should be counted as matching/outperforming the baseline, but the paper's total of 11 omits it. Please correct the counts and clarify that the intended categories are disjoint.
- [Section 10, Tables 5-8] The headline statistics include the single-core '1C' columns as benchmark instances, but the central claim is about 'the best available conventional techniques' on a machine that provides a 32-core CPU and a GPU. If the 1C instances are excluded, the picture changes materially: only 14 of the 24 remaining parallel instances (58%) have at least one language at or above 80% of baseline, and only 3 (12.5%) match or exceed the baseline. The paper should either justify why single-core runs support a parallel-competitiveness claim or restrict the aggregate statistics to the 32C and GPU columns.
- [Sections 9.2, 9.8, 10] The FlashAttention baseline is not the best available conventional technique. Section 9.2 states that the paper writes its own GPU baseline in single precision using regular FPUs because the official FlashAttention implementation uses half-precision tensor cores, and Sections 9.3-9.7 show that all five functional languages implement the less efficient Custom Attention (Algorithm 6), not FlashAttention. Therefore the 12 FlashAttention instances cannot support the conclusion that mature functional array languages are competitive with the best available conventional techniques; they only support a claim of competitiveness with a self-defined single-precision baseline. The conclusion should be softened accordingly, or the baselines should be replaced by the official implementation for those languages that can target tensor cores. The same caveat applies to the self-written small-n N-body GPU baseline described in Section 6.2.
minor comments (4)
- [Table 9 (referenced as Table 4 in Section 10)] The SLOC totals do not match the row sums: Futhark's total is listed as 448 but its row entries sum to 46+136+161+90=433, and SaC's total of 479 is not reconcilable with the entries shown (61, —, 203, 79). Please audit the table.
- [Section 5.4] The measurement methodology reports averages but the tables do not show variance or the number of runs per cell; given that some results are close to the 80% threshold (for example, Futhark's 547 Gflops versus the N-body n=10^3 GPU baseline of 560 Gflops in Table 5), reporting standard deviations or per-cell run counts would substantially strengthen the comparison.
- [Abstract and Section 10] The abstract and Introduction say the evaluation covers 39 instances, while Section 10 says there are 36 benchmark instances with a baseline; please clarify that the difference is the three Quickhull GPU instances for which no baseline exists.
- [Figure 9] Figure 9 appears four times with subfigures (a)-(d) but has no caption; please add a caption that describes the DaCe optimization workflow illustrated in the four parts.
Circularity Check
No significant circularity: the headline performance claim is supported by direct measurements against external or published baselines, not by a self-referential derivation.
full rationale
This paper is an empirical benchmark and language-design comparison rather than a derivation of performance from language semantics. The central quantitative claim in Section 10 is supported by measured runtimes in Tables 5-8 against external or published baselines: NAS MG uses the Fortran+OpenMP and NPB-GPU implementations, Quickhull uses PBBS, and N-body uses a CUDA sample plus a self-written OpenMP baseline. No equation in the paper defines a predicted quantity in terms of a fitted parameter, and no result is asserted solely because a same-author citation says so. The language self-citations (Futhark, SaC, Accelerate, DaCe, Co-dfns) provide background descriptions of compilers and optimization techniques, but they are not the evidence for the headline performance claim. The main caveat is methodological rather than circular: Section 9.2 states that the FlashAttention GPU baseline was written by the authors using single-precision arithmetic and regular FPUs because the official implementation uses half-precision tensor cores, and Section 9.8 concedes that 'we cannot meaningfully compare most languages with the baseline because it implements the more efficient Flash Attention.' This weakens the phrase 'best available conventional techniques' for those instances, but it is a baseline-suitability and overclaim concern, not a reduction of the conclusion to its inputs. Similarly, authorship overlap between benchmark writers and language teams is a potential bias, but it does not make the measured performance numbers logically dependent on the conclusion. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- N-body dataset pairs (n, t) =
(1e3, 1e5), (1e4, 1e3), (1e5, 1e1)
- FlashAttention dataset pairs (d, N) =
(64, 16384), (64, 32768), (128, 8192), (128, 16384)
- Aggregation threshold 'at least one language' and the 80%/50% levels =
N/A
assumptions (4)
- domain assumption The selected baselines represent the best available conventional techniques.
- domain assumption Implementation effort and expertise are comparable across the five languages.
- domain assumption SLOC is a valid proxy for programming effort and expressiveness.
- domain assumption The four chosen benchmarks are representative of the parallel computational models relevant to functional array languages.
Cite this review
Pith. "Pith review of Comparing Parallel Functional Array Languages: Programming and Performance." pith.science (2026). https://pith.science/paper/ARWX4LZ7
@misc{pith2026250508906,
author = {Pith},
title = {Pith review of: Comparing Parallel Functional Array Languages: Programming and Performance},
year = {2026},
howpublished = {\url{https://pith.science/paper/ARWX4LZ7}},
note = {Machine review of arXiv:2505.08906}
}
read the original abstract
Parallel functional array languages are an emerging class of programming languages that promise to combine low-effort parallel programming with good performance and performance portability. We systematically compare the designs and implementations of five different functional array languages: Accelerate, APL, DaCe, Futhark, and SaC. We demonstrate the expressiveness of functional array programming by means of four challenging benchmarks, namely N-body simulation, MultiGrid, Quickhull, and Flash Attention. These benchmarks represent a range of application domains and parallel computational models. We argue that the functional array code is much shorter and more comprehensible than the hand-optimized baseline implementations because it omits architecture-specific aspects. Instead, the language implementations generate both multicore and GPU executables from a single source code base. Hence, we further argue that functional array code could more easily be ported to, and optimized for, new parallel architectures than conventional implementations of numerical kernels. We demonstrate this potential by reporting the performance of the five parallel functional array languages on a total of 39 instances of the four benchmarks on both a 32-core AMD EPYC 7313 multicore system and on an NVIDIA A30 GPU. We explore in-depth why each language performs well or not so well on each benchmark and architecture. We argue that the results demonstrate that mature functional array languages have the potential to deliver performance competitive with the best available conventional techniques.
Figures
Figures from the paper (14 more)
Forward citations
Cited by 1 Pith paper
-
High-Level Big Integer Arithmetic in Futhark for GPUs
High-level Futhark code for GPU big-integer arithmetic can approach hand-written CUDA performance once arrays are automatically placed in registers.
Reference graph
Works this paper leans on
-
[1]
R. K. W. Hui, M. J. Kromberg, APL since 1978, Proc. ACM Program. Lang. 4 (HOPL) (Jun. 2020).doi:10.1145/3386319
-
[2]
J. Backus, The history of Fortran I, II, and III, Association for Computing Machinery, New York, NY, USA, 1978, Ch. 2, p. 25–74.doi:10.1145/ 800025.1198345
arXiv 1978
- [3]
-
[4]
T. L. McDonell, M. M. Chakravarty, G. Keller, B. Lippmeier, Optimis- ing purely functional GPU programs, in: Proceedings of the 18th ACM SIGPLAN International Conference on Functional Programming, ICFP ’13, Association for Computing Machinery, New York, NY, USA, 2013, p. 49–60. doi:10.1145/2500365.2500595
arXiv 2013
-
[5]
A. N. Ziogas, T. Schneider, T. Ben-Nun, A. Calotoiu, T. De Matteis, J. de Fine Licht, L. Lavarini, T. Hoefler, Productivity, portability, perfor- mance: data-centric Python, in: Proceedings of the International Confer- ence for High Performance Computing, Networking, Storage and Analysis, 13https://github.com/diku-dk/CF AL-bench/ 63 SC ’21, Association fo...
arXiv 2021
-
[6]
T. Henriksen, N. G. W. Serup, M. Elsman, F. Henglein, C. E. Oancea, Futhark: Purely Functional GPU-programming with Nested Parallelism and In-place Array Updates, in: Proceedings of the 38th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2017, ACM, New York, NY, USA, 2017, pp. 556–571. doi:10.1145/ 3062341.3062354
arXiv 2017
-
[7]
C. Grelck, S. Scholz, SAC - A functional array language for efficient multi- threaded execution, Int. J. Parallel Program. 34 (4) (2006) 383–427.doi: 10.1007/S10766-006-0018-X
-
[8]
J. Rosenberg, Some misconceptions about lines of code, in: Proceed- ings fourth international software metrics symposium, IEEE, IEEE, Al- buquerque, NM, USA, 1997, pp. 137–142.doi:10.1109/METRIC.1997. 637174
Show all 104 references
-
[9]
A. K. Hovgaard, T. Henriksen, M. Elsman, High-performance defunction- alisation in futhark, in: M. Pałka, M. Myreen (Eds.), Trends in Func- tional Programming, Springer International Publishing, Cham, 2019, pp. 136–156. doi:10.1007/978-3-030-18506-0_7
2019 doi
-
[10]
G. E. Blelloch, Programming parallel algorithms, Commun. ACM 39 (3) (1996) 85–97. doi:10.1145/227234.227246
1996
-
[11]
T. L. McDonell, J. D. Meredith, G. Keller, Embedded pattern matching, in: Proceedings of the 15th ACM SIGPLAN International Haskell Sym- posium, Haskell 2022, Association for Computing Machinery, New York, NY, USA, 2022, p. 123–136.doi:10.1145/3546189.3549917
2022
-
[12]
Grelck, S.-B
C. Grelck, S.-B. Scholz, Classes and objects as basis for i/o in sac, in: T. Johnsson (Ed.), 7th International Workshop on Implementation of Functional Languages (IFL’95), Båstad, Sweden, Chalmers University of Technology, Gothenburg, Sweden, 1995, pp. 30–44. URL sac-classes-o...
1995
-
[13]
Huijben, J
R. Huijben, J. Aaldering, P. Achten, S.-B. Scholz, Flattening combina- tions of arrays and records, in: J. Hemann, S. Chang (Eds.), Trends in Functional Programming, Springer Nature Switzerland, Cham, 2025, pp. 220–240. doi:10.1007/978-3-031-74558-4_10
2025 doi
-
[14]
Calotoiu, T
A. Calotoiu, T. Ben-Nun, G. Kwasniewski, J. de Fine Licht, T. Schneider, P. Schaad, T. Hoefler, Lifting C semantics for dataflow optimization, in: Proceedings of the 36th ACM International Conference on Supercomput- ing, ICS ’22, Association for Computing Machinery, New York, ...
2022
-
[16]
URL https://spcldace.readthedocs.io
Scalable Parallel Computing Laboratory, ETH Zurich, DaCe: Data- Centric Parallel Programming (2024). URL https://spcldace.readthedocs.io
2024
-
[17]
Gonnord, L
L. Gonnord, L. Henrio, L. Morel, G. Radanne, A Survey on Parallelism and Determinism, ACM Comput. Surv. 55 (10) (Feb. 2023). doi:10. 1145/3564529
2023
-
[18]
1–14.doi:10.1109/SC41405.2020.00101
T.Henriksen, S.Hellfritzsch, P.Sadayappan, C.Oancea, Compiling gener- alized histograms for gpu, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’20, IEEE Press, 2020, pp. 1–14.doi:10.1109/SC41405.2020.00101
2020 arXiv
-
[19]
Wadler, Deforestation: transforming programs to eliminate trees, Theoretical Computer Science 73 (2) (1990) 231–248
P. Wadler, Deforestation: transforming programs to eliminate trees, Theoretical Computer Science 73 (2) (1990) 231–248. doi:10.1016/ 0304-3975(90)90147-A
1990
-
[20]
G. E. Blelloch, Vector models for data-parallel computing, Vol. 75, MIT press Cambridge, 1990
1990
-
[21]
Henriksen, F
T. Henriksen, F. Thorøe, M. Elsman, C. Oancea, Incremental Flattening for Nested Data Parallelism, in: Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming, PPoPP ’19, ACM, New York, NY, USA, 2019, pp. 53–67.doi:10.1145/3293883.3295707
2019
-
[22]
Munksgaard, S
P. Munksgaard, S. L. Breddam, T. Henriksen, F. C. Gieseke, C. Oancea, Dataset sensitive autotuning of multi-versioned code based on mono- tonic properties, in: V. Zsók, J. Hughes (Eds.), Trends in Functional Programming, Springer International Publishing, Cham, 2021, pp. 3–23....
2021 doi
-
[23]
Schenck, O
R. Schenck, O. Rønning, T. Henriksen, C. E. Oancea, Ad for an array language with nested parallelism, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’22, IEEE Press, 2022, pp. 1–15.doi:10.1109/SC41404. 2022.00063
2022
-
[24]
Munksgaard, T
P. Munksgaard, T. Henriksen, P. Sadayappan, C. Oancea, Memory op- timizations in an array language, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’22, IEEE Press, 2022, pp. 1–15.doi:10.1109/SC41404. 2022.00036. 65
2022
-
[25]
A.Nicolaisen, M.A. Persson, Implementingsingle-pass scaninthe futhark compiler, Master’s thesis, Department of Computer Science, Faculty of Science, University of Copenhagen, https://futhark-lang.org/student- projects/marco-andreas-scan.pdf (2020). URL https://futhark-lang.org...
2020
-
[26]
M. T. Clausen, Regular segmented single-pass scan in futhark, Master’s thesis, Department of Computer Science, Faculty of Science, University of Copenhagen, https://futhark-lang.org/student-projects/morten-msc- thesis.pdf (2021). URL https://futhark-lang.org/student-projects/ ...
2021
-
[27]
Merrill, M
D. Merrill, M. Garland, Single-pass parallel prefix scan with decoupled lookback, Nvidia technical report nvr-2016-002, march 2016, NVIDIA (2016). URL https://research.nvidia.com/sites/default/files/pubs/ 2016-03_Single-pass-Parallel-Prefix/nvr-2016-002.pdf
2016
-
[28]
T.Schrijvers, S.PeytonJones, M.Chakravarty, M.Sulzmann, Typecheck- ing with open type functions, in: Proceedings of the 13th ACM SIGPLAN International Conference on Functional Programming, ICFP ’08, Associ- ation for Computing Machinery, New York, NY, USA, 2008, p. 51–62. doi:...
2008
-
[29]
M. M. T. Chakravarty, G. Keller, S. P. Jones, Associated type synonyms, in: Proceedings of the Tenth ACM SIGPLAN International Conference on Functional Programming, ICFP ’05, Association for Computing Ma- chinery, New York, NY, USA, 2005, p. 241–253.doi:10.1145/1086365. 1086397
2005 doi
-
[30]
Grelck, S.-B
C. Grelck, S.-B. Scholz, Merging compositions of array skeletons in SAC, Journal of Parallel Computing 32 (7+8) (2006) 507–522.doi:10.1016/ j.parco.2006.08.003
2006
-
[31]
Grelck, K
C. Grelck, K. Trojahner, Implicit Memory Management for SaC, in: C. Grelck, F. Huch (Eds.), Implementation and Application of Functional Languages, 16th International Workshop, IFL’04, University of Kiel, In- stitute of Computer Science and Applied Mathematics, 2004, pp. 335–3...
2004
-
[32]
Grelck, Single Assignment C (SAC): the compilation technology per- spective, in: V
C. Grelck, Single Assignment C (SAC): the compilation technology per- spective, in: V. Zsók, Z. Porkoláb, Z. Horváth (Eds.), 6th Central Eu- ropean Functional Programming Summer School (CEFP’15), Budapest, Hungary, Vol. 10094 of Lecture Notes in Computer Science, Springer, 201...
2019 doi
-
[33]
A. W.-y. Hsu, A data parallel compiler hosted on the GPU, Ph.D. thesis, Indiana University (2019). URL https://hdl.handle.net/2022/24749
2019
-
[34]
A. W. Keep, A nanopass framework for commercial compiler development, Ph.D. thesis, Indiana University (2012). URL https://andykeep.com/pubs/dissertation.pdf
2012
-
[35]
Malcolm, P
J. Malcolm, P. Yalamanchili, C. McClanahan, V. Venugopalakrishnan, K. Patel, J. Melonakos, ArrayFire: a GPU acceleration platform, in: E. J. Kelmelis (Ed.), Modeling and Simulation for Defense Systems and Appli- cations VII, Vol. 8403, International Society for Optics and Phot...
2012 doi
-
[36]
A. N. Ziogas, T. Ben-Nun, G. I. Fernández, T. Schneider, M. Luisier, T. Hoefler, A Data-Centric Approach to Extreme-Scale Ab initio Dissipa- tive Quantum Transport Simulations, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage ...
2019
-
[37]
Ben-Nun, L
T. Ben-Nun, L. Groner, F. Deconinck, T. Wicky, E. Davis, J. Dahm, O. Elbert, R. George, J. McGibbon, L. Trümper, E. Wu, O. Fuhrer, T. Schulthess, T. Hoefler, Productive Performance Engineering for Weather and Climate Modeling with Python, in: Proceedings of the Inter- national...
2022 arXiv
-
[38]
Schaad, T
P. Schaad, T. Ben-Nun, T. Hoefler, Boosting Performance Optimization with Interactive Data Movement Visualization, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC’22), 2022, pp. 1–16.doi:10.1109/SC41404. 2022.00069
2022
-
[39]
A. N. Ziogas, G. Kwasniewski, T. Ben-Nun, T. Schneider, T. Hoe- fler, Deinsum: Practically I/O Optimal Multilinear Algebra, in: Pro- ceedings of the International Conference for High Performance Comput- ing, Networking, Storage and Analysis (SC’22), 2022, pp. 1–15. doi: 10.555...
2022
-
[40]
Trümper, T
L. Trümper, T. Ben-Nun, P. Schaad, A. Calotoiu, T. Hoefler, Performance Embeddings: A Similarity-Based Transfer Tuning Approach to Perfor- mance Optimization, in: Proceedings of the 37th International Conference on Supercomputing, ICS ’23, Association for Computing Machinery, ...
2023
-
[41]
D. H. Bailey, E. Barszcz, J. T. Barton, D. S. Browning, R. L. Carter, L. Dagum, R. A. Fatoohi, P. O. Frederickson, T. A. Lasinski, R. S. Schreiber, H. D. Simon, V. Venkatakrishnan, S. K. Weeratunga, The NAS 67 parallel benchmarks—summary and preliminary results, in: Proceeding...
1991
-
[42]
G.Araujo, D.Griebler, D.A.Rockenbach, M.Danelutto, L.G.Fernandes, NAS Parallel Benchmarks with CUDA and beyond, Software: Practice and Experience 53 (1) (2023) 53–80.doi:10.1002/spe.3056
2023 doi
-
[43]
Anderson, G
D. Anderson, G. E. Blelloch, L. Dhulipala, M. Dobson, Y. Sun, The problem-based benchmark suite (pbbs), v2, in: Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Pro- gramming, PPoPP ’22, Association for Computing Machinery, New York, NY, USA...
2022
-
[44]
T. Dao, D. Fu, S. Ermon, A. Rudra, C. Ré, FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, in: S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, A. Oh (Eds.), Advances in Neural Information Processing Systems, Vol. 35, Curran Associates, Inc.,...
2022
-
[45]
Bailey, T
D. Bailey, T. Harris, W. Saphir, R. Van Der Wijngaart, A. Woo, M. Yarrow, The nas parallel benchmarks 2.0, Tech. rep., Technical Re- port NAS-95-020, NASA Ames Research Center (1995)
1995
-
[46]
Di Domenico, G
D. Di Domenico, G. G. H. Cavalheiro, J. V. F. Lima, NAS Parallel Bench- mark Kernels with Python: A performance and programming effort anal- ysis focusing on GPUs, in: 2022 30th Euromicro International Conference on Parallel, Distributed and Network-based Processing (PDP), 202...
2022
-
[47]
M. M. T. Chakravarty, R. Leshchinskiy, S. Peyton Jones, G. Keller, S. Marlow, Data parallel Haskell: a status report, in: Proceedings of the 2007 Workshop on Declarative Aspects of Multicore Programming, DAMP ’07, Association for Computing Machinery, New York, NY, USA, 2007, p...
2007
-
[48]
Zhang, F
Y. Zhang, F. Mueller, CuNesl: Compiling nested data-parallel languages for SIMT architectures, in: Proceedings of the 2012 41st International Conference on Parallel Processing, ICPP’12, IEEE Computer Society, Washington, DC, USA, 2012, pp. 340–349.doi:10.1109/ICPP.2012.21
2012 doi
-
[49]
Bergstrom, J
L. Bergstrom, J. Reppy, Nested data-parallelism on the gpu, in: Proceed- ings of the 17th ACM SIGPLAN International Conference on Functional Programming, ICFP ’12, ACM, New York, NY, USA, 2012, pp. 247–258. doi:10.1145/2364527.2364563. 68
2012
-
[50]
Edelkamp, A
S. Edelkamp, A. Weiß, Blockquicksort: Avoiding branch mispredictions in quicksort, ACM J. Exp. Algorithmics 24 (Jan. 2019).doi:10.1145/ 3274660
2019
-
[51]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin, Attention is all you need, in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems,...
2017
-
[52]
J. S. Bridle, Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition, in: F. F. Soulié, J. Hérault (Eds.), Neurocomputing, Springer Berlin Heidelberg, Berlin, Heidelberg, 1990, pp. 227–236. doi:10.1007/...
1990
-
[53]
Milakov, N
M. Milakov, N. Gimelshein, Online normalizer calculation for softmax (2018). arXiv:1805.02867. URL https://arxiv.org/abs/1805.02867
2018 arXiv
-
[54]
Dao AI Lab, FlashAttention, https://github.com/Dao-AILab/ flash-attention (2024)
2024
-
[55]
Henriksen, C
T. Henriksen, C. E. Oancea, A T2 graph-reduction approach to fusion, in: Proceedings of the 2Nd ACM SIGPLAN Workshop on Functional High- performance Computing, FHPC ’13, ACM, New York, NY, USA, 2013, pp. 47–58. doi:10.1145/2502323.2502328
2013
-
[56]
Henriksen, K
T. Henriksen, K. F. Larsen, C. E. Oancea, Design and GPGPU perfor- mance of futhark’s redomap construct, in: Proceedings of the 3rd ACM SIGPLAN International Workshop on Libraries, Languages, and Compil- ers for Array Programming, ARRAY 2016, ACM, New York, NY, USA, 2016, pp. ...
2016
-
[57]
D. W. Walker, J. J. Dongarra, MPI: a standard message passing interface, Supercomputer 12 (1996) 56–68.doi:10.1145/169627.169855
1996
-
[58]
P. W. Trinder, K. Hammond, H.-W. Loidl, S. P. Jones, Algorithm+ strat- egy= parallelism, Journal of functional programming 8 (1) (1998) 23–60. doi:10.1017/S0956796897002967
1998 doi
-
[59]
C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Vir- tanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Hal- dane, J. F. del Río, M. Wiebe, P. Peterson, P. Gérard-Marchant, K. Shep- pa...
2020 doi
-
[60]
Sidelnik, S
A. Sidelnik, S. Maleki, B. L. Chamberlain, M. J. Garzar’n, D. Padua, Performance portability with the chapel language, in: 2012 IEEE 26th international parallel and distributed processing symposium, IEEE, 2012, pp. 582–594. doi:10.1109/IPDPS.2012.60
2012 doi
-
[61]
L. V. Kale, S. Krishnan, CHARM++: a portable concurrent object ori- ented system based on C++, in: Proceedings of the Eighth Annual Con- ference on Object-Oriented Programming Systems, Languages, and Appli- cations, OOPSLA ’93, Association for Computing Machinery, New York, NY...
1993
-
[62]
Rice University, High performance fortran language specification, SIG- PLAN Fortran Forum 12 (4) (1993) 1–86.doi:10.1145/174223.158909
C. Rice University, High performance fortran language specification, SIG- PLAN Fortran Forum 12 (4) (1993) 1–86.doi:10.1145/174223.158909
1993
-
[63]
von Praun, V
P.Charles, C.Grothoff, V.Saraswat, C.Donawa, A.Kielstra, K.Ebcioglu, C. von Praun, V. Sarkar, X10: an object-oriented approach to non- uniform cluster computing, SIGPLAN Not. 40 (10) (2005) 519–538. doi:10.1145/1103845.1094852
2005
-
[64]
Chandra, Parallel programming in OpenMP, Morgan Kaufmann, 2001
R. Chandra, Parallel programming in OpenMP, Morgan Kaufmann, 2001
2001
-
[65]
Sanders, E
J. Sanders, E. Kandrot, CUDA by example: an introduction to general- purpose GPU programming, Addison-Wesley Professional, 2010
2010
-
[66]
Munshi, The OpenCL specification, in: 2009 IEEE Hot Chips 21 Symposium (HCS), 2009, pp
A. Munshi, The OpenCL specification, in: 2009 IEEE Hot Chips 21 Symposium (HCS), 2009, pp. 1–314. doi:10.1109/HOTCHIPS.2009. 7478342
2009 doi
-
[67]
Bauer, M
M. Bauer, M. Garland, Legate NumPy: Accelerated and Distributed Ar- ray Computing, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’19, As- sociation for Computing Machinery, New York, NY, USA, 2019, pp. 1–23...
2019
-
[68]
Okuta, Y
R. Okuta, Y. Unno, D. Nishino, S. Hido, C. Loomis, CuPy: A NumPy- Compatible Library for NVIDIA GPU Calculations (2017). URL http://learningsys.org/nips17/assets/papers/paper_16. pdf
2017
-
[69]
Carter Edwards, C
H. Carter Edwards, C. R. Trott, D. Sunderland, Kokkos: Enabling manycore performance portability through polymorphic memory ac- cess patterns, Journal of Parallel and Distributed Computing 74 (12) (2014)3202–3216, domain-SpecificLanguagesandHigh-LevelFrameworks for High-Perfor...
2014 doi
-
[70]
URL https://github.com/LLNL/RAJA 70
LLNL, RAJA Performance Portability Layer (2019). URL https://github.com/LLNL/RAJA 70
2019
-
[71]
Kotsifakou, P
M. Kotsifakou, P. Srivastava, M. D. Sinclair, R. Komuravelli, V. Adve, S. Adve, HPVM: heterogeneous parallel virtual machine, SIGPLAN Not. 53 (1) (2018) 68–80.doi:10.1145/3200691.3178493
2018
-
[72]
Bauer, S
M. Bauer, S. Treichler, E. Slaughter, A. Aiken, Legion: expressing locality andindependencewithlogicalregions, in: ProceedingsoftheInternational Conference on High Performance Computing, Networking, Storage and Analysis, SC ’12, IEEE Computer Society Press, Washington, DC, USA...
2012
-
[73]
Pouchet, U
L.-N. Pouchet, U. Bondhugula, C. Bastoul, A. Cohen, J. Ramanujam, P. Sadayappan, N. Vasilache, Loop transformations: Convexity, pruning and optimization, in: Proceedings of the 38th Annual ACM SIGPLAN- SIGACT Symposium on Principles of Programming Languages, POPL ’11, ACM, New...
2011
-
[74]
Bondhugula, A
U. Bondhugula, A. Hartono, J. Ramanujam, P. Sadayappan, A practical automatic polyhedral parallelizer and locality optimizer, in: Proceedings of the 29th ACM SIGPLAN Conference on Programming Language De- sign and Implementation, PLDI ’08, ACM, New York, NY, USA, 2008, pp. 101...
2008
-
[75]
Verdoolaege, J
S. Verdoolaege, J. Carlos Juega, A. Cohen, J. Ignacio Gómez, C. Ten- llado, F. Catthoor, Polyhedral parallel code generation for cuda, ACM TransactionsonArchitectureandCodeOptimization(TACO)9(4)(2013) 54:1–54:23. doi:10.1145/2400682.2400713
2013
-
[76]
T.Grosser, A.Größlinger, C.Lengauer, Polly-PerformingPolyhedralOp- timizations on a Low-Level Intermediate Representation, Parallel Process- ing Letters 22 (04) (2012) 1250010.doi:10.1142/S0129626412500107
2012 doi
-
[77]
Baghdadi, J
R. Baghdadi, J. Ray, M. B. Romdhane, E. Del Sozzo, A. Akkas, Y. Zhang, P.Suriana, S.Kamil, S.Amarasinghe, Tiramisu: apolyhedralcompiler for expressing fast and portable code, in: Proceedings of the 2019 IEEE/ACM International Symposium on Code Generation and Optimization, CGO ...
2019
-
[78]
M. M. Strout, Performance transformations for irregular applications, Ph.D. thesis, University of California, aAI3094622 (2003)
2003
-
[79]
M. M. Strout, A. LaMielle, L. Carter, J. Ferrante, B. Kreaseck, C. Olschanowsky, An approach for code generation in the sparse poly- hedral framework, Parallel Computing 53 (2016) 32–57.doi:10.1016/ j.parco.2016.02.004
2016
-
[80]
C. E. Oancea, L. Rauchwerger, A Hybrid Approach to Proving Mem- ory Reference Monotonicity, in: S. Rajopadhye, M. Mills Strout (Eds.), Languages and Compilers for Parallel Computing, Springer 71 Berlin Heidelberg, Berlin, Heidelberg, 2013, pp. 61–75. doi:10.1007/ 978-3-642-36036-7_5
2013
-
[81]
S. Moon, M. W. Hall, Evaluation of Predicated Array Data-Flow Analysis for Automatic Parallelization, in: Proceedings of the Seventh ACM SIG- PLAN Symposium on Principles and Practice of Parallel Programming, PPoPP ’99, Association for Computing Machinery, New York, NY, USA, 1...
1999
-
[82]
C. E. Oancea, A. Mycroft, Set-congruence dynamic analysis for thread- level speculation (tls), in: J. N. Amaral (Ed.), Languages and Compilers for Parallel Computing, Springer Berlin Heidelberg, Berlin, Heidelberg, 2008, pp. 156–171
2008
-
[83]
F. Dang, H. Yu, L. Rauchwerger, The R-LRPD Test: Speculative Par- allelization of Partially Parallel Loops, in: Proceedings 16th Interna- tional Parallel and Distributed Processing Symposium, 2002, pp. 10 pp–. doi:10.1109/IPDPS.2002.1015493
2002 arXiv
-
[84]
M. Hall, C. Oancea, A. C. Elster, A. Rasch, S. Joshi, A. M. Tavakkoli, R. Schulze, Scheduling languages: A past, present, and future taxonomy (2024). arXiv:2410.19927. URL https://arxiv.org/abs/2410.19927
2024 arXiv
-
[85]
Donadio, J
S. Donadio, J. Brodman, T. Roeder, K. Yotov, D. Barthou, A. Cohen, M. J. Garzarán, D. Padua, K. Pingali, A language for the compact repre- sentation of multiple program versions, in: E. Ayguadé, G. Baumgartner, J. Ramanujam, P. Sadayappan (Eds.), Languages and Compilers for Pa...
2006 doi
-
[86]
C. Chen, J. Chame, M. W. Hall, Chill : A framework for composing high-level loop transformations, Tech. rep., Technical Report 08-897, U. of Southern California (2008)
2008
-
[87]
Girbal, N
S. Girbal, N. Vasilache, C. Bastoul, A. Cohen, D. Parello, M. Sigler, O. Temam, Semi-automatic composition of loop transformations for deep parallelism and memory hierarchies, International Journal of Parallel Pro- gramming 34 (Jun. 2006).doi:10.1007/s10766-006-0012-3
2006 doi
-
[88]
Ragan-Kelley, C
J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, S. Amaras- inghe, Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines, SIGPLAN Not. 48 (6) (2013) 519–530. doi:10.1145/2499370.2462176
2013
-
[89]
R. T. Mullapudi, V. Vasista, U. Bondhugula, Polymage: Automatic op- timization for image processing pipelines, in: Proceedings of the Twen- tieth International Conference on Architectural Support for Program- ming Languages and Operating Systems, ASPLOS ’15, Association for 72...
2015
-
[90]
Hegarty, J
J. Hegarty, J. Brunhaver, Z. DeVito, J. Ragan-Kelley, N. Cohen, S. Bell, A. Vasilyev, M. Horowitz, P. Hanrahan, Darkroom: Compiling high-level image processing code into hardware pipelines, ACM Trans. Graph. 33 (4) (2014) 144:1–144:11. doi:10.1145/2601097.2601174
2014
-
[91]
Nelson, A
T. Nelson, A. Rivera, P. Balaprakash, M. Hall, P. D. Hovland, E. Jessup, B. Norris, Generating efficient tensor contractions for gpus, in: 2015 44th International Conference on Parallel Processing, 2015, pp. 969–978.doi: 10.1109/ICPP.2015.106
2015 doi
-
[92]
Yadav, A
R. Yadav, A. Aiken, F. Kjolstad, DISTAL: the distributed tensor alge- bra compiler, in: Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation, PLDI 2022, Association for Computing Machinery, New York, NY, USA, 2022, p. 2...
2022
-
[93]
Venkat, M
A. Venkat, M. Hall, M. Strout, Loop and data transformations for sparse matrix code, in: Proceedings of the 36th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI ’15, Associa- tion for Computing Machinery, New York, NY, USA, 2015, p. 521–532
2015
-
[94]
Venkat, M
A. Venkat, M. S. Mohammadi, J. Park, H. Rong, R. Barik, M. M. Strout, M. Hall, Automating wavefront parallelization for sparse matrix compu- tations, in: Proceedings of the International Conference for High Per- formance Computing, Networking, Storage and Analysis, SC ’16, IEE...
2016
-
[95]
Senanayake, C
R. Senanayake, C. Hong, Z. Wang, A. Wilson, S. Chou, S. Kamil, S. Ama- rasinghe, F. Kjolstad, A sparse iteration space transformation framework for sparse tensor algebra, Proc. ACM Program. Lang. 4 (OOPSLA) (Nov. 2020). doi:10.1145/3428226
2020 doi
-
[96]
Bansal, O
M. Bansal, O. Hsu, K. Olukotun, F. Kjolstad, Mosaic: An interoperable compiler for tensor algebra, Proc. ACM Program. Lang. 7 (PLDI) (Jun. 2023). doi:10.1145/3591236
2023 doi
-
[97]
Tillet, H
P. Tillet, H. T. Kung, D. Cox, Triton: an intermediate language and compiler for tiled neural network computations, in: Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL 2019, Association for Computing Ma- chinery, Ne...
2019 doi
-
[98]
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L.Wang, Y.Hu, L.Ceze, etal.,{TVM}: Anautomated{End-to-End}op- timizing compiler for deep learning, in: Proceedings of the 13th USENIX 73 Conference on Operating Systems Design and Implementation, OSDI’18, USENI...
2018
-
[99]
Venkat, T
A. Venkat, T. Rusira, R. Barik, M. Hall, L. Truong, Swirl: High- performance many-core cpu code generation for deep neural networks, The International Journal of High Performance Computing Applications 33 (6) (2019) 1275–1289.doi:10.1177/1094342019866247
2019 doi
-
[100]
Paszke, D
A. Paszke, D. D. Johnson, D. Duvenaud, D. Vytiniotis, A. Radul, M. J. Johnson, J. Ragan-Kelley, D. Maclaurin, Getting to the point: index sets and parallelism-preserving autodiff for pointful array programming, Proc. ACM Program. Lang. 5 (ICFP) (Aug. 2021).doi:10.1145/3473593
2021 doi
-
[101]
Steuwer, T
M. Steuwer, T. Remmelg, C. Dubach, Lift: a functional data-parallel IR for high-performance GPU code generation, in: Proceedings of the 2017 International Symposium on Code Generation and Optimization, CGO ’17, IEEE Press, 2017, p. 74–85
2017
-
[102]
Steuwer, T
M. Steuwer, T. Koehler, B. Köpcke, F. Pizzuti, RISE & shine: Language- oriented compiler design, CoRR abs/2201.03611 (2022). arXiv:2201. 03611. URL https://arxiv.org/abs/2201.03611
2022 arXiv
-
[103]
61–72.doi:10.1145/3578360.3580269
A.Rasch, R.Schulze, D.Shabalin, A.Elster, S.Gorlatch, M.Hall, (de/re)- compositions expressed systematically via mdh-based schedules, in: Pro- ceedings of the 32nd ACM SIGPLAN International Conference on Com- piler Construction, CC 2023, Association for Computing Machinery, Ne...
2023
-
[104]
Kundefinedhler, A
T. Kundefinedhler, A. Goens, S. Bhat, T. Grosser, P. Trinder, M. Steuwer, Guided Equality Saturation, Proc. ACM Program. Lang. 8 (POPL) (Jan. 2024). doi:10.1145/3632900
2024 doi
-
[105]
Holk, Region-based memory management for expressive gpu program- ming, Ph.D
E. Holk, Region-based memory management for expressive gpu program- ming, Ph.D. thesis, Indiana University (2016). 74
2016
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.