Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Comparing Parallel Functional Array Languages: Programming and Performance

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Five functional array languages, compiled from a single high-level source to both multicore and GPU, reach 80% or more of hand-optimized OpenMP, Fortran, and CUDA performance on 25 of 36 baseline instances, with codebases at least 2x…

desk verdict A serious, useful five-language benchmark study whose headline performance claims are inflated by counting single-core baseline wins and by a self-written FlashAttention baseline; the underlying Futhark/DaCe results are credible and worth publishing after the claims are softened. read the letter →

arxiv 2505.08906 v1 pith:ARWX4LZ7 submitted 2025-05-13 cs.PL cs.DCcs.PF

classification cs.PLcs.DCcs.PF
keywords parallelfunctionalarraylanguagesdata-parallelprogrammingperformanceportabilityGPUcompilationnestedparallelismbenchmarkingsourcelinesofcode
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that functional array languages are not just expressive but can be fast enough to compete with hand-optimized conventional HPC code. On four benchmarks spanning regular and irregular nested parallelism, five languages — Accelerate, APL, DaCe, Futhark, and SaC — each compiled a single high-level source to both a 32-core multicore system and an NVIDIA A30 GPU. Against OpenMP, Fortran, and CUDA baselines across 36 baseline instances, at least one functional language matched or beat baseline performance in 30% of instances and reached 80% or more of baseline performance in 70% of instances. The code is also far shorter: the functional codebases are at least 2 times smaller than the CPU baseline and at least 8 times smaller than the GPU baseline, while the combined baseline codebase is at least 10 times larger than any functional one. The paper concludes that mature functional array languages have the potential to deliver performance competitive with the best available conventional techniques, with remaining gaps largely attributed to backend engineering maturity rather than a fundamental limitation of the approach.

What carries the argument

The load-bearing object is the benchmark comparison protocol: each of the four problems is implemented once per language from a correct starting point, then repeatedly refactored for performance, while the same source is compiled to both multicore and GPU. The underlying mechanism is the data-parallel compilation pipeline, in which second-order array combinators such as map, reduce, scan, and scatter are fused, flattened, and tiled by the compiler; Futhark's incremental flattening, a multi-version compilation technique that maps nested application parallelism onto GPU grid and block levels, and DaCe's graph-based SDFG transformations are the most developed examples. The performance ratios against hand-optimized baselines, together with source-line-of-code counts, are the quantities that carry the argument.

What would settle it

Run the same four benchmarks with baselines that are independently maintained and unmodified public releases, and have the functional implementations written by programmers who did not design the languages under a fixed, realistic time budget; if the fraction of baseline instances where at least one functional language reaches 80% of baseline performance falls below a majority, the central claim would be undermined.

Watch

Extended reading notes

Core claim

The central claim is that mature functional array languages can deliver performance competitive with the best available conventional techniques. The evidence is a systematic comparison on 39 instances of four benchmarks: N-body simulation, MultiGrid, Quickhull, and Flash Attention, run on a 32-core AMD EPYC 7313 multicore system and an NVIDIA A30 GPU. Across the 36 instances with hand-optimized baselines, at least one functional language matched or outperformed the baseline in 30% of cases, exceeded 80% of baseline performance in 70% of cases, and only in 2 instances — Quickhull on the multicore CPU — did no language exceed 50% of baseline performance. This is achieved from a single high-level source per language, compiled to both CPU and GPU, with functional codebases at least 10 times smaller than the combined CPU-plus-GPU baseline codebase. The authors argue that the primary reason some languages underperform is incomplete engineering of their backends, not a fundamental limitation of the functional array approach, and they note that only DaCe and Futhark consistently achieve good performance on both architectures.

Load-bearing premise

The comparison assumes the functional implementations and the hand-optimized baselines received comparable and equally expert tuning effort, but the benchmark teams largely wrote their own languages and some baselines were produced by the same group, so unequal effort could explain the performance results.

Editorial extensions

If this is right

  • A single functional array source can replace separate CPU and GPU implementations, cutting code by at least 2x versus the CPU baseline and at least 8x versus the GPU baseline.
  • Porting a functional array program to new hardware becomes the task of writing a good compiler backend, not rewriting the application.
  • The two languages with the most developed backends on both targets, DaCe and Futhark, are the ones that deliver portable performance, suggesting backend quality is the current bottleneck.
  • On Quickhull, Futhark's GPU version ran the 100-million-point datasets 2.3x to 4.3x faster than the multicore baseline, showing that at least some irregular divide-and-conquer workloads can be won by flat data-parallel code.
  • Because every functional language implemented the simpler Custom Attention rather than the full Flash Attention algorithm, the Flash Attention shortfall is partly an algorithmic choice and adopting the tiled algorithm should close much of the gap.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if backend engineering is the main constraint, a follow-up study of the same languages five years later should show CPU-side gaps closing without source changes, which would confirm the portability claim.
  • Beyond the paper: the benchmark set implies a minimal primitive checklist for a performant array language — scan or prefix-sum, reduce-by-index, recursion or its unfolding, and tiled matrix multiplication — since languages missing any of these were unable to compete on the corresponding benchmark.
  • Beyond the paper: the source-line-of-code evidence suggests that even at half the baseline speed, functional array code may win on total development cost; a direct test would measure end-to-end programmer time, including debugging and autotuning, to first correct and then fast versions.
  • Beyond the paper: both test architectures share the same programming model family, so a stronger portability test would compile the same sources to a non-NVIDIA GPU or an FPGA, where the paper's backend-engineering explanation predicts DaCe and Futhark would still lead.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents a systematic comparison of five parallel functional array languages (Accelerate, APL, DaCe, Futhark, and SaC) on four benchmarks (N-body, MultiGrid, Quickhull, and Flash Attention), targeting a 32-core AMD EPYC 7313 and an NVIDIA A30 GPU. For each language it describes design and implementation choices, reports SLOC as a proxy for programming effort, and gives performance tables for multicore and GPU execution. The central claim is that mature functional array languages have the potential to deliver performance competitive with the best available conventional techniques, supported mainly by aggregate statistics over 36 benchmark instances and by the observation that a single high-level source replaces separate CPU and GPU baselines.

Significance. If the performance and code-size claims survive scrutiny, this would be a valuable reference for the functional array programming community: it covers five mature implementations, uses two external baseline suites (NAS MG and PBBS Quickhull), publishes an open-source benchmark repository, and honestly reports many negative results (e.g., SaC's missing GPU results for MG and FlashAttention, APL's generally low performance, and the universal FlashAttention shortfall). The design and implementation comparison in Sections 3 and 4 is informative and largely independent of the performance claims. However, the headline statistics in Section 10 are not reproducible from the tables, and the FlashAttention baseline is not the best available conventional technique, so the central conclusion is currently overstated and needs correction before the paper can be accepted.

major comments (3)
  1. [Section 10, Tables 5-8] The aggregate statistics in Section 10 do not match the data in Tables 5-8. An independent recount of all 36 baseline instances gives 24/36 (67%) with at least one language at or above 80% of baseline, 12/36 (33%) at or above 100%, 10/36 (28%) strictly between 50% and 80%, and 2/36 (6%) at or below 50%. The paper reports 25 (70%), 11 (30%), 9 (25%), and 2 (6%). One concrete discrepancy is the FlashAttention dataset (d=64, N=32768) on 1 CPU core, where Accelerate reaches 150 Gflops against a baseline of 77 Gflops (ratio 1.95): this instance should be counted as matching/outperforming the baseline, but the paper's total of 11 omits it. Please correct the counts and clarify that the intended categories are disjoint.
  2. [Section 10, Tables 5-8] The headline statistics include the single-core '1C' columns as benchmark instances, but the central claim is about 'the best available conventional techniques' on a machine that provides a 32-core CPU and a GPU. If the 1C instances are excluded, the picture changes materially: only 14 of the 24 remaining parallel instances (58%) have at least one language at or above 80% of baseline, and only 3 (12.5%) match or exceed the baseline. The paper should either justify why single-core runs support a parallel-competitiveness claim or restrict the aggregate statistics to the 32C and GPU columns.
  3. [Sections 9.2, 9.8, 10] The FlashAttention baseline is not the best available conventional technique. Section 9.2 states that the paper writes its own GPU baseline in single precision using regular FPUs because the official FlashAttention implementation uses half-precision tensor cores, and Sections 9.3-9.7 show that all five functional languages implement the less efficient Custom Attention (Algorithm 6), not FlashAttention. Therefore the 12 FlashAttention instances cannot support the conclusion that mature functional array languages are competitive with the best available conventional techniques; they only support a claim of competitiveness with a self-defined single-precision baseline. The conclusion should be softened accordingly, or the baselines should be replaced by the official implementation for those languages that can target tensor cores. The same caveat applies to the self-written small-n N-body GPU baseline described in Section 6.2.
minor comments (4)
  1. [Table 9 (referenced as Table 4 in Section 10)] The SLOC totals do not match the row sums: Futhark's total is listed as 448 but its row entries sum to 46+136+161+90=433, and SaC's total of 479 is not reconcilable with the entries shown (61, —, 203, 79). Please audit the table.
  2. [Section 5.4] The measurement methodology reports averages but the tables do not show variance or the number of runs per cell; given that some results are close to the 80% threshold (for example, Futhark's 547 Gflops versus the N-body n=10^3 GPU baseline of 560 Gflops in Table 5), reporting standard deviations or per-cell run counts would substantially strengthen the comparison.
  3. [Abstract and Section 10] The abstract and Introduction say the evaluation covers 39 instances, while Section 10 says there are 36 benchmark instances with a baseline; please clarify that the difference is the three Quickhull GPU instances for which no baseline exists.
  4. [Figure 9] Figure 9 appears four times with subfigures (a)-(d) but has no caption; please add a caption that describes the DaCe optimization workflow illustrated in the four parts.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline performance claim is supported by direct measurements against external or published baselines, not by a self-referential derivation.

full rationale

This paper is an empirical benchmark and language-design comparison rather than a derivation of performance from language semantics. The central quantitative claim in Section 10 is supported by measured runtimes in Tables 5-8 against external or published baselines: NAS MG uses the Fortran+OpenMP and NPB-GPU implementations, Quickhull uses PBBS, and N-body uses a CUDA sample plus a self-written OpenMP baseline. No equation in the paper defines a predicted quantity in terms of a fitted parameter, and no result is asserted solely because a same-author citation says so. The language self-citations (Futhark, SaC, Accelerate, DaCe, Co-dfns) provide background descriptions of compilers and optimization techniques, but they are not the evidence for the headline performance claim. The main caveat is methodological rather than circular: Section 9.2 states that the FlashAttention GPU baseline was written by the authors using single-precision arithmetic and regular FPUs because the official implementation uses half-precision tensor cores, and Section 9.8 concedes that 'we cannot meaningfully compare most languages with the baseline because it implements the more efficient Flash Attention.' This weakens the phrase 'best available conventional techniques' for those instances, but it is a baseline-suitability and overclaim concern, not a reduction of the conclusion to its inputs. Similarly, authorship overlap between benchmark writers and language teams is a potential bias, but it does not make the measured performance numbers logically dependent on the conclusion. Accordingly, the circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No mathematical derivation is present, so the ledger records the empirical premises on which the comparative claim rests. The central claim does not depend on fitted constants or new postulated entities; it depends on the representativeness of the baselines, the fairness of implementation effort, the validity of SLOC as a proxy, and the choice of benchmark workloads.

free parameters (3)
  • N-body dataset pairs (n, t) = (1e3, 1e5), (1e4, 1e3), (1e5, 1e1)
    Chosen by hand so each dataset has the same nominal work while varying parallelism; Section 6.3 shows that the largest dataset benefits from sequentializing the inner loop, so this choice affects which optimization strategy wins.
  • FlashAttention dataset pairs (d, N) = (64, 16384), (64, 32768), (128, 8192), (128, 16384)
    Chosen to cover typical attention shapes; the performance conclusions for Flash Attention depend on these sizes.
  • Aggregation threshold 'at least one language' and the 80%/50% levels = N/A
    The headline percentages in Section 10 summarize the best language in each instance rather than a fixed language; this hand-chosen decision shapes how strong the class-level claim appears.
assumptions (4)
  • domain assumption The selected baselines represent the best available conventional techniques.
    The conclusion depends on NAS, PBBS, CUDA samples, and the self-written N-body and Flash Attention baselines being strong representatives of conventional practice; Sections 6.2 and 9.2 show some baselines were written by the authors.
  • domain assumption Implementation effort and expertise are comparable across the five languages.
    Section 5.3 says each benchmark was iteratively refactored by the language teams, but there is no independent control for effort; the paper itself notes that Accelerate code was not tuned and SaC's GPU backend is immature.
  • domain assumption SLOC is a valid proxy for programming effort and expressiveness.
    Section 10 uses Source Lines of Code as a 'crude but widely accepted measure' while acknowledging differences in style and tooling; comprehensibility claims are explicitly subjective.
  • domain assumption The four chosen benchmarks are representative of the parallel computational models relevant to functional array languages.
    The paper selects N-body, Multigrid, Quickhull, and Flash Attention to represent a range of application domains and parallel models (Section 5.1), and external validity depends on that selection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Comparing Parallel Functional Array Languages: Programming and Performance." pith.science (2026). https://pith.science/paper/ARWX4LZ7

@misc{pith2026250508906,
  author       = {Pith},
  title        = {Pith review of: Comparing Parallel Functional Array Languages: Programming and Performance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ARWX4LZ7}},
  note         = {Machine review of arXiv:2505.08906}
}
read the original abstract

Parallel functional array languages are an emerging class of programming languages that promise to combine low-effort parallel programming with good performance and performance portability. We systematically compare the designs and implementations of five different functional array languages: Accelerate, APL, DaCe, Futhark, and SaC. We demonstrate the expressiveness of functional array programming by means of four challenging benchmarks, namely N-body simulation, MultiGrid, Quickhull, and Flash Attention. These benchmarks represent a range of application domains and parallel computational models. We argue that the functional array code is much shorter and more comprehensible than the hand-optimized baseline implementations because it omits architecture-specific aspects. Instead, the language implementations generate both multicore and GPU executables from a single source code base. Hence, we further argue that functional array code could more easily be ported to, and optimized for, new parallel architectures than conventional implementations of numerical kernels. We demonstrate this potential by reporting the performance of the five parallel functional array languages on a total of 39 instances of the four benchmarks on both a 32-core AMD EPYC 7313 multicore system and on an NVIDIA A30 GPU. We explore in-depth why each language performs well or not so well on each benchmark and architecture. We argue that the results demonstrate that mature functional array languages have the potential to deliver performance competitive with the best available conventional techniques.

Figures

Figures reproduced from arXiv: 2505.08906 by the authors.

Figure 1
Figure 1. Some Futhark syntax examples. not support general recursion, but merely provides tail recursion through a special syntactic construct, named loop. The loop construct binds the loop parameter (induction variable) x to 0, then evaluates the body n times; binding the result of each evaluation to x, and returning the final value of x, as illustrated in [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Futhark N-body simulation, representing vectors as records. In a larger application, we would likely define a reusable module for vectors, and perhaps parameterize over the number type. Only the functions calc_accels, step, and nbody involve parallelism, and the latter only by virtue of invoking step. 7 [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Accelerate N-body simulation: data type definitions and auxiliary functions. array arguments to zipWith have the same rank, but they can be of different sizes. The result array also has the same rank, and the size is the minimum of the two arrays in each dimension. Values of Shape type can be used to specify the size of an array, or to index into an array. For example, the function generate :: (Shape sh, Elt a) => E… view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: Accelerate N-body simulation: main functions. accel function is then applied on each point-wise pair, and the result is folded into a vector again. Generating the intermediate arrays may seem inefficient, but the implementation eliminates, or fuses, the arrays, and the…
Figure 5
Figure 5. Figure 5: APL N-body simulation. See text for description of each element. high-level description of the application’s models, unencumbered by hardware￾specific code and performance optimizations in general. The performance en￾gineer does not improve the high-level description d…
Figure 6
Figure 6. Figure 6: Elements of the SDFG IR, for details see [16]. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: DaCe N-body simulation in data-centric Python and its SDFG IR representation. The double-headed arrows match selected computations’ Python code with their correspond￾ing SDFG representation. Cloud-like shapes abstract away several SDFG subgraphs containing Map scopes f…
Figure 8
Figure 8. Figure 8: Structural organization of the SaC compiler sac2c [PITH_FULL_IMAGE:figures/full_fig_p028_8.png]
Figure 9
Figure 9. Figure 9: Optimization workflow for the acceleration computation of the [PITH_FULL_IMAGE:figures/full_fig_p031_9.png]
Figure 10
Figure 10. Figure 10: Multigrid method for numerically solving [PITH_FULL_IMAGE:figures/full_fig_p040_10.png]
Figure 11
Figure 11. Figure 11: The left-hand side shows a simplified implementation of the optimized computa￾tional kernel of MG, named relaxNAS, which is made generic by parameterizing it over an index function (elmAt) rather than a 3D array. The right-hand side shows the instantiations required t…
Figure 12
Figure 12. Figure 12: A Rank-Polymorphic stencil in SaC. For suitably chosen weights [PITH_FULL_IMAGE:figures/full_fig_p042_12.png]
Figure 13
Figure 13. Figure 13: DaCe Python implementation of norm2u3 and its optimization. 7.8. Summary The NAS MG performance results for the five languages on both multicore CPU and GPU are summarized in [PITH_FULL_IMAGE:figures/full_fig_p044_13.png]
Figure 14
Figure 14. Figure 14: In the code, KT denotes the transpose of K and × denotes matrix multiplication. forall i ∈ 0 .. N − 1 by d denotes a parallel (map-like) computation in which i starts from 0 and advances with a step of d, i.e., i = 0, d, 2 · d, . . . . Qb iterates over the d × d slice…
Figure 15
Figure 15. Figure 15: Algorithm 8 uses the same input and output as Algorithm 7. [PITH_FULL_IMAGE:figures/full_fig_p051_15.png]
Figure 16
Figure 16. Figure 16: Custom Attention in Futhark. The implementation of onlineSoftmax (not shown) [PITH_FULL_IMAGE:figures/full_fig_p052_16.png]
Figure 17
Figure 17. Figure 17: DaCe data-centric Python implementation of Flash Attention. [PITH_FULL_IMAGE:figures/full_fig_p053_17.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. High-Level Big Integer Arithmetic in Futhark for GPUs

    cs.SC 2026-07 conditional novelty 6.0 of 10

    High-level Futhark code for GPU big-integer arithmetic can approach hand-written CUDA performance once arrays are automatically placed in registers.

Reference graph

Works this paper leans on

104 extracted references · 38 canonical work pages · cited by 1 Pith paper

  1. [1]

    R. K. W. Hui, M. J. Kromberg, APL since 1978, Proc. ACM Program. Lang. 4 (HOPL) (Jun. 2020).doi:10.1145/3386319

  2. [2]

    Backus, The history of Fortran I, II, and III, Association for Computing Machinery, New York, NY, USA, 1978, Ch

    J. Backus, The history of Fortran I, II, and III, Association for Computing Machinery, New York, NY, USA, 1978, Ch. 2, p. 25–74.doi:10.1145/ 800025.1198345

  3. [3]

    Ihaka, R

    R. Ihaka, R. Gentleman, R: A Language for Data Analysis and Graphics, Journal of Computational and Graphical Statistics 5 (3) (1996) 299–314. doi:10.1080/10618600.1996.10474713

  4. [4]

    T. L. McDonell, M. M. Chakravarty, G. Keller, B. Lippmeier, Optimis- ing purely functional GPU programs, in: Proceedings of the 18th ACM SIGPLAN International Conference on Functional Programming, ICFP ’13, Association for Computing Machinery, New York, NY, USA, 2013, p. 49–60. doi:10.1145/2500365.2500595

  5. [5]

    A. N. Ziogas, T. Schneider, T. Ben-Nun, A. Calotoiu, T. De Matteis, J. de Fine Licht, L. Lavarini, T. Hoefler, Productivity, portability, perfor- mance: data-centric Python, in: Proceedings of the International Confer- ence for High Performance Computing, Networking, Storage and Analysis, 13https://github.com/diku-dk/CF AL-bench/ 63 SC ’21, Association fo...

  6. [6]

    Henriksen, N

    T. Henriksen, N. G. W. Serup, M. Elsman, F. Henglein, C. E. Oancea, Futhark: Purely Functional GPU-programming with Nested Parallelism and In-place Array Updates, in: Proceedings of the 38th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2017, ACM, New York, NY, USA, 2017, pp. 556–571. doi:10.1145/ 3062341.3062354

  7. [7]

    Grelck, S

    C. Grelck, S. Scholz, SAC - A functional array language for efficient multi- threaded execution, Int. J. Parallel Program. 34 (4) (2006) 383–427.doi: 10.1007/S10766-006-0018-X

  8. [8]

    Rosenberg, Some misconceptions about lines of code, in: Proceed- ings fourth international software metrics symposium, IEEE, IEEE, Al- buquerque, NM, USA, 1997, pp

    J. Rosenberg, Some misconceptions about lines of code, in: Proceed- ings fourth international software metrics symposium, IEEE, IEEE, Al- buquerque, NM, USA, 1997, pp. 137–142.doi:10.1109/METRIC.1997. 637174

Show all 104 references
  1. [9]

    A. K. Hovgaard, T. Henriksen, M. Elsman, High-performance defunction- alisation in futhark, in: M. Pałka, M. Myreen (Eds.), Trends in Func- tional Programming, Springer International Publishing, Cham, 2019, pp. 136–156. doi:10.1007/978-3-030-18506-0_7

  2. [10]

    G. E. Blelloch, Programming parallel algorithms, Commun. ACM 39 (3) (1996) 85–97. doi:10.1145/227234.227246

  3. [11]

    T. L. McDonell, J. D. Meredith, G. Keller, Embedded pattern matching, in: Proceedings of the 15th ACM SIGPLAN International Haskell Sym- posium, Haskell 2022, Association for Computing Machinery, New York, NY, USA, 2022, p. 123–136.doi:10.1145/3546189.3549917

  4. [12]

    Grelck, S.-B

    C. Grelck, S.-B. Scholz, Classes and objects as basis for i/o in sac, in: T. Johnsson (Ed.), 7th International Workshop on Implementation of Functional Languages (IFL’95), Båstad, Sweden, Chalmers University of Technology, Gothenburg, Sweden, 1995, pp. 30–44. URL sac-classes-o...

  5. [13]

    Huijben, J

    R. Huijben, J. Aaldering, P. Achten, S.-B. Scholz, Flattening combina- tions of arrays and records, in: J. Hemann, S. Chang (Eds.), Trends in Functional Programming, Springer Nature Switzerland, Cham, 2025, pp. 220–240. doi:10.1007/978-3-031-74558-4_10

  6. [14]

    Calotoiu, T

    A. Calotoiu, T. Ben-Nun, G. Kwasniewski, J. de Fine Licht, T. Schneider, P. Schaad, T. Hoefler, Lifting C semantics for dataflow optimization, in: Proceedings of the 36th ACM International Conference on Supercomput- ing, ICS ’22, Association for Computing Machinery, New York, ...

  7. [16]

    URL https://spcldace.readthedocs.io

    Scalable Parallel Computing Laboratory, ETH Zurich, DaCe: Data- Centric Parallel Programming (2024). URL https://spcldace.readthedocs.io

  8. [17]

    Gonnord, L

    L. Gonnord, L. Henrio, L. Morel, G. Radanne, A Survey on Parallelism and Determinism, ACM Comput. Surv. 55 (10) (Feb. 2023). doi:10. 1145/3564529

  9. [18]

    1–14.doi:10.1109/SC41405.2020.00101

    T.Henriksen, S.Hellfritzsch, P.Sadayappan, C.Oancea, Compiling gener- alized histograms for gpu, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’20, IEEE Press, 2020, pp. 1–14.doi:10.1109/SC41405.2020.00101

  10. [19]

    Wadler, Deforestation: transforming programs to eliminate trees, Theoretical Computer Science 73 (2) (1990) 231–248

    P. Wadler, Deforestation: transforming programs to eliminate trees, Theoretical Computer Science 73 (2) (1990) 231–248. doi:10.1016/ 0304-3975(90)90147-A

  11. [20]

    G. E. Blelloch, Vector models for data-parallel computing, Vol. 75, MIT press Cambridge, 1990

  12. [21]

    Henriksen, F

    T. Henriksen, F. Thorøe, M. Elsman, C. Oancea, Incremental Flattening for Nested Data Parallelism, in: Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming, PPoPP ’19, ACM, New York, NY, USA, 2019, pp. 53–67.doi:10.1145/3293883.3295707

  13. [22]

    Munksgaard, S

    P. Munksgaard, S. L. Breddam, T. Henriksen, F. C. Gieseke, C. Oancea, Dataset sensitive autotuning of multi-versioned code based on mono- tonic properties, in: V. Zsók, J. Hughes (Eds.), Trends in Functional Programming, Springer International Publishing, Cham, 2021, pp. 3–23....

  14. [23]

    Schenck, O

    R. Schenck, O. Rønning, T. Henriksen, C. E. Oancea, Ad for an array language with nested parallelism, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’22, IEEE Press, 2022, pp. 1–15.doi:10.1109/SC41404. 2022.00063

  15. [24]

    Munksgaard, T

    P. Munksgaard, T. Henriksen, P. Sadayappan, C. Oancea, Memory op- timizations in an array language, in: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, SC ’22, IEEE Press, 2022, pp. 1–15.doi:10.1109/SC41404. 2022.00036. 65

  16. [25]

    A.Nicolaisen, M.A. Persson, Implementingsingle-pass scaninthe futhark compiler, Master’s thesis, Department of Computer Science, Faculty of Science, University of Copenhagen, https://futhark-lang.org/student- projects/marco-andreas-scan.pdf (2020). URL https://futhark-lang.org...

  17. [26]

    M. T. Clausen, Regular segmented single-pass scan in futhark, Master’s thesis, Department of Computer Science, Faculty of Science, University of Copenhagen, https://futhark-lang.org/student-projects/morten-msc- thesis.pdf (2021). URL https://futhark-lang.org/student-projects/ ...

  18. [27]

    Merrill, M

    D. Merrill, M. Garland, Single-pass parallel prefix scan with decoupled lookback, Nvidia technical report nvr-2016-002, march 2016, NVIDIA (2016). URL https://research.nvidia.com/sites/default/files/pubs/ 2016-03_Single-pass-Parallel-Prefix/nvr-2016-002.pdf

  19. [28]

    T.Schrijvers, S.PeytonJones, M.Chakravarty, M.Sulzmann, Typecheck- ing with open type functions, in: Proceedings of the 13th ACM SIGPLAN International Conference on Functional Programming, ICFP ’08, Associ- ation for Computing Machinery, New York, NY, USA, 2008, p. 51–62. doi:...

  20. [29]

    M. M. T. Chakravarty, G. Keller, S. P. Jones, Associated type synonyms, in: Proceedings of the Tenth ACM SIGPLAN International Conference on Functional Programming, ICFP ’05, Association for Computing Ma- chinery, New York, NY, USA, 2005, p. 241–253.doi:10.1145/1086365. 1086397

  21. [30]

    Grelck, S.-B

    C. Grelck, S.-B. Scholz, Merging compositions of array skeletons in SAC, Journal of Parallel Computing 32 (7+8) (2006) 507–522.doi:10.1016/ j.parco.2006.08.003

  22. [31]

    Grelck, K

    C. Grelck, K. Trojahner, Implicit Memory Management for SaC, in: C. Grelck, F. Huch (Eds.), Implementation and Application of Functional Languages, 16th International Workshop, IFL’04, University of Kiel, In- stitute of Computer Science and Applied Mathematics, 2004, pp. 335–3...

  23. [32]

    Grelck, Single Assignment C (SAC): the compilation technology per- spective, in: V

    C. Grelck, Single Assignment C (SAC): the compilation technology per- spective, in: V. Zsók, Z. Porkoláb, Z. Horváth (Eds.), 6th Central Eu- ropean Functional Programming Summer School (CEFP’15), Budapest, Hungary, Vol. 10094 of Lecture Notes in Computer Science, Springer, 201...

  24. [33]

    A. W.-y. Hsu, A data parallel compiler hosted on the GPU, Ph.D. thesis, Indiana University (2019). URL https://hdl.handle.net/2022/24749

  25. [34]

    A. W. Keep, A nanopass framework for commercial compiler development, Ph.D. thesis, Indiana University (2012). URL https://andykeep.com/pubs/dissertation.pdf

  26. [35]

    Malcolm, P

    J. Malcolm, P. Yalamanchili, C. McClanahan, V. Venugopalakrishnan, K. Patel, J. Melonakos, ArrayFire: a GPU acceleration platform, in: E. J. Kelmelis (Ed.), Modeling and Simulation for Defense Systems and Appli- cations VII, Vol. 8403, International Society for Optics and Phot...

  27. [36]

    A. N. Ziogas, T. Ben-Nun, G. I. Fernández, T. Schneider, M. Luisier, T. Hoefler, A Data-Centric Approach to Extreme-Scale Ab initio Dissipa- tive Quantum Transport Simulations, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage ...

  28. [37]

    Ben-Nun, L

    T. Ben-Nun, L. Groner, F. Deconinck, T. Wicky, E. Davis, J. Dahm, O. Elbert, R. George, J. McGibbon, L. Trümper, E. Wu, O. Fuhrer, T. Schulthess, T. Hoefler, Productive Performance Engineering for Weather and Climate Modeling with Python, in: Proceedings of the Inter- national...

  29. [38]

    Schaad, T

    P. Schaad, T. Ben-Nun, T. Hoefler, Boosting Performance Optimization with Interactive Data Movement Visualization, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC’22), 2022, pp. 1–16.doi:10.1109/SC41404. 2022.00069

  30. [39]

    A. N. Ziogas, G. Kwasniewski, T. Ben-Nun, T. Schneider, T. Hoe- fler, Deinsum: Practically I/O Optimal Multilinear Algebra, in: Pro- ceedings of the International Conference for High Performance Comput- ing, Networking, Storage and Analysis (SC’22), 2022, pp. 1–15. doi: 10.555...

  31. [40]

    Trümper, T

    L. Trümper, T. Ben-Nun, P. Schaad, A. Calotoiu, T. Hoefler, Performance Embeddings: A Similarity-Based Transfer Tuning Approach to Perfor- mance Optimization, in: Proceedings of the 37th International Conference on Supercomputing, ICS ’23, Association for Computing Machinery, ...

  32. [41]

    D. H. Bailey, E. Barszcz, J. T. Barton, D. S. Browning, R. L. Carter, L. Dagum, R. A. Fatoohi, P. O. Frederickson, T. A. Lasinski, R. S. Schreiber, H. D. Simon, V. Venkatakrishnan, S. K. Weeratunga, The NAS 67 parallel benchmarks—summary and preliminary results, in: Proceeding...

  33. [42]

    G.Araujo, D.Griebler, D.A.Rockenbach, M.Danelutto, L.G.Fernandes, NAS Parallel Benchmarks with CUDA and beyond, Software: Practice and Experience 53 (1) (2023) 53–80.doi:10.1002/spe.3056

  34. [43]

    Anderson, G

    D. Anderson, G. E. Blelloch, L. Dhulipala, M. Dobson, Y. Sun, The problem-based benchmark suite (pbbs), v2, in: Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Pro- gramming, PPoPP ’22, Association for Computing Machinery, New York, NY, USA...

  35. [44]

    T. Dao, D. Fu, S. Ermon, A. Rudra, C. Ré, FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, in: S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, A. Oh (Eds.), Advances in Neural Information Processing Systems, Vol. 35, Curran Associates, Inc.,...

  36. [45]

    Bailey, T

    D. Bailey, T. Harris, W. Saphir, R. Van Der Wijngaart, A. Woo, M. Yarrow, The nas parallel benchmarks 2.0, Tech. rep., Technical Re- port NAS-95-020, NASA Ames Research Center (1995)

  37. [46]

    Di Domenico, G

    D. Di Domenico, G. G. H. Cavalheiro, J. V. F. Lima, NAS Parallel Bench- mark Kernels with Python: A performance and programming effort anal- ysis focusing on GPUs, in: 2022 30th Euromicro International Conference on Parallel, Distributed and Network-based Processing (PDP), 202...

  38. [47]

    M. M. T. Chakravarty, R. Leshchinskiy, S. Peyton Jones, G. Keller, S. Marlow, Data parallel Haskell: a status report, in: Proceedings of the 2007 Workshop on Declarative Aspects of Multicore Programming, DAMP ’07, Association for Computing Machinery, New York, NY, USA, 2007, p...

  39. [48]

    Zhang, F

    Y. Zhang, F. Mueller, CuNesl: Compiling nested data-parallel languages for SIMT architectures, in: Proceedings of the 2012 41st International Conference on Parallel Processing, ICPP’12, IEEE Computer Society, Washington, DC, USA, 2012, pp. 340–349.doi:10.1109/ICPP.2012.21

  40. [49]

    Bergstrom, J

    L. Bergstrom, J. Reppy, Nested data-parallelism on the gpu, in: Proceed- ings of the 17th ACM SIGPLAN International Conference on Functional Programming, ICFP ’12, ACM, New York, NY, USA, 2012, pp. 247–258. doi:10.1145/2364527.2364563. 68

  41. [50]

    Edelkamp, A

    S. Edelkamp, A. Weiß, Blockquicksort: Avoiding branch mispredictions in quicksort, ACM J. Exp. Algorithmics 24 (Jan. 2019).doi:10.1145/ 3274660

  42. [51]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin, Attention is all you need, in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems,...

  43. [52]

    J. S. Bridle, Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition, in: F. F. Soulié, J. Hérault (Eds.), Neurocomputing, Springer Berlin Heidelberg, Berlin, Heidelberg, 1990, pp. 227–236. doi:10.1007/...

  44. [53]

    Milakov, N

    M. Milakov, N. Gimelshein, Online normalizer calculation for softmax (2018). arXiv:1805.02867. URL https://arxiv.org/abs/1805.02867

  45. [54]

    Dao AI Lab, FlashAttention, https://github.com/Dao-AILab/ flash-attention (2024)

  46. [55]

    Henriksen, C

    T. Henriksen, C. E. Oancea, A T2 graph-reduction approach to fusion, in: Proceedings of the 2Nd ACM SIGPLAN Workshop on Functional High- performance Computing, FHPC ’13, ACM, New York, NY, USA, 2013, pp. 47–58. doi:10.1145/2502323.2502328

  47. [56]

    Henriksen, K

    T. Henriksen, K. F. Larsen, C. E. Oancea, Design and GPGPU perfor- mance of futhark’s redomap construct, in: Proceedings of the 3rd ACM SIGPLAN International Workshop on Libraries, Languages, and Compil- ers for Array Programming, ARRAY 2016, ACM, New York, NY, USA, 2016, pp. ...

  48. [57]

    D. W. Walker, J. J. Dongarra, MPI: a standard message passing interface, Supercomputer 12 (1996) 56–68.doi:10.1145/169627.169855

  49. [58]

    P. W. Trinder, K. Hammond, H.-W. Loidl, S. P. Jones, Algorithm+ strat- egy= parallelism, Journal of functional programming 8 (1) (1998) 23–60. doi:10.1017/S0956796897002967

  50. [59]

    C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Vir- tanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Hal- dane, J. F. del Río, M. Wiebe, P. Peterson, P. Gérard-Marchant, K. Shep- pa...

  51. [60]

    Sidelnik, S

    A. Sidelnik, S. Maleki, B. L. Chamberlain, M. J. Garzar’n, D. Padua, Performance portability with the chapel language, in: 2012 IEEE 26th international parallel and distributed processing symposium, IEEE, 2012, pp. 582–594. doi:10.1109/IPDPS.2012.60

  52. [61]

    L. V. Kale, S. Krishnan, CHARM++: a portable concurrent object ori- ented system based on C++, in: Proceedings of the Eighth Annual Con- ference on Object-Oriented Programming Systems, Languages, and Appli- cations, OOPSLA ’93, Association for Computing Machinery, New York, NY...

  53. [62]

    Rice University, High performance fortran language specification, SIG- PLAN Fortran Forum 12 (4) (1993) 1–86.doi:10.1145/174223.158909

    C. Rice University, High performance fortran language specification, SIG- PLAN Fortran Forum 12 (4) (1993) 1–86.doi:10.1145/174223.158909

  54. [63]

    von Praun, V

    P.Charles, C.Grothoff, V.Saraswat, C.Donawa, A.Kielstra, K.Ebcioglu, C. von Praun, V. Sarkar, X10: an object-oriented approach to non- uniform cluster computing, SIGPLAN Not. 40 (10) (2005) 519–538. doi:10.1145/1103845.1094852

  55. [64]

    Chandra, Parallel programming in OpenMP, Morgan Kaufmann, 2001

    R. Chandra, Parallel programming in OpenMP, Morgan Kaufmann, 2001

  56. [65]

    Sanders, E

    J. Sanders, E. Kandrot, CUDA by example: an introduction to general- purpose GPU programming, Addison-Wesley Professional, 2010

  57. [66]

    Munshi, The OpenCL specification, in: 2009 IEEE Hot Chips 21 Symposium (HCS), 2009, pp

    A. Munshi, The OpenCL specification, in: 2009 IEEE Hot Chips 21 Symposium (HCS), 2009, pp. 1–314. doi:10.1109/HOTCHIPS.2009. 7478342

  58. [67]

    Bauer, M

    M. Bauer, M. Garland, Legate NumPy: Accelerated and Distributed Ar- ray Computing, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’19, As- sociation for Computing Machinery, New York, NY, USA, 2019, pp. 1–23...

  59. [68]

    Okuta, Y

    R. Okuta, Y. Unno, D. Nishino, S. Hido, C. Loomis, CuPy: A NumPy- Compatible Library for NVIDIA GPU Calculations (2017). URL http://learningsys.org/nips17/assets/papers/paper_16. pdf

  60. [69]

    Carter Edwards, C

    H. Carter Edwards, C. R. Trott, D. Sunderland, Kokkos: Enabling manycore performance portability through polymorphic memory ac- cess patterns, Journal of Parallel and Distributed Computing 74 (12) (2014)3202–3216, domain-SpecificLanguagesandHigh-LevelFrameworks for High-Perfor...

  61. [70]

    URL https://github.com/LLNL/RAJA 70

    LLNL, RAJA Performance Portability Layer (2019). URL https://github.com/LLNL/RAJA 70

  62. [71]

    Kotsifakou, P

    M. Kotsifakou, P. Srivastava, M. D. Sinclair, R. Komuravelli, V. Adve, S. Adve, HPVM: heterogeneous parallel virtual machine, SIGPLAN Not. 53 (1) (2018) 68–80.doi:10.1145/3200691.3178493

  63. [72]

    Bauer, S

    M. Bauer, S. Treichler, E. Slaughter, A. Aiken, Legion: expressing locality andindependencewithlogicalregions, in: ProceedingsoftheInternational Conference on High Performance Computing, Networking, Storage and Analysis, SC ’12, IEEE Computer Society Press, Washington, DC, USA...

  64. [73]

    Pouchet, U

    L.-N. Pouchet, U. Bondhugula, C. Bastoul, A. Cohen, J. Ramanujam, P. Sadayappan, N. Vasilache, Loop transformations: Convexity, pruning and optimization, in: Proceedings of the 38th Annual ACM SIGPLAN- SIGACT Symposium on Principles of Programming Languages, POPL ’11, ACM, New...

  65. [74]

    Bondhugula, A

    U. Bondhugula, A. Hartono, J. Ramanujam, P. Sadayappan, A practical automatic polyhedral parallelizer and locality optimizer, in: Proceedings of the 29th ACM SIGPLAN Conference on Programming Language De- sign and Implementation, PLDI ’08, ACM, New York, NY, USA, 2008, pp. 101...

  66. [75]

    Verdoolaege, J

    S. Verdoolaege, J. Carlos Juega, A. Cohen, J. Ignacio Gómez, C. Ten- llado, F. Catthoor, Polyhedral parallel code generation for cuda, ACM TransactionsonArchitectureandCodeOptimization(TACO)9(4)(2013) 54:1–54:23. doi:10.1145/2400682.2400713

  67. [76]

    T.Grosser, A.Größlinger, C.Lengauer, Polly-PerformingPolyhedralOp- timizations on a Low-Level Intermediate Representation, Parallel Process- ing Letters 22 (04) (2012) 1250010.doi:10.1142/S0129626412500107

  68. [77]

    Baghdadi, J

    R. Baghdadi, J. Ray, M. B. Romdhane, E. Del Sozzo, A. Akkas, Y. Zhang, P.Suriana, S.Kamil, S.Amarasinghe, Tiramisu: apolyhedralcompiler for expressing fast and portable code, in: Proceedings of the 2019 IEEE/ACM International Symposium on Code Generation and Optimization, CGO ...

  69. [78]

    M. M. Strout, Performance transformations for irregular applications, Ph.D. thesis, University of California, aAI3094622 (2003)

  70. [79]

    M. M. Strout, A. LaMielle, L. Carter, J. Ferrante, B. Kreaseck, C. Olschanowsky, An approach for code generation in the sparse poly- hedral framework, Parallel Computing 53 (2016) 32–57.doi:10.1016/ j.parco.2016.02.004

  71. [80]

    C. E. Oancea, L. Rauchwerger, A Hybrid Approach to Proving Mem- ory Reference Monotonicity, in: S. Rajopadhye, M. Mills Strout (Eds.), Languages and Compilers for Parallel Computing, Springer 71 Berlin Heidelberg, Berlin, Heidelberg, 2013, pp. 61–75. doi:10.1007/ 978-3-642-36036-7_5

  72. [81]

    S. Moon, M. W. Hall, Evaluation of Predicated Array Data-Flow Analysis for Automatic Parallelization, in: Proceedings of the Seventh ACM SIG- PLAN Symposium on Principles and Practice of Parallel Programming, PPoPP ’99, Association for Computing Machinery, New York, NY, USA, 1...

  73. [82]

    C. E. Oancea, A. Mycroft, Set-congruence dynamic analysis for thread- level speculation (tls), in: J. N. Amaral (Ed.), Languages and Compilers for Parallel Computing, Springer Berlin Heidelberg, Berlin, Heidelberg, 2008, pp. 156–171

  74. [83]

    F. Dang, H. Yu, L. Rauchwerger, The R-LRPD Test: Speculative Par- allelization of Partially Parallel Loops, in: Proceedings 16th Interna- tional Parallel and Distributed Processing Symposium, 2002, pp. 10 pp–. doi:10.1109/IPDPS.2002.1015493

  75. [84]

    M. Hall, C. Oancea, A. C. Elster, A. Rasch, S. Joshi, A. M. Tavakkoli, R. Schulze, Scheduling languages: A past, present, and future taxonomy (2024). arXiv:2410.19927. URL https://arxiv.org/abs/2410.19927

  76. [85]

    Donadio, J

    S. Donadio, J. Brodman, T. Roeder, K. Yotov, D. Barthou, A. Cohen, M. J. Garzarán, D. Padua, K. Pingali, A language for the compact repre- sentation of multiple program versions, in: E. Ayguadé, G. Baumgartner, J. Ramanujam, P. Sadayappan (Eds.), Languages and Compilers for Pa...

  77. [86]

    C. Chen, J. Chame, M. W. Hall, Chill : A framework for composing high-level loop transformations, Tech. rep., Technical Report 08-897, U. of Southern California (2008)

  78. [87]

    Girbal, N

    S. Girbal, N. Vasilache, C. Bastoul, A. Cohen, D. Parello, M. Sigler, O. Temam, Semi-automatic composition of loop transformations for deep parallelism and memory hierarchies, International Journal of Parallel Pro- gramming 34 (Jun. 2006).doi:10.1007/s10766-006-0012-3

  79. [88]

    Ragan-Kelley, C

    J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, S. Amaras- inghe, Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines, SIGPLAN Not. 48 (6) (2013) 519–530. doi:10.1145/2499370.2462176

  80. [89]

    R. T. Mullapudi, V. Vasista, U. Bondhugula, Polymage: Automatic op- timization for image processing pipelines, in: Proceedings of the Twen- tieth International Conference on Architectural Support for Program- ming Languages and Operating Systems, ASPLOS ’15, Association for 72...

  81. [90]

    Hegarty, J

    J. Hegarty, J. Brunhaver, Z. DeVito, J. Ragan-Kelley, N. Cohen, S. Bell, A. Vasilyev, M. Horowitz, P. Hanrahan, Darkroom: Compiling high-level image processing code into hardware pipelines, ACM Trans. Graph. 33 (4) (2014) 144:1–144:11. doi:10.1145/2601097.2601174

  82. [91]

    Nelson, A

    T. Nelson, A. Rivera, P. Balaprakash, M. Hall, P. D. Hovland, E. Jessup, B. Norris, Generating efficient tensor contractions for gpus, in: 2015 44th International Conference on Parallel Processing, 2015, pp. 969–978.doi: 10.1109/ICPP.2015.106

  83. [92]

    Yadav, A

    R. Yadav, A. Aiken, F. Kjolstad, DISTAL: the distributed tensor alge- bra compiler, in: Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation, PLDI 2022, Association for Computing Machinery, New York, NY, USA, 2022, p. 2...

  84. [93]

    Venkat, M

    A. Venkat, M. Hall, M. Strout, Loop and data transformations for sparse matrix code, in: Proceedings of the 36th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI ’15, Associa- tion for Computing Machinery, New York, NY, USA, 2015, p. 521–532

  85. [94]

    Venkat, M

    A. Venkat, M. S. Mohammadi, J. Park, H. Rong, R. Barik, M. M. Strout, M. Hall, Automating wavefront parallelization for sparse matrix compu- tations, in: Proceedings of the International Conference for High Per- formance Computing, Networking, Storage and Analysis, SC ’16, IEE...

  86. [95]

    Senanayake, C

    R. Senanayake, C. Hong, Z. Wang, A. Wilson, S. Chou, S. Kamil, S. Ama- rasinghe, F. Kjolstad, A sparse iteration space transformation framework for sparse tensor algebra, Proc. ACM Program. Lang. 4 (OOPSLA) (Nov. 2020). doi:10.1145/3428226

  87. [96]

    Bansal, O

    M. Bansal, O. Hsu, K. Olukotun, F. Kjolstad, Mosaic: An interoperable compiler for tensor algebra, Proc. ACM Program. Lang. 7 (PLDI) (Jun. 2023). doi:10.1145/3591236

  88. [97]

    Tillet, H

    P. Tillet, H. T. Kung, D. Cox, Triton: an intermediate language and compiler for tiled neural network computations, in: Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL 2019, Association for Computing Ma- chinery, Ne...

  89. [98]

    T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L.Wang, Y.Hu, L.Ceze, etal.,{TVM}: Anautomated{End-to-End}op- timizing compiler for deep learning, in: Proceedings of the 13th USENIX 73 Conference on Operating Systems Design and Implementation, OSDI’18, USENI...

  90. [99]

    Venkat, T

    A. Venkat, T. Rusira, R. Barik, M. Hall, L. Truong, Swirl: High- performance many-core cpu code generation for deep neural networks, The International Journal of High Performance Computing Applications 33 (6) (2019) 1275–1289.doi:10.1177/1094342019866247

  91. [100]

    Paszke, D

    A. Paszke, D. D. Johnson, D. Duvenaud, D. Vytiniotis, A. Radul, M. J. Johnson, J. Ragan-Kelley, D. Maclaurin, Getting to the point: index sets and parallelism-preserving autodiff for pointful array programming, Proc. ACM Program. Lang. 5 (ICFP) (Aug. 2021).doi:10.1145/3473593

  92. [101]

    Steuwer, T

    M. Steuwer, T. Remmelg, C. Dubach, Lift: a functional data-parallel IR for high-performance GPU code generation, in: Proceedings of the 2017 International Symposium on Code Generation and Optimization, CGO ’17, IEEE Press, 2017, p. 74–85

  93. [102]

    Steuwer, T

    M. Steuwer, T. Koehler, B. Köpcke, F. Pizzuti, RISE & shine: Language- oriented compiler design, CoRR abs/2201.03611 (2022). arXiv:2201. 03611. URL https://arxiv.org/abs/2201.03611

  94. [103]

    61–72.doi:10.1145/3578360.3580269

    A.Rasch, R.Schulze, D.Shabalin, A.Elster, S.Gorlatch, M.Hall, (de/re)- compositions expressed systematically via mdh-based schedules, in: Pro- ceedings of the 32nd ACM SIGPLAN International Conference on Com- piler Construction, CC 2023, Association for Computing Machinery, Ne...

  95. [104]

    Kundefinedhler, A

    T. Kundefinedhler, A. Goens, S. Bhat, T. Grosser, P. Trinder, M. Steuwer, Guided Equality Saturation, Proc. ACM Program. Lang. 8 (POPL) (Jan. 2024). doi:10.1145/3632900

  96. [105]

    Holk, Region-based memory management for expressive gpu program- ming, Ph.D

    E. Holk, Region-based memory management for expressive gpu program- ming, Ph.D. thesis, Indiana University (2016). 74

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.