Chen, Y. and Xia, C. orcid.org/0000-0003-2014-5453 (2026) From Naive CUDA to Triton: A Systematic Evaluation of AI-Era HPC Operator Development. In: ICS Workshops '26: Proceedings of the 40th ACM International Conference on Supercomputing - Workshops. ICS Workshops '26: 2026 International Conference on Supercomputing Workshops, 06-09 Jul 2026, Belfast, Northern Ireland, United Kingdom. Association for Computing Machinery (ACM), pp. 60-64. ISBN: 979-8-4007-2300-1.
Abstract
High-performance operator development is central to AI-driven scientific computing, yet manual CUDA optimization remains time-consuming and expertise-intensive. Large language models (LLMs) and domain-specific languages such as OpenAI Triton have recently emerged as promising alternatives for GPU kernel development. This paper presents a systematic evaluation of these programming paradigms using a four operators considering computational intensity and memory access regularity. We study four representative kernels from real HPC applications: Stencil (miniGhost), SpMV (Siesta), GEMM (GAMESS), and N-Body molecular dynamics.
Experiments on an NVIDIA RTX 3090 show that Triton performs strongly on regular compute-bound workloads, outperforming cuBLAS by 1.28 × for GEMM through auto-tuning and software pipelining. However, its performance degrades on irregular access patterns and collapses in all-to-all interaction patterns: Triton-based N-Body is 13.5 × slower than naive CUDA. We further show that high SM occupancy does not necessarily imply high performance, and that compute-unit utilization provides a more reliable indicator of kernel efficiency. These results characterize where LLM-assisted CUDA and Triton are effective, where expert-level optimized CUDA remains necessary.
Metadata
| Item Type: | Proceedings Paper |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | Copyright © 2026 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution 4.0 International License. |
| Keywords: | HPC, GPU, CUDA, Triton, Large Language Models, Performance Evaluation, Benchmarks |
| Dates: |
|
| Institution: | The University of Leeds |
| Academic Units: | The University of Leeds > Faculty of Engineering & Physical Sciences (Leeds) > School of Computing (Leeds) > Distributed Systems & Services |
| Date Deposited: | 04 Aug 2026 15:18 |
| Last Modified: | 04 Aug 2026 15:18 |
| Published Version: | https://dl.acm.org/doi/10.1145/3774895.3812198 |
| Status: | Published |
| Publisher: | Association for Computing Machinery (ACM) |
| Identification Number: | 10.1145/3774895.3812198 |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:244123 |
Download
Filename: From Naive CUDA to Triton A Systematic Evaluation of AI-EraHPC Operator Development.pdf
Licence: CC-BY 4.0

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)