Khattak, F. and Mikaitis, M. orcid.org/0000-0001-8706-1436 (2026) Accurate Models of NVIDIA Tensor Cores. ACM Transactions on Architecture and Code Optimization (TACO). ISSN: 1544-3566
Abstract
Matrix multiplication is a fundamental operation in both training of neural networks and inference. To accelerate matrix multiplication, Graphical Processing Units (GPUs) provide it implemented in hardware. Due to the increased throughput over the software-based matrix multiplication, the multipliers are increasingly used outside AI, to accelerate various applications in scientific computing. However, matrix multipliers targeted at AI are at present not compliant with IEEE 754 floating-point arithmetic behaviour, with different vendors offering different numerical features. This leads to non-reproducible results across different generations of GPU architectures, at the matrix multiply-accumulate instruction level. To study numerical characteristics of matrix multipliers—such as rounding behaviour, accumulator width, normalization points, extra carry bits, and others—test vectors are typically constructed. Yet, these vectors may or may not distinguish between different hardware models, and due to limited hardware availability, their reliability across many different platforms remains largely untested. We present software models for emulating the inner product behavior of low- and mixed-precision matrix multipliers in the V100, A100, H100 and B200 data center GPUs in most supported input formats of interest to mixed-precision algorithm developers: 8-, 16-, and 19-bit floating point. These matrix multiplier models are first approximated by determining the numerical features via test vectors designed to trigger outputs sensitive to bit level differences in the implementation, followed by semi-exhaustive comparison (randomised input vectors of 107 values) between the models and the actual GPU matrix multipliers—this process is repeated until the model is bit-accurate. These models enable verification of test vectors before applying them to real hardware and also support computational scientists and mixed-precision algorithm developers with easy-to-use accurate models available in MATLAB—we demonstrate their use in multi-word emulation algorithms for matrix multiplication.
Metadata
| Item Type: | Article |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | © 2026 Copyright held by the owner/author(s). This is an open access article under the terms of the Creative Commons Attribution License (CC-BY 4.0), which permits unrestricted use, distribution and reproduction in any medium, provided the original work is properly cited. |
| Keywords: | Tensor cores, mixed-precision computing, matrix multiply, inner product, IEEE 754 standard arithmetic |
| Dates: |
|
| Institution: | The University of Leeds |
| Academic Units: | The University of Leeds > Faculty of Engineering & Physical Sciences (Leeds) > School of Computing (Leeds) > Computation Science & Engineering |
| Funding Information: | Funder Grant number EPSRC Accounts Payable OPP136 |
| Date Deposited: | 14 Jul 2026 13:37 |
| Last Modified: | 14 Jul 2026 13:37 |
| Status: | Published online |
| Publisher: | Association for Computing Machinery (ACM) |
| Identification Number: | 10.1145/3830409 |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:243066 |
Download
Filename: 3830409.pdf
Licence: CC-BY 4.0

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)