A Parallel Scan Algorithm in the Tensor Core Unit Model

Anastasios Zouzias; William F. McColl

A Parallel Scan Algorithm in the Tensor Core Unit Model

Anastasios Zouzias, William F. McColl

TL;DR

A parallel scan (prefix sum) algorithm in the Tensor Core Unit (TCU) model of computation that performs multiplications of square matrices of size s and has depth at most $2\lfloor \log_s (n)$ for inputs of size n.

Abstract

We present a parallel scan (prefix sum) algorithm in the Tensor Core Unit (TCU) model of computation. The TCU model assumes that multiplication between two square matrices of constant size $s$ is a basic operation. In the $(s^2, \ell)$-TCU model, we show that for inputs of size $n$, the algorithm has depth at most $2\lfloor \log_s (n)\rfloor$ and runs in $O(n(1 + \ell /s^2)/p + (s^2 + \ell) \log_s (n))$ time assuming $p$ tensor core units. Equivalently, the algorithm performs $O(n/s^2)$ multiplications of square matrices of size s.

A Parallel Scan Algorithm in the Tensor Core Unit Model

TL;DR

A parallel scan (prefix sum) algorithm in the Tensor Core Unit (TCU) model of computation that performs multiplications of square matrices of size s and has depth at most

for inputs of size n.

Abstract

We present a parallel scan (prefix sum) algorithm in the Tensor Core Unit (TCU) model of computation. The TCU model assumes that multiplication between two square matrices of constant size

is a basic operation. In the

-TCU model, we show that for inputs of size

, the algorithm has depth at most

and runs in

time assuming

tensor core units. Equivalently, the algorithm performs

multiplications of square matrices of size s.

A Parallel Scan Algorithm in the Tensor Core Unit Model

TL;DR

Abstract

A Parallel Scan Algorithm in the Tensor Core Unit Model

TL;DR

Abstract

Paper Structure

Table of Contents

Key Result

Figures (3)

Theorems & Definitions (5)