Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks

Tejas Raja

Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks

Tejas Raja

TL;DR

The results show that selecting the right data flow for specific matrix configurations can drastically reduce energy consumption, and provide helpful insights into optimizing hardware for AI and machine learning applications, offering potential improvements in designing energy-efficient DNN accelerators.

Abstract

The paper discusses how Systolic Arrays can improve matrix multiplication for deep neural networks (DNNs). With AI models like OpenAI's GPT now containing trillions of parameters, the need for efficient matrix multiplication is more critical than ever. In this paper, the three main systolic array data flows: Weight Stationary (WS), Input Stationary (IS), and Output Stationary (OS) are discussed. Each data flow's energy consumption and efficiency across various matrix sizes are calculated using the SCALE-Sim simulator. The results show that selecting the right data flow for specific matrix configurations can drastically reduce energy consumption. The conclusions provide helpful insights into optimizing hardware for AI and machine learning applications, offering potential improvements in designing energy-efficient DNN accelerators.

Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks

TL;DR

Abstract

Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (6)