An FPGA-Based Optimizer Design for Distributed Deep Learning with Multiple GPUs

1 December 2021

journal article
research article
Published by Institute of Electronics, Information and Communications Engineers (IEICE) in IEICE Transactions on Information and Systems

Vol. E104.D (12), 2057-2067
https://doi.org/10.1587/transinf.2021pap0008

Abstract

Since deep learning workloads perform a large number of matrix operations on training data, GPUs (Graphics Processing Units) are efficient especially for the training phase. A cluster of computers each of which equips multiple GPUs can significantly accelerate the deep learning workloads. More specifically, a back-propagation algorithm following a gradient descent approach is used for the training. Although the gradient computation is still a major bottleneck of the training, gradient aggregation and optimization impose both communication and computation overheads, which should also be reduced for further shortening the training time. To address this issue, in this paper, multiple GPUs are interconnected with a PCI Express (PCIe) over 10Gbit Ethernet (10GbE) technology. Since these remote GPUs are interconnected with network switches, gradient aggregation and optimizers (e.g., SGD, AdaGrad, Adam, and SMORMS3) are offloaded to FPGA-based 10GbE switches between remote GPUs; thus, the gradient aggregation and parameter optimization are completed in the network. The proposed FPGA-based 10GbE switches with the four optimizers are implemented on NetFPGA-SUME board. Their resource utilizations are increased by PEs for the optimizers, and they consume up to 56% of the resources. Evaluation results using four remote GPUs connected via the proposed FPGA-based switch demonstrate that these optimizers are accelerated by up to 3.0x and 1.25x compared to CPU and GPU implementations, respectively. Also, the gradient aggregation throughput by the FPGA-based switch achieves up to 98.3% of the 10GbE line rate.

Keywords

This publication has 4 references indexed in Scilit:

Accelerating Deep Learning using Multiple GPUs and FPGA-Based 10GbE Switch
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2020
A Network-Centric Hardware/Algorithm Co-Design to Accelerate Distributed Training of Deep Neural Networks
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2018
Densely Connected Convolutional Networks
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2017
End-to-End Adaptive Packet Aggregation for High-Throughput I/O Bus Network Using Ethernet
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2014