GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Huang, Yanping; Cheng, Yonglong; Chen, Dehao; Lee, HyoukJoong; Ngiam, Jiquan; Le, Quoc V.; Chen, Zhifeng

Computer Science > Computer Vision and Pattern Recognition

arXiv:1811.06965v1 (cs)

[Submitted on 16 Nov 2018 (this version), latest version 25 Jul 2019 (v5)]

Title:GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Authors:Yanping Huang, Yonglong Cheng, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Zhifeng Chen

View PDF

Abstract:GPipe is a scalable pipeline parallelism library that enables learning of giant deep neural networks. It partitions network layers across accelerators and pipelines execution to achieve high hardware utilization. It leverages recomputation to minimize activation memory usage. For example, using partitions over 8 accelerators, it is able to train networks that are 25x larger, demonstrating its scalability. It also guarantees that the computed gradients remain consistent regardless of the number of partitions. It achieves an almost linear speed up without any changes in the model parameters: when using 4x more accelerators, training the same model is up to 3.5x faster. We train a 557 million parameters AmoebaNet model on ImageNet and achieve a new state-of-the-art 84.3% top-1 / 97.0% top-5 accuracy on ImageNet. Finally, we use this learned model as an initialization for training 7 different popular image classification datasets and obtain results that exceed the best published ones on 5 of them, including pushing the CIFAR-10 accuracy to 99% and CIFAR-100 accuracy to 91.3%.

Comments:	10 pages. Work in progress. Copyright 2018 by the authors
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1811.06965 [cs.CV]
	(or arXiv:1811.06965v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1811.06965

Submission history

From: Yanping Huang [view email]
[v1] Fri, 16 Nov 2018 18:43:28 UTC (653 KB)
[v2] Mon, 19 Nov 2018 18:32:58 UTC (542 KB)
[v3] Tue, 20 Nov 2018 17:25:46 UTC (543 KB)
[v4] Wed, 12 Dec 2018 17:45:02 UTC (544 KB)
[v5] Thu, 25 Jul 2019 21:42:58 UTC (779 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Submission history

Access Paper:

References & Citations

2 blog links

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Submission history

Access Paper:

References & Citations

2 blog links

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators