Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training

Li, Shenggui; Fang, Jiarui; Bian, Zhengda; Liu, Hongxin; Liu, Yuliang; Huang, Haichen; Wang, Boxiang; You, Yang

Computer Science > Machine Learning

arXiv:2110.14883v2 (cs)

[Submitted on 28 Oct 2021 (v1), revised 20 Sep 2022 (this version, v2), latest version 5 Oct 2023 (v3)]

Title:Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training

Authors:Shenggui Li, Jiarui Fang, Zhengda Bian, Hongxin Liu, Yuliang Liu, Haichen Huang, Boxiang Wang, Yang You

View PDF

Abstract:The success of Transformer models has pushed the deep learning model scale to billions of parameters. Due to the limited memory resource of a single GPU, However, the best practice for choosing the optimal parallel strategy is still lacking, since it requires domain expertise in both deep learning and parallel computing.
The Colossal-AI system addressed the above challenge by introducing a unified interface to scale your sequential code of model training to distributed environments. It supports parallel training methods such as data, pipeline, tensor, and sequence parallelism, as well as heterogeneous training methods integrated with zero redundancy optimizer. Compared to the baseline system, Colossal-AI can achieve up to 2.76 times training speedup on large-scale models.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as:	arXiv:2110.14883 [cs.LG]
	(or arXiv:2110.14883v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2110.14883

Submission history

From: Yang You [view email]
[v1] Thu, 28 Oct 2021 04:45:55 UTC (214 KB)
[v2] Tue, 20 Sep 2022 12:54:20 UTC (665 KB)
[v3] Thu, 5 Oct 2023 04:09:09 UTC (620 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2021-10

Change to browse by:

cs
cs.AI
cs.CL
cs.CV
cs.DC

References & Citations

DBLP - CS Bibliography

listing | bibtex

Boxiang Wang
Yongbin Li
Fan Cui
Yang You

export BibTeX citation

Computer Science > Machine Learning

Title:Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators