Trained Ternary Quantization

Zhu, Chenzhuo; Han, Song; Mao, Huizi; Dally, William J.

Computer Science > Machine Learning

arXiv:1612.01064v2 (cs)

[Submitted on 4 Dec 2016 (v1), revised 23 Jan 2017 (this version, v2), latest version 23 Feb 2017 (v3)]

Title:Trained Ternary Quantization

Authors:Chenzhuo Zhu, Song Han, Huizi Mao, William J. Dally

View PDF

Abstract:Deep neural networks are widely used in machine learning applications. However, the deployment of large neural networks models can be difficult to deploy on mobile devices with limited power budgets. To solve this problem, we propose Trained Ternary Quantization (TTQ), a method that can reduce the precision of weights in neural networks to ternary values. This method has very little accuracy degradation and can even improve the accuracy of some models (32, 44, 56-layer ResNet) on CIFAR-10 and AlexNet on ImageNet. And our AlexNet model is trained from scratch, which means it's as easy as to train normal full precision model. We highlight our trained quantization method that can learn both ternary values and ternary assignment. During inference, only ternary values (2-bit weights) and scaling factors are needed, therefore our models are nearly 16x smaller than full-precision models. Our ternary models can also be viewed as sparse binary weight networks, which can potentially be accelerated with custom circuit. Experiments on CIFAR-10 show that the ternary models obtained by trained quantization method outperform full-precision models of ResNet-32,44,56 by 0.04%, 0.16%, 0.36%, respectively. On ImageNet, our model outperforms full-precision AlexNet model by 0.3% of Top-1 accuracy and outperforms previous ternary models by 3%.

Comments:	Submitted to ICLR 17
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:1612.01064 [cs.LG]
	(or arXiv:1612.01064v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1612.01064

Submission history

From: Chenzhuo Zhu [view email]
[v1] Sun, 4 Dec 2016 05:00:22 UTC (288 KB)
[v2] Mon, 23 Jan 2017 03:27:20 UTC (289 KB)
[v3] Thu, 23 Feb 2017 06:52:28 UTC (289 KB)

Computer Science > Machine Learning

Title:Trained Ternary Quantization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Trained Ternary Quantization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators