Understanding is Compression

Li, Ziguang; Huang, Chao; Wang, Xuliang; Hu, Haibo; Wyeth, Cole; Bu, Dongbo; Yu, Quan; Gao, Wen; Liu, Xingwu; Li, Ming

Computer Science > Information Theory

arXiv:2407.07723 (cs)

[Submitted on 24 Jun 2024 (v1), last revised 21 Aug 2024 (this version, v2)]

Title:Understanding is Compression

Authors:Ziguang Li, Chao Huang, Xuliang Wang, Haibo Hu, Cole Wyeth, Dongbo Bu, Quan Yu, Wen Gao, Xingwu Liu, Ming Li

View PDF HTML (experimental)

Abstract:Modern data compression methods are slowly reaching their limits after 80 years of research, millions of papers, and wide range of applications. Yet, the extravagant 6G communication speed requirement raises a major open question for revolutionary new ideas of data compression.
We have previously shown all understanding or learning are compression, under reasonable assumptions. Large language models (LLMs) understand data better than ever before. Can they help us to compress data?
The LLMs may be seen to approximate the uncomputable Solomonoff induction. Therefore, under this new uncomputable paradigm, we present LMCompress. LMCompress shatters all previous lossless compression algorithms, doubling the lossless compression ratios of JPEG-XL for images, FLAC for audios, and H.264 for videos, and quadrupling the compression ratio of bz2 for texts. The better a large model understands the data, the better LMCompress compresses.

Subjects:	Information Theory (cs.IT); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2407.07723 [cs.IT]
	(or arXiv:2407.07723v2 [cs.IT] for this version)
	https://doi.org/10.48550/arXiv.2407.07723

Submission history

From: Xingwu Liu [view email]
[v1] Mon, 24 Jun 2024 03:58:11 UTC (250 KB)
[v2] Wed, 21 Aug 2024 02:45:36 UTC (543 KB)

Computer Science > Information Theory

Title:Understanding is Compression

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Theory

Title:Understanding is Compression

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators