gzip 是否可以压缩数据而不将其全部加载到内存中,即流式传输/即时传输?

Mic*_*mza 1 python compression gzip

是否可以通过一定量的流来 gzip 数据,即无需立即将所有压缩数据加载到内存中?

例如,我可以在具有 2GB 内存的计算机上对一个 10GB 大小的文件进行 gzip 压缩吗?

在https://docs.python.org/3/library/gzip.html#gzip.compress,该gzip.compress函数返回 gzip 的字节,因此必须全部加载到内存中。但是......尚不清楚gzip.open内部是如何工作的:压缩的字节是否会同时全部存储在内存中。gzip 格式本身是否使得实现流式 gzip 变得特别困难?

[这个问题用Python标记,但也欢迎非Python答案]

Bar*_*mar 7

您不必一次压缩所有 10GB。您可以分块读取输入数据,并单独压缩每个块,因此不必一次全部装入内存。

chunksize = 100 * 1024 * 1024 # 100 mb chunks
with open("bigfile.txt") as infile:
    while True:
        chunk = infile.read(chunksize)
        if not chunk:
            break
        compressed = gzip.compress(chunk)
        # do something with compressed
Run Code Online (Sandbox Code Playgroud)

如果您要创建压缩文件,则可以将块直接写入 gzip 文件。

with open("bigfile.txt") as infile, gzip.open("bigfile.txt.gz", "w") as gzipfile:
    while True:
        chunk = infile.read(chunksize)
        if not chunk:
            break
        gzipfile.write(chunk)
Run Code Online (Sandbox Code Playgroud)


Mic*_*mza 5

[这是基于@Barmar的回答和评论]

可以实现流式gzip压缩。gzip模块使用zlib来实现流压缩,并且查看gzip 模块源代码,它似乎没有将所有输出字节加载到内存中。

您还可以直接使用 zlib 模块来执行此操作,例如使用小型生成器管道:

import zlib

def yield_uncompressed_bytes():
    # In a real case, would yield bytes pulled from the filesystem or the network
    chunk = b'*' * 65000
    for _ in range(0, 10000):
        print('In: ', len(chunk))
        yield chunk

def yield_compressed_bytes(_uncompressed_bytes):
    compress_obj = zlib.compressobj(wbits=zlib.MAX_WBITS + 16)
    for chunk in _uncompressed_bytes:
        if compressed_bytes := compress_obj.compress(chunk):
            yield compressed_bytes

    if compressed_bytes := compress_obj.flush():
        yield compressed_bytes

uncompressed_bytes = yield_uncompressed_bytes()
compressed_bytes = yield_compressed_bytes(uncompressed_bytes)

for chunk in compressed_bytes:
    # In a real case, could save to the filesystem, or send over the network
    print('Out:', len(chunk))
Run Code Online (Sandbox Code Playgroud)

您可以看到 与In:散布在一起Out:,表明 zlib compressobj 确实没有将整个输出存储在内存中。