Mic*_*mza 1 python compression gzip
是否可以通过一定量的流来 gzip 数据,即无需立即将所有压缩数据加载到内存中?
例如,我可以在具有 2GB 内存的计算机上对一个 10GB 大小的文件进行 gzip 压缩吗?
在https://docs.python.org/3/library/gzip.html#gzip.compress,该gzip.compress函数返回 gzip 的字节,因此必须全部加载到内存中。但是......尚不清楚gzip.open内部是如何工作的:压缩的字节是否会同时全部存储在内存中。gzip 格式本身是否使得实现流式 gzip 变得特别困难?
[这个问题用Python标记,但也欢迎非Python答案]
您不必一次压缩所有 10GB。您可以分块读取输入数据,并单独压缩每个块,因此不必一次全部装入内存。
chunksize = 100 * 1024 * 1024 # 100 mb chunks
with open("bigfile.txt") as infile:
while True:
chunk = infile.read(chunksize)
if not chunk:
break
compressed = gzip.compress(chunk)
# do something with compressed
Run Code Online (Sandbox Code Playgroud)
如果您要创建压缩文件,则可以将块直接写入 gzip 文件。
with open("bigfile.txt") as infile, gzip.open("bigfile.txt.gz", "w") as gzipfile:
while True:
chunk = infile.read(chunksize)
if not chunk:
break
gzipfile.write(chunk)
Run Code Online (Sandbox Code Playgroud)
[这是基于@Barmar的回答和评论]
可以实现流式gzip压缩。gzip模块使用zlib来实现流压缩,并且查看gzip 模块源代码,它似乎没有将所有输出字节加载到内存中。
您还可以直接使用 zlib 模块来执行此操作,例如使用小型生成器管道:
import zlib
def yield_uncompressed_bytes():
# In a real case, would yield bytes pulled from the filesystem or the network
chunk = b'*' * 65000
for _ in range(0, 10000):
print('In: ', len(chunk))
yield chunk
def yield_compressed_bytes(_uncompressed_bytes):
compress_obj = zlib.compressobj(wbits=zlib.MAX_WBITS + 16)
for chunk in _uncompressed_bytes:
if compressed_bytes := compress_obj.compress(chunk):
yield compressed_bytes
if compressed_bytes := compress_obj.flush():
yield compressed_bytes
uncompressed_bytes = yield_uncompressed_bytes()
compressed_bytes = yield_compressed_bytes(uncompressed_bytes)
for chunk in compressed_bytes:
# In a real case, could save to the filesystem, or send over the network
print('Out:', len(chunk))
Run Code Online (Sandbox Code Playgroud)
您可以看到 与In:散布在一起Out:,表明 zlib compressobj 确实没有将整个输出存储在内存中。
| 归档时间: |
|
| 查看次数: |
3275 次 |
| 最近记录: |