Fir*_*ger 1 python sorting python-3.x
我有一个大小合适的.tsv文件,其中包含以下格式的文档
ID DocType NormalizedName DisplayName Year Description
12648 Book a fancy title A FaNcY-Title 2005 This is a short description of the book
1867453 Essay on the history of humans On the history of humans 2016 This is another short description, this time of the essay
...
Run Code Online (Sandbox Code Playgroud)
该文件的压缩版本大小约为 67 GB,压缩后约为 22 GB。
我想根据 ID(大约 3 亿行)按升序对文件的行进行排序。每行的 ID 都是唯一的,范围为 1 - 2147483647(正数部分long),可能存在间隙。
不幸的是,我最多只有 8GB 可用内存,所以我无法一次加载整个文件。
对该列表进行排序并将其写回磁盘的最省时的方式是什么?
我使用以下方法进行了概念验证heapq.merge:
步骤一:生成测试文件
生成包含3亿行的测试文件:
from random import randint
row = '{} Essay on the history of humans On the history of humans 2016 This is another short description, this time of the essay\n'
with open('large_file.tsv', 'w') as f_out:
for i in range(300_000_000):
f_out.write(row.format(randint(1, 2147483647)))
Run Code Online (Sandbox Code Playgroud)
步骤2:分成块并对每个块进行排序
每个块有 100 万行:
import glob
path = "chunk_*.tsv"
chunksize = 1_000_000
fid = 1
lines = []
with open('large_file.tsv', 'r') as f_in:
f_out = open('chunk_{}.tsv'.format(fid), 'w')
for line_num, line in enumerate(f_in, 1):
lines.append(line)
if not line_num % chunksize:
lines = sorted(lines, key=lambda k: int(k.split()[0]))
f_out.writelines(lines)
print('splitting', fid)
f_out.close()
lines = []
fid += 1
f_out = open('chunk_{}.tsv'.format(fid), 'w')
# last chunk
if lines:
print('splitting', fid)
lines = sorted(lines, key=lambda k: int(k.split()[0]))
f_out.writelines(lines)
f_out.close()
lines = []
Run Code Online (Sandbox Code Playgroud)
第三步:合并每个块
from heapq import merge
chunks = []
for filename in glob.glob(path):
chunks += [open(filename, 'r')]
with open('sorted.tsv', 'w') as f_out:
f_out.writelines(merge(*chunks, key=lambda k: int(k.split()[0])))
Run Code Online (Sandbox Code Playgroud)
时间:
我的机器是Ubuntu Linux 18.04,AMD 2400G,便宜的WD SSD Green)
第 2 步- 分割和排序块 - 花费约 12 分钟
第 3 步- 合并块 - 大约需要10 分钟
我预计这些值在具有更好磁盘(NVME?)和 CPU 的机器上要低得多。
| 归档时间: |
|
| 查看次数: |
2888 次 |
| 最近记录: |