Julia - 读取大文件的并行性

JKH*_*KHA 3 parallel-processing multithreading julia

在 Julia v1.1 中,假设我有一个非常大的文本文件 (30GB) 并且我想要并行性(多线程)来读取每一行,我该怎么办?

这段代码是在检查Julia 的 multi-threading 文档后尝试这样做的,但它根本不起作用

open("pathtofile", "r") do file
    # Count number of lines in file
    seekend(file)
    fileSize = position(file)
    seekstart(file)

    # skip nseekchars first characters of file
    seek(file, nseekchars)

    # progress bar, because it's a HUGE file
    p = Progress(fileSize, 1, "Reading file...", 40)
    Threads.@threads for ln in eachline(file)
        # do something on ln
        u, v = map(x->parse(UInt32, x), split(ln))
        .... # other interesting things
        update!(p, position(file))
    end    
end
Run Code Online (Sandbox Code Playgroud)

注意 1 :您需要using ProgressMeter(我希望我的代码在并行读取文件时显示进度条)

注 2:nseekchars 是一个 Int 和我想在文件开头跳过的字符数

注意 3:代码正在运行,但Threads.@threads在 for 循环旁边没有宏的情况下不会执行并行

Prz*_*fel 5

为了获得最大的 I/O 性能:

  1. 并行化硬件 - 即使用磁盘阵列而不是单个驱动器。尝试搜索raid 性能以获得许多出色的解释(或提出单独的问题)

  2. 使用 Julia内存映射机制

s = open("my_file.txt","r")
using Mmap
a = Mmap.mmap(s)
Run Code Online (Sandbox Code Playgroud)
  1. 一旦有了内存映射,并行处理。注意线程的错误共享(取决于您的实际情况)。