Numba 快速数学并不能提高速度

use*_*621 3 math cpu performance numpy numba

fastmath我在启用和禁用选项的情况下运行以下代码。

import numpy as np
from numba import jit
from threading import Thread
import time
import psutil
from tqdm import tqdm


@jit(nopython=True, fastmath=True)
def compute_angle(vectors):
    return 180 + np.degrees(np.arctan2(vectors[:, :, 1], vectors[:, :, 0]))



cpu_usage = list()
times = list()

# Log cpu usage
running = False
def threaded_function():
    while not running:
        time.sleep(0.1)
    print("Start logging CPU")
    while running:
        cpu_usage.append(psutil.cpu_percent())
    print("Stop logging CPU")
thread = Thread(target=threaded_function, args=())
thread.start()


iterations = 1000

# Generate frames
vectors_list = list()
for i in tqdm(range(iterations), total=iterations):
    vectors = np.random.randint(-50, 50, (500, 1000, 2))
    vectors_list.append(vectors)

for i in tqdm(range(iterations), total=iterations):
    s = time.time()
    compute_angle(vectors_list[i])
    e = time.time()
    times.append(e - s)
    # Do not count first iteration
    running = True

running = False

thread.join()

print("Average time per iteration", np.mean(times[1:]))
print("Average CPU usage:", np.mean(cpu_usage))
Run Code Online (Sandbox Code Playgroud)

结果fastmath=True是:

Average time per iteration 0.02076407738992044
Average CPU usage: 6.738916256157635`
Run Code Online (Sandbox Code Playgroud)

结果fastmath=False是:

Average time per iteration 0.020854528721149738
Average CPU usage: 6.676455696202531
Run Code Online (Sandbox Code Playgroud)

由于我使用的是数学运算,我应该期待一些收益吗?我也尝试安装icc-rt,但我不知道如何检查它是否启用。谢谢你!

max*_*111 5

要使 SIMD 矢量化正常工作,还缺少一些东西。为了获得最大性能,还必须避免昂贵的临时数组,如果您使用部分矢量化函数,则可能无法对其进行优化。

\n
    \n
  • 函数调用必须内联
  • \n
  • 内存访问模式必须在编译时已知。在下面的示例中,这是通过 完成的assert vectors.shape[2]==2。一般来说,最后一个数组的形状也可能大于 2,这对于 SIMD 向量化来说会复杂得多。
  • \n
  • 除以零检查也可以避免 SIMD 向量化,如果不优化它们的话,速度会很慢。我通过计算div_pi=1/np.pi一次而不是循环内的简单乘法来手动完成此操作。如果重复除法无法避免,您可以使用error_model="numpy"零检查来避免除法。
  • \n
\n

例子

\n
import numpy as np\nimport numba as nb\n\n@nb.njit(fastmath=True)\ndef your_function(vectors):\n    return 180 + np.degrees(np.arctan2(vectors[:, :, 1], vectors[:, :, 0]))\n\n@nb.njit(fastmath=True)#False\ndef optimized_function(vectors):\n    assert vectors.shape[2]==2\n\n    res=np.empty((vectors.shape[0],vectors.shape[1]),dtype=vectors.dtype)\n    div_pi=180/np.pi\n    for i in range(vectors.shape[0]):\n        for j in range(vectors.shape[1]):\n            res[i,j]=np.arctan2(vectors[i,j,1],vectors[i,j,0])*div_pi+180\n    return res\n
Run Code Online (Sandbox Code Playgroud)\n

时间安排

\n
vectors=np.random.rand(1000,1000,2)\n\n%timeit your_function(vectors)\n#no difference between fastmath=True or False, no SIMD-vectorization at all\n#23.3 ms \xc2\xb1 241 \xc2\xb5s per loop (mean \xc2\xb1 std. dev. of 7 runs, 10 loops each)\n\n%timeit optimized_function(vectors)\n#with fastmath=False #SIMD-vectorized, but with the slower (more accurate) SVML algorithm\n#9.03 ms \xc2\xb1 120 \xc2\xb5s per loop (mean \xc2\xb1 std. dev. of 7 runs, 100 loops each)\n#with fastmath=True  #SIMD-vectorized, but with the faster(less accurate) SVML algorithm\n#4.45 ms \xc2\xb1 14.5 \xc2\xb5s per loop (mean \xc2\xb1 std. dev. of 7 runs, 100 loops each)\n
Run Code Online (Sandbox Code Playgroud)\n