Tar*_*rik 5 python performance julia
我刚刚翻译了一组涉及矩阵的科学计算,其中的元素是符号表达式,它们被区分并与各种其他数学表达式组合,然后进行数值积分。为了重现我遇到的性能差距,下面的代码片段构成了一个最小的示例。我知道先用符号区分然后用数字积分是没有意义的,但同样,重点是性能差距。需要注意的是,导入库并不代表太多时间,也不能解释性能差距。
朱莉娅代码:
using Symbolics, QuadGK
@variables x
m = [i * 10*x^3 + 1/i * sin(x) + 5*i*x^3 * cos(x) - 8i*x^2 + 2/sin(i*3.0)*x + exp(1/(x+10)) for i in 1:500]
m_d = expand_derivatives.(Differential(x).(m))
m_d_expr = build_function(m, x)
m_d_f = eval(m_d_expr[1])
v = quadgk(m_d_f, 0, 1)
print(v[1])
Run Code Online (Sandbox Code Playgroud)
蟒蛇代码:
import pandas as pd
import numpy as np
from sympy import sin, diff, pi, lambdify, integrate, cos, exp
from sympy.abc import x
from sympy.matrices import Matrix
from scipy.integrate import quad
def integrate_matrix(m, v, a, b):
mi = np.zeros((m.rows, m.cols))
for i in range(m.rows):
for j in range(m.cols):
f = lambdify(v, m[i, j])
integral_value = quad(f, a, b)[0]
mi[i, j] = integral_value
return mi
m = Matrix([i * 10*x**3 + 1/i * sin(x) + 5*i*x**3 * cos(x) - 8*i*x**2 + 2/sin(i*3.0)*x + exp(1/(x+10)) for i in range(1, 501)])
v = integrate_matrix(m, x, 0, 1)
print(v)
Run Code Online (Sandbox Code Playgroud)
我的问题:有没有办法改进 Julia 代码以匹配 Python 代码性能。每次我试图让我的码头对 Julia 的表现留下深刻印象时,我都会感到尴尬。我仍然是一个 Julia noob,但我真的不知道该怎么做。
大约时间:Python:6 秒 Julia:30+ 秒
朱莉娅 1.6 版 Python 3.7
注意:由于差距很大,我发布了这个。不,CAS 并没有解释这一切。此外,我们正在做一个符号差异而不是集成,更不用说 sympy 是缓慢的。我可以将代码添加到精确的时间,然后呢?我遇到问题的原始科学代码是 6 秒 Python vs 75 秒 Julia。多可惜。
更快地运行整个事情是任何理智的人关心的。
据我了解,Julia 关心的是将运行速度提高几倍,而只运行一次总是比较慢,因为 Julia 代码需要在执行前进行编译。与 Julia 不同,Python 不进行任何 JIT 编译,并且随时准备以相同的速度运行。
因此,我将您的 Julia 代码粘贴到其中code.jl并在同一会话中多次运行:
# New Julia session!
julia> @time include("code.jl")
[long array...]
24.660636 seconds (42.99 M allocations: 2.607 GiB, 3.82% gc time, 0.02% compilation time)
julia> @time include("code.jl")
[long array...]
2.761062 seconds (5.61 M allocations: 240.159 MiB, 10.39% gc time, 57.06% compilation time)
julia> @time include("code.jl")
[long array...]
2.608917 seconds (5.61 M allocations: 240.164 MiB, 4.47% gc time, 61.75% compilation time)
# Restarted Julia
julia> @time include("code.jl")
25.538249 seconds (42.99 M allocations: 2.607 GiB, 3.76% gc time, 0.02% compilation time)
julia> @time include("code.jl")
2.740550 seconds (5.61 M allocations: 240.159 MiB, 9.94% gc time, 56.72% compilation time)
Run Code Online (Sandbox Code Playgroud)
因此,第一次运行代码大约需要 25 秒,再次运行大约需要 3 秒(!),即使这 3 秒中有 50% 用于编译内容。但是,最初的 25 秒中只有 0.02% 用于编译。显然,放缓不是由于编译时间?还要注意它在第一次运行时执行了多少内存分配:4300 万 vs 下一次运行的大约 550 万(少7倍!)。但无论如何,第一次运行真的很慢,而随后的运行却是闪电般的快。
julia> @time using Symbolics
3.503349 seconds (6.42 M allocations: 460.519 MiB, 3.53% gc time, 0.13% compilation time)
julia> @time using Symbolics
0.000241 seconds (136 allocations: 9.641 KiB)
0.000280 seconds (136 allocations: 9.641 KiB)
0.000249 seconds (136 allocations: 9.641 KiB)
0.000251 seconds (136 allocations: 9.641 KiB)
0.000252 seconds (136 allocations: 9.641 KiB)
0.000246 seconds (136 allocations: 9.641 KiB)
# I didn't import it before,
# but apparently `Symbolics` did
julia> @time using QuadGK
0.000276 seconds (137 allocations: 9.688 KiB)
0.000276 seconds (136 allocations: 9.641 KiB)
0.000240 seconds (136 allocations: 9.641 KiB)
0.000251 seconds (136 allocations: 9.641 KiB)
Run Code Online (Sandbox Code Playgroud)
也就是说,仅运行带有导入的代码的第一行就花费了 3.5 秒。我想,由于缓存,后续导入显然更快。
julia> @time m = [i * 10*x^3 + 1/i * sin(x) + 5*i*x^3 * cos(x) - 8i*x^2 + 2/sin(i*3.0)*x + exp(1/(x+10)) for i in 1:500];
2.590259 seconds (4.69 M allocations: 284.672 MiB, 10.86% gc time, 98.69% compilation time)
julia> @time m = [i * 10*x^3 + 1/i * sin(x) + 5*i*x^3 * cos(x) - 8i*x^2 + 2/sin(i*3.0)*x + exp(1/(x+10)) for i in 1:500];
0.102573 seconds (231.21 k allocations: 12.507 MiB, 72.61% compilation time)
0.098871 seconds (231.21 k allocations: 12.508 MiB, 72.39% compilation time)
0.108458 seconds (231.21 k allocations: 12.512 MiB, 7.93% gc time, 67.73% compilation time)
0.099787 seconds (231.22 k allocations: 12.508 MiB, 72.99% compilation time)
0.098378 seconds (231.21 k allocations: 12.507 MiB, 73.80% compilation time)
Run Code Online (Sandbox Code Playgroud)
同样,启动缓慢(98.69% 编译时间),但接下来的运行速度要快得多。
# New Julia session!
julia> @time include("code.jl")
[long array...]
24.660636 seconds (42.99 M allocations: 2.607 GiB, 3.82% gc time, 0.02% compilation time)
julia> @time include("code.jl")
[long array...]
2.761062 seconds (5.61 M allocations: 240.159 MiB, 10.39% gc time, 57.06% compilation time)
julia> @time include("code.jl")
[long array...]
2.608917 seconds (5.61 M allocations: 240.164 MiB, 4.47% gc time, 61.75% compilation time)
# Restarted Julia
julia> @time include("code.jl")
25.538249 seconds (42.99 M allocations: 2.607 GiB, 3.76% gc time, 0.02% compilation time)
julia> @time include("code.jl")
2.740550 seconds (5.61 M allocations: 240.159 MiB, 9.94% gc time, 56.72% compilation time)
Run Code Online (Sandbox Code Playgroud)
因此,Python 代码始终运行约6 秒。
这比后续运行 Julia 代码慢2 倍!但是,一旦启动 Python 解释器,您就会获得这种速度,而 Julia 将花时间编译代码并做……其他需要 4300 万内存分配的事情。但是 Julia 用糟糕的启动时间换取的是编译代码的性能(在这个例子中,Julia比 Python 快 2 倍)。
include("your_code.jl")在修改代码之后。这可能会导致奇怪的错误,因为环境将由以前运行的数据填充。| 归档时间: |
|
| 查看次数: |
145 次 |
| 最近记录: |