OpenCV GPU库对矩阵运算有多好?

Ale*_*xey 7 c++ opencv cuda gpu thrust

我正在将OpenCV用于计算机视觉应用.我想在GPU上加速一些矩阵运算(矩阵相当大),并且如果可能的话,希望避免直接在CUDA C中进行编码.OpenCV 2.4.1具有许多GPU加速功能.他们的体验表现如何?我最好还是使用另一个库(例如Thrust)吗?

EDIT 示例应用:计算GPU上的平方欧几里德距离矩阵.目前,我在Matlab中使用并行计算工具箱(PCT)进行的GPU加速(和矢量化)实现比使用OpenCV的C++实现快5到10倍.

Matlab实现:

function K = sqEuclideanDist(P_cpu,Q_cpu)
% Vectorized method to compute pairwise squared Euclidean distance on GPU
% Returns K(i,j) = (P(i,:) - Q(j,:))'*(P(i,:) - Q(j,:))

P_gpu = gpuArray(P_cpu);
Q_gpu = gpuArray(Q_cpu);

[nP, d] = size(P_gpu);
[nQ, d] = size(Q_gpu);

pmag = sum(P_gpu .* P_gpu, 2);
qmag = sum(Q_gpu .* Q_gpu, 2);

% note that K is on GPU
K = ones(nP,1)*qmag' + pmag*ones(1,nQ) - 2*P_gpu*Q_gpu';

end
Run Code Online (Sandbox Code Playgroud)

更新这是另一个完成相同的Matlab实现(感谢/sf/answers/544202641/).但它仅在CPU上运行,因为bsxfunPCT不支持.仍然在寻找C++替代品.

function K = sqEuclideanDist(P_cpu,Q_cpu)
% Returns K(i,j) = (P(i,:) - Q(j,:))'*(P(i,:) - Q(j,:))
% Runs on CPU only.

K = bsxfun(@plus,sum(p.^2,2),sum(q.^2,2)') - 2*(p*q');

end
Run Code Online (Sandbox Code Playgroud)

小智 5

我发现ArrayFire速度更快,并且已经开始使用它代替 OpenCV 中的 GPU 内核进行图像处理。以下是我发现的一些基准测试,将 ArrayFire(曾经在一个名为 LibJacket 的不同接口中)与 OpenCV 进行比较,并且在我的基准测试中也是如此,ArrayFire 比 OpenCV 中的 GPU 功能快 2-4 倍。据我所知,NVIDIA 并没有在 OpenCV 中编写 GPU 内核,而是将其外包给某人,这可能就是它们如此缓慢的原因。因为我只使用 1 个 GPU,所以我可以免费使用 ArrayFire。

更新,考虑到@Alex 发布的新 MATLAB 代码: 我在我的系统上运行了此代码的基准测试。我知道并行计算工具箱 gpuArray 比 CPU 慢,但 Jacket 和 ArrayFire 踢屁股。硬件规格是:

Intel(R) Xeon(R) CPU X5660  @ 2.80GHz
NVIDIA Tesla M2090
Run Code Online (Sandbox Code Playgroud)

使用 Parallel Computing Toolbox gpuArray(完全预热)的 CPU 与 GPU 的结果。 CPU 比 PCT gpuArray 快:

>> tic; sqEuclideanDist(gpuArray(rand(1581,3)),gpuArray(rand(189,3))); toc;
Elapsed time is 0.006859 seconds.
>> tic; sqEuclideanDist(rand(1581,3),rand(189,3)); toc;
Elapsed time is 0.005712 seconds.
Run Code Online (Sandbox Code Playgroud)

使用 Jacket 的 CPU 与 GPU 的结果(完全预热)。 Jacket 比 PCT gpuArray 快 3.7 倍,比 CPU 快 3 倍

>> tic; sqEuclideanDist(gdouble(rand(1581,3)),gdouble(rand(189,3))); toc;
Elapsed time is 0.001876 seconds.
Run Code Online (Sandbox Code Playgroud)

这是修改后的代码,让您可以轻松运行所有内容:

function K = sqEuclideanDist(P,Q)
% Vectorized method to compute pairwise squared Euclidean distance on GPU
% Returns K(i,j) = (P(i,:) - Q(j,:))'*(P(i,:) - Q(j,:))

[nP, d] = size(P);
[nQ, d] = size(Q);

pmag = sum(P .* P, 2);
qmag = sum(Q .* Q, 2);

K = ones(nP,1)*qmag' + pmag*ones(1,nQ) - 2*P*Q';

end
Run Code Online (Sandbox Code Playgroud)

Jacket 确实在 GPU 上支持 BSXFUN,并且确实在一定程度上提高了速度:

>> tic; sqEuclideanDist(gdouble(rand(1581,3)),gdouble(rand(189,3))); toc;
Elapsed time is 0.001420 seconds.
Run Code Online (Sandbox Code Playgroud)

请注意,此处使用的尺寸非常小,因此尝试在这些小尺寸上运行的大多数 CUDA 代码可能性能不佳。这就是我喜欢使用 AccelerEyes 的东西的原因,因为这些人已经优化了 GPU,这与 PCT gpuArray、Thrust、OpenCV 不同,我过去曾尝试过这些。

这是 ArrayFire Free C++ 结果:

Time:  0.0003577 seconds
Speedups:  19.2X faster than PCT gpuArray, 16X faster than the CPU, 5.2X faster
than Jacket in MATLAB original version, 4X faster than Jacket in MATLAB using
BSXFUN
Run Code Online (Sandbox Code Playgroud)

这是我为此编写的 ArrayFire 代码:

static array SqEuclideanDist(array P, array Q)
{
    // 0 based indexing
    array pmag = sum(P * P, 1);
    array qmag = sum(Q * Q, 1);

    int np = P.dims(0);
    int nq = Q.dims(0);

    array K = tile(qmag.T(), np, 1) + tile(pmag, 1, nq) - 2 * matmul(P, Q.T());
    return K;
}

int main(int argc, char **argv)
{
    double *P_cpu = new double[1581 * 3];
    double *Q_cpu = new double[189 * 3];

    array P = array(1581, 3, P_cpu);
    array Q = array(189 , 3, Q_cpu);
    af::sync();

    int iter = 1000;

    timer::tic();
    for (int i = 0; i < iter; i++) {
        array K = SqEuclideanDist(P, Q);
        af::eval(K);
    }

    af::sync();
    printf("Time taken: %2.4lfms\n", (1000 * timer::toc()) / iter);

    delete[] P_cpu;
    delete[] Q_cpu;
}
Run Code Online (Sandbox Code Playgroud)