如何向量化嵌套循环

Question

如何向量化嵌套循环

erb*_*dge 9 matlab loops for-loop vectorization octave

我无法想象如何对这组循环进行矢量化.任何指导将不胜感激.

ind_1 = [1,2,3];
ind_2 = [1,2,4];
K = zeros(3,3,3,3,3,3,3,3,3);
pp = rand(4,4,4);

for s = 1:3
 for t = 1:3
  for k = 1:3
   for l = 1:3
    for m = 1:3
     for n = 1:3
      for o = 1:3
       for p = 1:3
        for r = 1:3
         % the following loops are singular valued except when
         % y=3 for ind_x(y) in this case
         for a_s = ind_1(s):ind_2(s)
          for a_t = ind_1(t):ind_2(t)
           for a_k = ind_1(k):ind_2(k)
            for a_l = ind_1(l):ind_2(l)
             for a_m = ind_1(m):ind_2(m)
              for a_n = ind_1(n):ind_2(n)
               for a_o = ind_1(o):ind_2(o)
                for a_p = ind_1(p):ind_2(p)
                 for a_r = ind_1(r):ind_2(r)
                  K(s,t,k,l,m,n,o,p,r) = K(s,t,k,l,m,n,o,p,r) + ...
                    pp(a_s, a_t, a_r) * pp(a_k, a_l, a_r) * ...
                    pp(a_n, a_m, a_s) * pp(a_o, a_n, a_t) * ...
                    pp(a_p, a_o, a_k) * pp(a_m, a_p, a_l);
                 end
                end
               end
              end
             end
            end
           end
          end
         end
        end
       end
      end
     end
    end
   end
  end
 end
end

Run Code Online (Sandbox Code Playgroud)

编辑:

该代码是由产品的值求和产生秩9张量索引为1至3个pp■一个或两次,每次索引,这取决于的值ind_1和ind_2.

编辑:

这是一个3d示例,但请记住,pp在9d版本中不保留简单置换索引的事实:

ind_1 = [1,2,3];
ind_2 = [1,2,4];
K = zeros(3,3,3);
pp = rand(4,4,4);

for s = 1:3
 for t = 1:3
  for k = 1:3
   % the following loops are singular valued except when
   % y=3 for ind_x(y) in this case
   for a_s = ind_1(s):ind_2(s)
    for a_t = ind_1(t):ind_2(t)
     for a_k = ind_1(k):ind_2(k)
      K(s,t,k) = K(s,t,k) + ...
        pp(a_s, a_t, a_r) * pp(a_t, a_s, a_k) * ...
        pp(a_k, a_t, a_s) * pp(a_k, a_s, a_t);
     end
    end
   end
  end
 end
end

Run Code Online (Sandbox Code Playgroud)

Answer 1

Yve*_*reY 5

哇!非常简单的解决方案,但不容易找到.顺便说一句,我想知道你的公式来自哪里.

如果你不介意暂时失去一点内存(2次4 ^ 9阵列比之前的3 ^ 9),你可以推迟在第3和第4个超平面的累积.

与八度3.2.4在UNIX盒测试,它下降23S(67MB)至0.17s(98MB) .

function K = tensor9_opt(pp)

  ppp = repmat(pp, [1 1 1 4 4 4 4 4 4]) ;
  % The 3 first numbers are variable indices (eg 1 for a_s to 9 for a_r)
  % Other numbers must complete 1:9 indices in any order
  T = ipermute(ppp, [1 2 9 3 4 5 6 7 8]) .* ...
      ipermute(ppp, [3 4 9 1 2 5 6 7 8]) .* ...
      ipermute(ppp, [6 5 1 2 3 4 7 8 9]) .* ...
      ipermute(ppp, [7 6 2 1 3 4 5 8 9]) .* ...
      ipermute(ppp, [8 7 3 1 2 4 5 6 9]) .* ...
      ipermute(ppp, [5 8 4 1 2 3 6 7 9]) ;

  % I have not found how to manipulate 'multi-ranges' programmatically. 
  T1 = T (:,:,:,:,:,:,:,:,1:end-1) ; T1(:,:,:,:,:,:,:,:,end) += T (:,:,:,:,:,:,:,:,end) ;
  T  = T1(:,:,:,:,:,:,:,1:end-1,:) ; T (:,:,:,:,:,:,:,end,:) += T1(:,:,:,:,:,:,:,end,:) ;
  T1 = T (:,:,:,:,:,:,1:end-1,:,:) ; T1(:,:,:,:,:,:,end,:,:) += T (:,:,:,:,:,:,end,:,:) ;
  T  = T1(:,:,:,:,:,1:end-1,:,:,:) ; T (:,:,:,:,:,end,:,:,:) += T1(:,:,:,:,:,end,:,:,:) ;
  T1 = T (:,:,:,:,1:end-1,:,:,:,:) ; T1(:,:,:,:,end,:,:,:,:) += T (:,:,:,:,end,:,:,:,:) ;
  T  = T1(:,:,:,1:end-1,:,:,:,:,:) ; T (:,:,:,end,:,:,:,:,:) += T1(:,:,:,end,:,:,:,:,:) ;
  T1 = T (:,:,1:end-1,:,:,:,:,:,:) ; T1(:,:,end,:,:,:,:,:,:) += T (:,:,end,:,:,:,:,:,:) ;
  T  = T1(:,1:end-1,:,:,:,:,:,:,:) ; T (:,end,:,:,:,:,:,:,:) += T1(:,end,:,:,:,:,:,:,:) ;
  K  = T (1:end-1,:,:,:,:,:,:,:,:) ; K (end,:,:,:,:,:,:,:,:) += T (end,:,:,:,:,:,:,:,:) ;
endfunction

pp = rand(4,4,4);
K = tensor9_opt(pp) ;

Run Code Online (Sandbox Code Playgroud)

归档时间：	13 年，3 月前
查看次数：	960 次
最近记录：	10 年，6 月前