doo*_*oma 5 python arrays numpy vectorization
我有两个 2D numpy 数组,例如:
A = numpy.array([[1, 2, 4, 8], [16, 32, 32, 8], [64, 32, 16, 8]])
和
B = numpy.array([[1, 2], [32, 32]])
我想要拥有A可以从 的任何行中找到所有元素的所有行B。如果 的一行中有 2 个相同的元素B,则 from 的行也A必须至少包含 2 个。就我的例子而言,我想实现这一目标:
A_filtered = [[1, 2, 4, 8], [16, 32, 32, 8]]
我可以控制值表示,因此我选择了二进制表示只有一个位置的数字1(例如:0b00000001和0b00000010等...)这样我可以使用函数轻松检查所有类型的值是否都在行中np.logical_or.reduce(),但是我无法检查一行中相同元素的数量是否更大或相等A。我真的希望我可以避免简单的for循环和数组的深层复制,因为性能对我来说是一个非常重要的方面。
我怎样才能在 numpy 中以有效的方式做到这一点?
更新:
这里的解决方案可能有效,但我认为性能对我来说是一个大问题,它A可以很大(> 300000 行),也B可以中等(> 30 行):
[set(row).issuperset(hand) for row in A.tolist() for hand in B.tolist()]
更新2:
该set()解决方案不起作用,因为set()删除了所有重复的值。
我希望我答对了你的问题。至少它可以解决您在问题中描述的问题。如果输出的顺序应与输入保持相同,请更改就地排序。
该代码看起来很丑陋,但应该表现良好并且不应该很难理解。
代码
import time
import numba as nb
import numpy as np
@nb.njit(fastmath=True,parallel=True)
def filter(A,B):
iFilter=np.zeros(A.shape[0],dtype=nb.bool_)
for i in nb.prange(A.shape[0]):
break_loop=False
for j in range(B.shape[0]):
ind_to_B=0
for k in range(A.shape[1]):
if A[i,k]==B[j,ind_to_B]:
ind_to_B+=1
if ind_to_B==B.shape[1]:
iFilter[i]=True
break_loop=True
break
if break_loop==True:
break
return A[iFilter,:]
Run Code Online (Sandbox Code Playgroud)
测量性能
####First call has some compilation overhead####
A=np.random.randint(low=0, high=60, size=300_000*4).reshape(300_000,4)
B=np.random.randint(low=0, high=60, size=30*2).reshape(30,2)
t1=time.time()
#At first sort the arrays
A.sort()
B.sort()
A_filtered=filter(A,B)
print(time.time()-t1)
####Let's measure the second call too####
A=np.random.randint(low=0, high=60, size=300_000*4).reshape(300_000,4)
B=np.random.randint(low=0, high=60, size=30*2).reshape(30,2)
t1=time.time()
#At first sort the arrays
A.sort()
B.sort()
A_filtered=filter(A,B)
print(time.time()-t1)
Run Code Online (Sandbox Code Playgroud)
结果
46ms after the first run on a dual-core Notebook (sorting included)
32ms (sorting excluded)
Run Code Online (Sandbox Code Playgroud)
| 归档时间: |
|
| 查看次数: |
598 次 |
| 最近记录: |