非常大数据集的余弦相似度

sgo*_*les 4 python numpy dataframe cosine-similarity

我在计算 100 维向量的大列表之间的余弦相似度时遇到问题。当我使用 时from sklearn.metrics.pairwise import cosine_similarity,我得到MemoryError了我的 16 GB 机器。每个数组都非常适合我的记忆,但我MemoryError在np.dot()内部通话中得到了

这是我的用例以及我目前如何解决它。

这是我的 100 维父向量,我需要将它与其他 500,000 个相同维度(即 100)的不同向量进行比较

parent_vector = [1, 2, 3, 4 ..., 100]
Run Code Online (Sandbox Code Playgroud)

这是我的子向量(在这个例子中有一些虚构的随机数)

child_vector_1 = [2, 3, 4, ....., 101]
child_vector_2 = [3, 4, 5, ....., 102]
child_vector_3 = [4, 5, 6, ....., 103]
.......
.......
child_vector_500000 = [3, 4, 5, ....., 103]
Run Code Online (Sandbox Code Playgroud)

我的最终目标是获得child_vector_1与父向量具有非常高的余弦相似度的前 N 个子向量(具有它们的名称和相应的余弦分数)。

我目前的方法(我知道这是低效且消耗内存的):

第 1 步:创建以下形状的超级数据框

parent_vector         1,    2,    3, .....,    100   
child_vector_1        2,    3,    4, .....,    101   
child_vector_2        3,    4,    5, .....,    102   
child_vector_3        4,    5,    6, .....,    103   
......................................   
child_vector_500000   3,    4,    5, .....,    103
Run Code Online (Sandbox Code Playgroud)

第 2 步:使用

from sklearn.metrics.pairwise import cosine_similarity
cosine_similarity(df)
Run Code Online (Sandbox Code Playgroud)

获得所有向量之间的成对余弦相似度(如上图所示)

第 3 步:制作一个元组列表来存储所有此类组合的key此类child_vector_1和值(如余弦相似度数)。

第 4 步:使用sort()列表获取前 N 个——这样我就可以得到子向量名称以及它与父向量的余弦相似度分数。

PS:我知道这是非常低效的,但我想不出更好的方法来更快地计算每个子向量和父向量之间的余弦相似度并获得前 N 个值。

任何帮助将不胜感激。

小智 6

即使您的 (500000, 100) 数组(父级及其子级)适合内存,但其上的任何成对度量都不会。原因是,顾名思义,成对度量计算任何两个孩子的距离。为了存储这些距离,您需要一个 (500000,500000) 大小的浮点数组,如果我的计算是正确的,它将需要大约 100 GB 的内存。

幸运的是,您的问题有一个简单的解决方案。如果我理解正确,您只想拥有孩子和父母之间的距离,这将导致长度为 500000 的向量很容易存储在内存中。

为此,您只需要为仅包含 parent_vector 的 cosine_similarity 提供第二个参数

import pandas as pd
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity

df = pd.DataFrame(np.random.rand(500000,100)) 
df['distances'] = cosine_similarity(df, df.iloc[0:1]) # Here I assume that the parent vector is stored as the first row in the dataframe, but you could also store it separately

n = 10 # or however many you want
n_largest = df['distances'].nlargest(n + 1) # this contains the parent itself as the most similar entry, hence n+1 to get n children
Run Code Online (Sandbox Code Playgroud)

希望能解决你的问题。

  • 即使我面临同样的问题,我的数据帧的大小为`(32593, 12)` 我需要计算所有对的余弦相似度,即 32593*32593 并且它不适合内存。我该如何处理这种情况? (2认同)