在多维立方体上的Postgresql k-最近邻(KNN)

MD *_*ffy 6 postgresql nearest-neighbor postgresql-9.1

我有一个有8个维度的立方体.我想做最近邻居匹配.我对postgresql完全不熟悉.我读到9.1支持多维上的最近邻匹配.如果有人能给出一个完整的例子,我真的很感激:

  1. 如何使用8D立方体创建表?

  2. 样本插入

  3. 查找 - 完全匹配

  4. 查找 - 最近邻居匹配

样本数据:

为简单起见,我们可以假设所有值的范围都是0-100.

第1点:(1,1,1,1,1,1,1,1)

第2点:(2,2,2,2,2,2,2,2)

查找值:(1,1,1,1,1,1,1,2)

这应该与Point1匹配,而不是Point2.

参考文献:

What's_new_in_PostgreSQL_9.1

https://en.wikipedia.org/wiki/K-d_tree#Nearest_neighbour_search

Tom*_*eif 6

PostgreSQL支持距离运算符<->,据我所知,这可以用于分析文本(使用pg_trgrm模块)和几何数据类型.

我不知道如何使用它超过1维.也许您必须定义自己的距离函数或以某种方式将数据转换为具有文本或几何类型的一列.例如,如果您有8列(8维立方体)的表:

c1 c2 c3 c4 c5 c6 c7 c8
 1  0  1  0  1  0  1  2
Run Code Online (Sandbox Code Playgroud)

你可以将它转换为:

c1 c2 c3 c4 c5 c6 c7 c8
 a  b  a  b  a  b  a  c
Run Code Online (Sandbox Code Playgroud)

然后用一列表格:

c1
abababac
Run Code Online (Sandbox Code Playgroud)

然后你可以使用(创建gist 索引后):

SELECT c1, c1 <-> 'ababab'
 FROM test_trgm 
 ORDER BY c1 <-> 'ababab';
Run Code Online (Sandbox Code Playgroud)

例

创建样本数据

-- Create some temporary data
-- ! Note that table are created in tmp schema (change sql to your scheme) and deleted if exists !
drop table if exists tmp.test_data;

-- Random integer matrix 100*8 
create table tmp.test_data as (
   select 
      trunc(random()*100)::int as input_variable_1,
      trunc(random()*100)::int as input_variable_2, 
      trunc(random()*100)::int as input_variable_3,
      trunc(random()*100)::int as input_variable_4, 
      trunc(random()*100)::int as input_variable_5, 
      trunc(random()*100)::int as input_variable_6, 
      trunc(random()*100)::int as input_variable_7, 
      trunc(random()*100)::int as input_variable_8
   from 
      generate_series(1,100,1)
);
Run Code Online (Sandbox Code Playgroud)

将输入数据转换为文本

drop table if exists tmp.test_data_trans;

create table tmp.test_data_trans as (
select 
   input_variable_1 || ';' ||
   input_variable_2 || ';' ||
   input_variable_3 || ';' ||
   input_variable_4 || ';' ||
   input_variable_5 || ';' ||
   input_variable_6 || ';' ||
   input_variable_7 || ';' ||
   input_variable_8 as trans_variable
from 
   tmp.test_data
);
Run Code Online (Sandbox Code Playgroud)

这将为您提供一个trans_variable存储所有8个维度的变量:

trans_variable
40;88;68;29;19;54;40;90
80;49;56;57;42;36;50;68
29;13;63;33;0;18;52;77
44;68;18;81;28;24;20;89
80;62;20;49;4;87;54;18
35;37;32;25;8;13;42;54
8;58;3;42;37;1;41;49
70;1;28;18;47;78;8;17
Run Code Online (Sandbox Code Playgroud)

||您也可以使用以下语法代替运算符(更短,但更神秘):

select 
   array_to_string(string_to_array(t.*::text,''),'') as trans_variable
from 
   tmp.test_data t
Run Code Online (Sandbox Code Playgroud)

添加索引

create index test_data_gist_index on tmp.test_data_trans using gist(trans_variable);
Run Code Online (Sandbox Code Playgroud)

测试距离注意:我从表中选择了一行52;42;18;50;68;29;8;55- 并使用稍微改变的值(42;42;18;52;98;29;8;55)来测试距离.当然,测试数据中的值将完全不同,因为它是RANDOM矩阵.

select 
   *, 
   trans_variable <->  '42;42;18;52;98;29;8;55' as distance,
   similarity(trans_variable, '42;42;18;52;98;29;8;55') as similarity,
from 
   tmp.test_data_trans 
order by
   trans_variable <-> '52;42;18;50;68;29;8;55';
Run Code Online (Sandbox Code Playgroud)

您可以使用距离算子< - >或类似功能.距离= 1 - 相似度


goj*_*omo 5

A" 补丁引入了k近邻搜索欧几里得,出租车和切比雪夫距离立方体 "上最近提出的pgsql-黑客名单.如果您可以自定义PostgreSQL构建,它可能适合您的目的.

请注意,cube类型(PostgreSQL扩展)可用于表示n维中的点或立方体.(默认情况下,n的值最多可以达到100,如果cubedata.h提高限制,则更多.)因此,此补丁应该启用索引辅助的多维点/向量/立方体最近邻搜索.

(如果没有此补丁,该cube类型没有<->距离运算符,并且缺少支持函数(#8),OPERATOR CLASS gist_cube_ops这使得gist能够对这些值建立与距离相关的索引.)

我还没有尝试过这个补丁,并注意到其中一个讨论列表的回复表明它可能会破坏一些回归测试.