Rad*_*155 10 sql postgresql indexing
我有一个表,每年都会增加 10M 行。
该表有 10 列,称为 c1、c2、c3、...、c10。
我WHERE可能会在其中 8 个上使用该条款。
更具体地说:每次我查询表时,c10 列上总会有一个子句WHERE(它是一个日期,我可以搜索相等或范围)。
其他 7 个可能的可搜索列将不遵循任何模式。我可以搜索:
...以及所有其他可能的组合。
因此,在WHERE子句中,c10 将始终存在,而其他项可以以任意组合存在(甚至根本不存在)。
在这种情况下什么索引策略可以提高性能?我认为正确的做法是为每一列创建一个索引。使用多列索引可以提高性能吗?
据我所知,仅对于按此顺序使用 c1、c2、c3 或 c1、c2 或 c1 的查询,您将通过 (c1, c2, c3) 上的多列索引获得性能。但就像我说的,在我的场景中我唯一可以假设的是 c10 将始终出现在 WHERE 子句中(如果这有帮助的话,它也可以是第一个子句)
Edw*_*man 17
为了回答我们应该使用什么样的索引的问题,我们可以创建一个简单的测试。首先,我们创建数据库、表和索引。
CREATE DATABASE index_test;
CREATE TABLE single_column(a int, b int, c int);
CREATE TABLE multi_column(a int, b int, c int);
CREATE INDEX single_column_a_idx ON single_column (a);
CREATE INDEX single_column_b_idx ON single_column (b);
CREATE INDEX single_column_c_idx ON single_column (c);
CREATE INDEX multi_column_idx ON multi_column (a, b, c);
Run Code Online (Sandbox Code Playgroud)
用随机数据填充表。
-- this function will be used for random number generation
CREATE OR REPLACE FUNCTION random_in_range(INTEGER, INTEGER) RETURNS INTEGER AS $$
SELECT floor(($1 + ($2 - $1 + 1) * random()))::INTEGER;
$$ LANGUAGE SQL;
INSERT INTO single_column(a, b, c)
SELECT random_in_range(1, 100),
random_in_range(1, 100),
random_in_range(1, 100)
FROM generate_series(1, 1000000);
INSERT INTO multi_column(a, b, c)
SELECT random_in_range(1, 100),
random_in_range(1, 100),
random_in_range(1, 100)
FROM generate_series(1, 1000000);
Run Code Online (Sandbox Code Playgroud)
运行测试。
CREATE DATABASE index_test;
CREATE TABLE single_column(a int, b int, c int);
CREATE TABLE multi_column(a int, b int, c int);
CREATE INDEX single_column_a_idx ON single_column (a);
CREATE INDEX single_column_b_idx ON single_column (b);
CREATE INDEX single_column_c_idx ON single_column (c);
CREATE INDEX multi_column_idx ON multi_column (a, b, c);
Run Code Online (Sandbox Code Playgroud)
结果
-- this function will be used for random number generation
CREATE OR REPLACE FUNCTION random_in_range(INTEGER, INTEGER) RETURNS INTEGER AS $$
SELECT floor(($1 + ($2 - $1 + 1) * random()))::INTEGER;
$$ LANGUAGE SQL;
INSERT INTO single_column(a, b, c)
SELECT random_in_range(1, 100),
random_in_range(1, 100),
random_in_range(1, 100)
FROM generate_series(1, 1000000);
INSERT INTO multi_column(a, b, c)
SELECT random_in_range(1, 100),
random_in_range(1, 100),
random_in_range(1, 100)
FROM generate_series(1, 1000000);
Run Code Online (Sandbox Code Playgroud)
single_column表在任何情况下都将始终使用索引。EXPLAIN ANALYZE SELECT * FROM single_column WHERE a < 3; -- index used
EXPLAIN ANALYZE SELECT * FROM single_column WHERE b < 3; -- index used
EXPLAIN ANALYZE SELECT * FROM single_column WHERE c < 3; -- index used
EXPLAIN ANALYZE SELECT * FROM single_column WHERE a < 3 AND b > 10 AND c <= 11; -- index used
Run Code Online (Sandbox Code Playgroud)
multi_column,仅当查询中的列与索引定义中的第一列相同时才会使用索引。EXPLAIN ANALYZE SELECT * FROM multi_column WHERE a < 3; -- index used
EXPLAIN ANALYZE SELECT * FROM multi_column WHERE b < 3; -- index not used
EXPLAIN ANALYZE SELECT * FROM multi_column WHERE c < 3; -- index not used
Run Code Online (Sandbox Code Playgroud)
single_column表可以在多列 WHERE 上使用索引,但multi_column表速度更快。multi_column表可以在单列 WHERE 上使用索引,但single_column表速度更快。小智 0
我强烈建议采取以下策略:
c10. 由于它是日期,因此您可以按范围进行分区,进行年度或每月分区。我发现分区带来了巨大的性能提升,特别是在 、 和大型表中始终使用一列或多列的情况下WHERE。
| 归档时间: |
|
| 查看次数: |
9875 次 |
| 最近记录: |