使用SQL查询识别趋势

Dan*_*sin 8 sql trend

我有一个表(让我们称之为数据)与一组对象ID,数值和日期.我想确定其值在过去X分钟(例如,一小时)内具有正趋势的对象.

示例数据:

entity_id | value | date

1234      | 15    | 2014-01-02 11:30:00

5689      | 21    | 2014-01-02 11:31:00

1234      | 16    | 2014-01-02 11:31:00
Run Code Online (Sandbox Code Playgroud)

我试着看类似的问题,但没有找到任何帮助不幸...

Joh*_*tom 30

您启发我在SQL Server中实现线性回归.这可以针对MySQL/Oracle/Whatever进行修改而不会有太多麻烦.这是确定每个entity_id的小时趋势的数学上最好的方法,它将只选择具有正趋势的趋势.

它实现了计算这里列出的B1的公式:https://en.wikipedia.org/wiki/Regression_analysis#Linear_regression

create table #temp
(
    entity_id int,
    value int,
    [date] datetime
)

insert into #temp (entity_id, value, [date])
values
(1,10,'20140102 07:00:00 AM'),
(1,20,'20140102 07:15:00 AM'),
(1,30,'20140102 07:30:00 AM'),
(2,50,'20140102 07:00:00 AM'),
(2,20,'20140102 07:47:00 AM'),
(3,40,'20140102 07:00:00 AM'),
(3,40,'20140102 07:52:00 AM')

select entity_id, 1.0*sum((x-xbar)*(y-ybar))/sum((x-xbar)*(x-xbar)) as Beta
from
(
    select entity_id,
        avg(value) over(partition by entity_id) as ybar,
        value as y,
        avg(datediff(second,'20140102 07:00:00 AM',[date])) over(partition by entity_id) as xbar,
        datediff(second,'20140102 07:00:00 AM',[date]) as x
    from #temp
    where [date]>='20140102 07:00:00 AM' and [date]<'20140102 08:00:00 AM'
) as Calcs
group by entity_id
having 1.0*sum((x-xbar)*(y-ybar))/sum((x-xbar)*(x-xbar))>0
Run Code Online (Sandbox Code Playgroud)

  • 您可能已经猜到了,这还允许您对结果集中的 Beta 列进行排序,以找到温度升高趋势最强的实体。如果你喜欢那种东西。 (2认同)