Emi*_*ily 0 datetime r subsampling
我想从datetime列以小时为间隔对数据帧进行子采样,从数据帧第一行的时间值开始.我的数据框从第一行到最后一行每隔10分钟运行一次.示例数据如下:
structure(list(datetime = structure(1:19, .Label = c("30/03/2011 05:09",
"30/03/2011 05:19", "30/03/2011 05:29", "30/03/2011 05:39", "30/03/2011 05:49",
"30/03/2011 05:59", "30/03/2011 06:09", "30/03/2011 06:19", "30/03/2011 06:29",
"30/03/2011 06:39", "30/03/2011 06:49", "30/03/2011 06:59", "30/03/2011 07:09",
"30/03/2011 07:19", "30/03/2011 07:29", "30/03/2011 07:39", "30/03/2011 07:49",
"30/03/2011 07:59", "30/03/2011 08:09"), class = "factor"), a_count = c(66L,
34L, 33L, 20L, 12L, 44L, 36L, 29L, 21L, 22L, 17L, 38L, 24L, 19L,
60L, 54L, 27L, 36L, 45L), b_count = c(166.49, 167.54, 168.31,
168.81, 169.24, 169.61, 169.96, 170.29, 170.63, 170.98, 171.31,
171.62, 171.94, 172.29, 172.68, 173.15, 173.71, 174.34, 174.99
)), .Names = c("datetime", "a_count", "b_count"), class = "data.frame", row.names = c(NA,
-19L))
Run Code Online (Sandbox Code Playgroud)
DF
datetime a_count b_count
1 30/09/2011 05:09 66 166.49
2 30/09/2011 05:19 34 167.54
3 30/09/2011 05:29 33 168.31
4 30/09/2011 05:39 20 168.81
5 30/09/2011 05:49 12 169.24
6 30/09/2011 05:59 44 169.61
7 30/09/2011 06:09 36 169.96
8 30/09/2011 06:19 29 170.29
9 30/09/2011 06:29 21 170.63
10 30/09/2011 06:39 22 170.98
11 30/09/2011 06:49 17 171.31
12 30/09/2011 06:59 38 171.62
13 30/09/2011 07:09 24 171.94
14 30/09/2011 07:19 19 172.29
15 30/09/2011 07:29 60 172.68
16 30/09/2011 07:39 54 173.15
17 30/09/2011 07:49 27 173.71
18 30/09/2011 07:59 36 174.34
19 30/09/2011 08:09 45 174.99
Run Code Online (Sandbox Code Playgroud)
我想最终得到以下数据框:
datetime a_count b_count
30/09/2011 05:09 66 166.49
30/09/2011 06:09 36 169.96
30/09/2011 07:09 24 171.94
30/09/2011 08:09 45 174.99
Run Code Online (Sandbox Code Playgroud)
任何建议将不胜感激!
很难猜出你有什么结构.是否保证您在第一次正确值+ x乘60分钟时有一个值?如果找不到值,会发生什么?如果您当时有两个值,会发生什么.你需要近似匹配吗?说,09:10算作09:09?
让你入门的想法如下:
# I will call your dataframe `d`.
# Transform datetime to a POSIXct object, R's datatype for timestamps
d$datetime <- as.POSIXct(as.character(d$datetime), format='%d/%m/%Y %H:%M')
# Extract the minutes
d$minute <- as.numeric(format(d$datetime, '%M'))
# And select by identical minute.
subset(d, minute == d$minute[1])
Run Code Online (Sandbox Code Playgroud)