考虑以下最小工作示例:
library(magrittr) # for the %>% pipe
library(data.table)
# test data.table contains common_column and two others
test_dt <- data.table(test_column_one = c(1, 2, 3), test_column_two = c("x","y","z"), common_column = c("ID1", "ID2", "ID3") )
# some other data.table that contains common_column
other_dt <- data.table( additional_info = c("US", "US", "GB"), common_column = c("ID1", "ID2", "ID3"))
example_function <- function(dt_column){
# does some things on the data tables based on the column parameter passed
merged_dt <- merge(other_dt, test_dt[,.(common_column, dt_column)], by = "common_column")[order(dt_column),] # order by the dt_column
return(merged_dt)
}
# calling the example function
example_function(test_dt$test_column_one)
Run Code Online (Sandbox Code Playgroud)
我怎样才能将代码修改为:
我想避免 for 循环并尽可能利用优化的 data.table 语法。
我尝试使用data.tableunlist()特定..的语法,但不知何故我总是收到奇怪的错误消息,并且不确定如何继续。
env 用于 data.table 编程的新接口解决了所有这些问题,并且不需要其他方法和 hack。
使用 SamR 的很好的例子,从 mt_cars(在 j 中)中选择“mpg” ,同时也在(在 i 中)some_col进行排序:some_col
setDT(mtcars)
some_col <- "wt"
mtcars[order(v), .(mpg, v), env=list(v=some_col)]
Run Code Online (Sandbox Code Playgroud)
请参阅此处的新闻第 10 项和此小插图。要更新到开发版本 1.14.99:data.table::update_dev_pkg().
我在下面列出了三种方法来做到这一点,但我认为最清楚的是创建一个与语法一起使用的列向量..。
example_function <- function(dt_column, dt1 = other_dt, dt2 = test_dt) {
cols_to_merge <- c("common_column", dt_column)
merged_dt <- merge(
dt1,
dt2[, ..cols_to_merge],
by = "common_column"
)
# No need to pipe - see explanation below
setorderv(merged_dt, dt_column)
return(merged_dt)
}
example_function("test_column_one")
# common_column additional_info test_column_one
# <char> <char> <num>
# 1: ID1 US 1
# 2: ID2 US 2
# 3: ID3 GB 3
Run Code Online (Sandbox Code Playgroud)
假设我们要从mpg中选择一个动态列mtcars,在本例中恰好是wt。
mtcars <- as.data.table(mtcars)
col_of_interest <- "wt"
Run Code Online (Sandbox Code Playgroud)
我认为基本上有三种data.table方法。
When
j是以 it 为前缀的符号,..将在调用范围中查找,其值取为列名或数字。
cols <- c("mpg", col_of_interest)
mtcars[, ..cols]
Run Code Online (Sandbox Code Playgroud)
这个 [the
..] 前缀现在扩展到出现在的所有符号j=
mtcars[, c("mpg", ..col_of_interest)]
Run Code Online (Sandbox Code Playgroud)
cols <- c("mpg", col_of_interest)
mtcars[, .SD, .SDcols = cols]
Run Code Online (Sandbox Code Playgroud)
我个人认为创建列名向量是最清晰的方法,但第二种方法意味着创建的变量少了一个。第三种方法是向后兼容的。
order()文档指定..仅适用于j. 以下语句是等效的:
mtcars[order(mpg)]
mtcars[order(mpg),]
mtcars[i = order(mpg),]
Run Code Online (Sandbox Code Playgroud)
正如order(mpg)提供的参数一样i,如果我们设置了x <- "mpg",那么这样做是无效的mtcars[order(..x)]。
如果您必须使用order(),可以通过以下一些方法来实现:
# Use ..
mtcars[order(mtcars[, ..x])]
# Use `[[`
mtcars[order(mtcars[[x]])]
# Use .SD
mtcars[mtcars[, sapply(.SD, order), .SDcols = x]]
# Use get() (this has been retired in favour of `env` - see comments)
mtcars[order(get(x))]
Run Code Online (Sandbox Code Playgroud)
然而,这些创建了数据的副本,它们并不美观,而且也有限制(如果您想按两列排序怎么办?)。
正如我上面提供的那样,通常最好不要通过管道和子集设置来创建副本,而是通过使用就地修改example_function()来利用效率:data.tablesetorderv()
setorderv(mtcars, x)
Run Code Online (Sandbox Code Playgroud)
这也可以轻松扩展到您想要按多列进行排序的情况。