使用字符串指定 data.table 列

Tee*_*Vee 6 r data.table

考虑以下最小工作示例:

library(magrittr) # for the %>% pipe 
library(data.table) 

# test data.table contains common_column and two others
test_dt <- data.table(test_column_one = c(1, 2, 3), test_column_two = c("x","y","z"), common_column = c("ID1", "ID2", "ID3") ) 

# some other data.table that contains common_column
other_dt <- data.table( additional_info = c("US", "US", "GB"), common_column = c("ID1", "ID2", "ID3")) 

example_function <- function(dt_column){
  # does some things on the data tables based on the column parameter passed
  merged_dt <- merge(other_dt, test_dt[,.(common_column, dt_column)], by = "common_column")[order(dt_column),] # order by the dt_column
  return(merged_dt)
} 

# calling the example function
example_function(test_dt$test_column_one)
Run Code Online (Sandbox Code Playgroud)

我怎样才能将代码修改为:

  1. 能够指定一个以列名作为参数的字符串
  2. 能够传递带有列名称的向量或字符串列表

我想避免 for 循环并尽可能利用优化的 data.table 语法。

我尝试使用data.tableunlist()特定..的语法,但不知何故我总是收到奇怪的错误消息,并且不确定如何继续。

Tob*_*obo 5

env 用于 data.table 编程的新接口解决了所有这些问题,并且不需要其他方法和 hack。

使用 SamR 的很好的例子,从 mt_cars(在 j 中)中选择“mpg” ,同时也在(在 i 中)some_col进行排序:some_col

setDT(mtcars)
some_col <- "wt"
mtcars[order(v), .(mpg, v), env=list(v=some_col)]
Run Code Online (Sandbox Code Playgroud)

请参阅此处的新闻第 10 项和此小插图。要更新到开发版本 1.14.99:data.table::update_dev_pkg().

  • 希望在几天之内来到 CRAN! (4认同)

Sam*_*amR 4

创建列名称向量

我在下面列出了三种方法来做到这一点,但我认为最清楚的是创建一个与语法一起使用的列向量..。

example_function <- function(dt_column, dt1 = other_dt, dt2 = test_dt) {
    cols_to_merge <- c("common_column", dt_column)
    merged_dt <- merge(
        dt1,
        dt2[, ..cols_to_merge],
        by = "common_column"
    )

    # No need to pipe - see explanation below
    setorderv(merged_dt, dt_column)
    return(merged_dt)
}

example_function("test_column_one")
#    common_column additional_info test_column_one
#           <char>          <char>           <num>
# 1:           ID1              US               1
# 2:           ID2              US               2
# 3:           ID3              GB               3
Run Code Online (Sandbox Code Playgroud)

与其他方法的比较

假设我们要从mpg中选择一个动态列mtcars,在本例中恰好是wt。

mtcars <- as.data.table(mtcars)
col_of_interest <- "wt"
Run Code Online (Sandbox Code Playgroud)

我认为基本上有三种data.table方法。

  1. 如上。正如data.table v1.10.2(2017 年 1 月 31 日)文档所述:

Whenj是以 it 为前缀的符号,..将在调用范围中查找,其值取为列名或数字。

cols  <- c("mpg", col_of_interest)
mtcars[, ..cols]
Run Code Online (Sandbox Code Playgroud)
  1. 自data.table v1.11.0(2018 年 5 月 1 日)起:

这个 [the ..] 前缀现在扩展到出现在的所有符号j=

mtcars[, c("mpg", ..col_of_interest)]
Run Code Online (Sandbox Code Playgroud)
  1. 如使用 .SD 进行数据分析小插图中所述。
cols <- c("mpg", col_of_interest)
mtcars[, .SD, .SDcols = cols]
Run Code Online (Sandbox Code Playgroud)

我个人认为创建列名向量是最清晰的方法,但第二种方法意味着创建的变量少了一个。第三种方法是向后兼容的。

使用注意事项order()

文档指定..仅适用于j. 以下语句是等效的:

mtcars[order(mpg)]
mtcars[order(mpg),]
mtcars[i = order(mpg),]
Run Code Online (Sandbox Code Playgroud)

正如order(mpg)提供的参数一样i,如果我们设置了x <- "mpg",那么这样做是无效的mtcars[order(..x)]。

如果您必须使用order(),可以通过以下一些方法来实现:

# Use ..
mtcars[order(mtcars[, ..x])]
# Use `[[`
mtcars[order(mtcars[[x]])]
# Use .SD
mtcars[mtcars[, sapply(.SD, order), .SDcols = x]]
# Use get() (this has been retired in favour of `env` - see comments)
mtcars[order(get(x))]
Run Code Online (Sandbox Code Playgroud)

然而,这些创建了数据的副本,它们并不美观,而且也有限制(如果您想按两列排序怎么办?)。

正如我上面提供的那样,通常最好不要通过管道和子集设置来创建副本,而是通过使用就地修改example_function()来利用效率:data.tablesetorderv()

setorderv(mtcars, x)
Run Code Online (Sandbox Code Playgroud)

这也可以轻松扩展到您想要按多列进行排序的情况。

  • 虽然这是一个非常好的答案,但这些接口(包括其他未提及的接口,例如“mget”和“with=F”)被认为已退役或至少即将退役。“data.table” 编程的新标准使用“env”,它以干净一致的方式跨“i”、“j”和“by”工作。 (3认同)
  • 查看您的示例,我们希望到达“mtcars[order(wt), .(mpg, wt)]”,其中“wt”是以编程方式定义的。在 `some_col &lt;- wt` 之后,实现与您所说的相同结果的一种方法是 `cols &lt;- c("mpg", some_col); setorderv(mtcars[, ..cols], some_col)`。但是 `env` 现在让我们可以直接执行 `mtcars[order(v), .(mpg, v), env = list(v = some_col)]` 。由于 `env` 是通用的,所以它会混合习惯用法来执行 `mtcars[order(v), ..cols, env = list(v = some_col)]` (但它有效)。另一方面,如果我们需要做的只是在 j 中进行替换,那么当然,为什么不继续使用“...”。 (3认同)
  • 与使用新环境变量相比,get、mget、eval(不确定 ..var)可能会导致代码速度变慢,因为它们可能会抑制某些代码优化。Env 的设计目的是通过从一开始就进行替换来绕过所有这些,生成的替换代码就像用户编写的一样,从而有利于 DT 中所有可能的优化。 (3认同)
  • 你好@SamR - 问题是关于以编程方式定义变量的通用方法,这明确是新的“env”最终设计要解决的问题。`..` (2018) 允许在 j 中进行替换,但 `env` (2023) 现在提供了一个通用接口,用于跨所有三个参数执行所有操作。该插图给出了“get”、“mget”和“eval”作为现在可以退役的注入/替换方法的示例,当然,没有任何内容被删除或以某种方式禁止。 (2认同)