在R 3.4.3和R 3.6.0之间更改data.table中存储的闭包的行为

rlh*_*lh2 5 r data.table

从R 3.4.3升级到R 3.6.0时,我注意到以下特殊行为(两者都使用data.table 1.12.6)。在3.4.3中,以下代码导致all.equal语句为TRUE,但在3.6.0中,存在一个相对平均差,这是因为即使我们试图访问从“ a”组计算出的近似值,使用“ b”组中的值(可能是由于延迟评估而造成的)。在3.6.0中,可以基于以下问题在对roxofun的调用中添加一个copy语句来解决此问题: data.table中闭包的处理

对我而言,令人着迷的是我在3.4.3中没有得到错误。知道发生了什么变化吗?

library(data.table)
data <- data.table(
  group = c(rep("a", 4), rep("b", 4)),
  x = rep(c(.02, .04, .12, .21), 2),
  y = c(
    0.0122, 0.01231, 0.01325, 0.01374, 0.01218, 0.01229, 0.0133, 0.01379)
)

dtFuncs <- data[ , list(
  func = list(stats::approxfun(x, y, rule = 2))
), by = group]

f <- function(group, x) {
  dtResults <- CJ(group = group, x = x)
  dtResults <- dtResults[ , {
   .g <- group
    f2 <- dtFuncs[group == .g, func][[1]]
    list(x = x, y = f2(x))
  }, by = group] 
  dtResults
}

x0 <- .07
g <- "a"
all.equal(
  with(data[group == g], approx(x, y, x0, rule = 2)$y),
  f(group = g, x = x0)$y
)
Run Code Online (Sandbox Code Playgroud)

rlh*_*lh2 1

在 r 源上运行 git bisect 后,我​​能够推断出正是此提交导致了该行为:https://github.com/wch/r-source/commit/adcf18b773149fa20f289f2c8f2e45e6f7b0dbfe

根本上发生的事情是,当 x 在 approxfun 中排序时,不再制作内部副本。如果数据是随机排序的,代码将继续工作!(见下面的片段)

对我来说,最好不要将复杂的对象与 data.table 混合在一起,因为每个“by”组都会一遍又一遍地使用相同的环境(或者非常谨慎地使用 data.table::copy)

## should be run under R > 3.6.0 to see disparity
library(data.table)

## original sorted x (does not work)
data <- data.table(
  group = c(rep("a", 4), rep("b", 4)),
  x = rep(c(.02, .04, .12, .21), 2),
  y = c(
    0.0122, 0.01231, 0.01325, 0.01374, 0.01218, 0.01229, 0.0133, 0.01379)
)

dtFuncs <- data[ , {
    print(environment())
    list(
        func = list(stats::approxfun(x, y, rule = 2))
    )
}, by = group]

f <- function(group, x) {
  dtResults <- CJ(group = group, x = x)
  dtResults <- dtResults[ , {
   .g <- group
    f2 <- dtFuncs[group == .g, func][[1]]
    list(x = x, y = f2(x))
  }, by = group] 
  dtResults
}

get("y", environment(dtFuncs$func[[1]]))
get("y", environment(dtFuncs$func[[2]]))

x0 <- .07
g <- "a"
all.equal(
  with(data[group == g], approx(x, y, x0, rule = 2)$y),
  f(group = g, x = x0)$y
)

## unsorted x (works)
data <- data.table(
  group = c(rep("a", 4), rep("b", 4)),
  x = rep(c(.02, .04, .12, .21), 2),
  y = c(
    0.0122, 0.01231, 0.01325, 0.01374, 0.01218, 0.01229, 0.0133, 0.01379)
)
set.seed(10)
data <- data[sample(1:.N, .N)]
dtFuncs <- data[ , {
    print(environment())
    list(
        func = list(stats::approxfun(x, y, rule = 2))
    )
}, by = group]

f <- function(group, x) {
  dtResults <- CJ(group = group, x = x)
  dtResults <- dtResults[ , {
   .g <- group
    f2 <- dtFuncs[group == .g, func][[1]]
    list(x = x, y = f2(x))
  }, by = group] 
  dtResults
}

get("y", environment(dtFuncs$func[[1]]))
get("y", environment(dtFuncs$func[[2]]))

x0 <- .07
g <- "a"
all.equal(
  with(data[group == g], approx(x, y, x0, rule = 2)$y),
  f(group = g, x = x0)$y
)

## better approach: maybe safer to avoid mixing objects treated by reference
## (data.table & closures) all together...
fList <- lapply(split(data, by = "group"), function(x){
    with(x, stats::approxfun(x, y, rule = 2))
})
fList
fList[[1]](.07) != fList[[2]](.07)
Run Code Online (Sandbox Code Playgroud)