dplyr case_when的data.table替代方案

arg*_*t91 6 if-statement r dplyr data.table

前段时间他们推出一个很好的类似SQL的替代ifelse之内dplyr,即case_when.

是否有相应的data.table允许您在一个[]语句中指定不同的条件,而不加载其他包?

例:

library(dplyr)

df <- data.frame(a = c("a", "b", "a"), b = c("b", "a", "a"))

df <- df %>% mutate(
    new = case_when(
    a == "a" & b == "b" ~ "c",
    a == "b" & b == "a" ~ "d",
    TRUE ~ "e")
    )

  a b new
1 a b   c
2 b a   d
3 a a   e
Run Code Online (Sandbox Code Playgroud)

它肯定会非常有用,并使代码更具可读性(这是我dplyr在这些情况下继续使用的原因之一).

ske*_*ook 27

仅供参考,对于那些在 2019 年之后遇到的人的最新答案data.table。1.13.0 以上的版本具有fcase可以使用的功能。请注意,dplyr::case_when由于语法不同,它不是替代品,而是一种“本机”data.table计算方式。

# Lazy evaluation
x = 1:10
data.table::fcase(
    x < 5L, 1L,
    x >= 5L, 3L,
    x == 5L, stop("provided value is an unexpected one!")
)
# [1] 1 1 1 1 3 3 3 3 3 3

dplyr::case_when(
    x < 5L ~ 1L,
    x >= 5L ~ 3L,
    x == 5L ~ stop("provided value is an unexpected one!")
)
# Error in eval_tidy(pair$rhs, env = default_env) :
#  provided value is an unexpected one!

# Benchmark
x = sample(1:100, 3e7, replace = TRUE) # 114 MB
microbenchmark::microbenchmark(
dplyr::case_when(
  x < 10L ~ 0L,
  x < 20L ~ 10L,
  x < 30L ~ 20L,
  x < 40L ~ 30L,
  x < 50L ~ 40L,
  x < 60L ~ 50L,
  x > 60L ~ 60L
),
data.table::fcase(
  x < 10L, 0L,
  x < 20L, 10L,
  x < 30L, 20L,
  x < 40L, 30L,
  x < 50L, 40L,
  x < 60L, 50L,
  x > 60L, 60L
),
times = 5L,
unit = "s")
# Unit: seconds
#               expr   min    lq  mean   median    uq    max neval
# dplyr::case_when   11.57 11.71 12.22    11.82 12.00  14.02     5
# data.table::fcase   1.49  1.55  1.67     1.71  1.73   1.86     5
Run Code Online (Sandbox Code Playgroud)

来源,data.table NEWS for 1.13.0,发布(2020 年 7 月 24 日)。


G. *_*eck 17

1)如果条件与默认值互斥,如果所有条件都为假,那么这有效:

library(data.table)
DT <- as.data.table(df) # df is from question

DT[, new := c("e", "c", "d")[1 +
                             1 * (a == "a" & b == "b") + 
                             2 * (a == "b" & b == "a")]
]
Run Code Online (Sandbox Code Playgroud)

赠送:

> DT
   a b new
1: a b   c
2: b a   d
3: a a   e
Run Code Online (Sandbox Code Playgroud)

2)如果条件的结果是数字,则更容易.例如,假设代替和c而d我们想要10和17,默认值为3.然后:

library(data.table)
DT <- as.data.table(df) # df is from question

DT[, new := 3 + 
            (10 - 3) * (a == "a" & b == "b") + 
            (17 - 3) * (a == "b" & b == "a")]
Run Code Online (Sandbox Code Playgroud)

3)注意添加1个衬管就足以实现这一点.它假定每行至少有一条TRUE支路.

when <- function(...) names(match.call()[-1])[apply(cbind(...), 1, which.max)]

# test
DT[, new := when(c = a == 'a' & b == 'b', 
                 d = a == 'b' & b == 'a', 
                 e = TRUE)]
Run Code Online (Sandbox Code Playgroud)

  • 用户只需使用该包,而不必查看其中的代码。维护它的是开发人员/维护人员。 (2认同)
  • 如果要手动分配,请使用原始的“何时”并将结果转换为数字。`when_num &lt;-function(...)as.numeric(when(...)); DT [,new:= when(“ 1” = a =='a'&b =='b',“ 2” = a =='b'&b =='a',“ 99” = TRUE) ]` (2认同)

Mau*_*ers 16

这不是一个真正的答案,但对评论来说有点太长了.如果认为不合适,我很乐意删除帖子.

RStudio社区上有一篇有趣的帖子,讨论了在dplyr::case_when没有通常tidyverse依赖性的情况下使用的选项.

总而言之,似乎存在三种替代方案:

  1. Stefan Fleckcase_when从中分离dplyr并构建了一个lest仅依赖于的新包base.
  2. yonicd开发noplyr,"提供基本dplyr和tidyr功能,没有整齐的依赖".
  3. Bob Rudis(hrbrmstr)创建了freebase一个"A'使用此类'的包,用于'Tidyverse'代码的Base R Pseudo-equivalent",这可能也值得一试.

如果只是case_when你所追求的,我想lest可能是一个有吸引力和最小的选择与结合data.table.

  • 非常感谢,这非常有帮助.我还没有听说过'lest`,这确实是我要开始测试的.我决不认为评论不合适,我只是让它停留一段时间而不接受看看其他解决方案`data.table`用户可能会带来什么. (2认同)