使用dplyr和add_row()在每个组中添加行

Dan*_*Dan 10 r dplyr tidyverse

如果我向ìris数据集添加一个新行:

iris <- as_tibble(iris)

> iris %>% 
    add_row(.before=0)

# A tibble: 151 × 5
    Sepal.Length Sepal.Width Petal.Length Petal.Width Species
          <dbl>       <dbl>        <dbl>       <dbl>   <chr>
1            NA          NA           NA          NA    <NA> <--- Good!
2           5.1         3.5          1.4         0.2  setosa
3           4.9         3.0          1.4         0.2  setosa
Run Code Online (Sandbox Code Playgroud)

有用.那么,为什么我不能在每个"子集"的顶部添加一个新行:

iris %>% 
 group_by(Species) %>% 
 add_row(.before=0)

Error: is.data.frame(df) is not TRUE
Run Code Online (Sandbox Code Playgroud)

kon*_*vas 14

如果你想使用分组操作,你需要do像JasonWang在他的评论中描述的那样,其他函数喜欢mutate或summarise期望与分组数据框(在你的情况下为50)具有相同行数的结果或者有一行(例如总结时).

正如您可能知道的那样,一般情况下do可能会很慢,如果您无法以其他方式实现结果,则应该是最后的手段.您的任务非常简单,因为它只涉及在数据框中添加额外的行,这可以通过简单的索引来完成,例如查看输出iris[NA, ].

你想要的主要是创建一个矢量

indices <- c(NA, 1:50, NA, 51:100, NA, 101:150)
Run Code Online (Sandbox Code Playgroud)

(因为第一组是第1到第50行,第二组是51到100,第三组是101到150).

结果是iris[indices, ].

构建此向量的更一般方法使用group_indices.

indices <- seq(nrow(iris)) %>% 
    split(group_indices(iris, Species)) %>% 
    map(~c(NA, .x)) %>%
    unlist
Run Code Online (Sandbox Code Playgroud)

(map来自purrr我假设你已加载,因为你已经标记了这一点tidyverse).


Ano*_*n R 6

稍有变化,也可以这样做:

library(purrr)
library(tibble)

iris %>%
  group_split(Species) %>%
  map_dfr(~ .x %>%
            add_row(.before = 1))

# A tibble: 153 x 5
   Sepal.Length Sepal.Width Petal.Length Petal.Width Species
          <dbl>       <dbl>        <dbl>       <dbl> <fct>  
 1         NA          NA           NA          NA   NA     
 2          5.1         3.5          1.4         0.2 setosa 
 3          4.9         3            1.4         0.2 setosa 
 4          4.7         3.2          1.3         0.2 setosa 
 5          4.6         3.1          1.5         0.2 setosa 
 6          5           3.6          1.4         0.2 setosa 
 7          5.4         3.9          1.7         0.4 setosa 
 8          4.6         3.4          1.4         0.3 setosa 
 9          5           3.4          1.5         0.2 setosa 
10          4.4         2.9          1.4         0.2 setosa 
# ... with 143 more rows
Run Code Online (Sandbox Code Playgroud)

这也可以用于分组数据框,但是,它有点冗长:

library(dplyr)

iris %>%
  group_by(Species) %>%
  summarise(Sepal.Length = c(NA, Sepal.Length), 
            Sepal.Width = c(NA, Sepal.Width), 
            Petal.Length = c(NA, Petal.Length),
            Petal.Width = c(NA, Petal.Width), 
            Species = c(NA, Species))
Run Code Online (Sandbox Code Playgroud)


Ale*_*lok 5

更新的版本将使用group_modify()而不是do().

iris %>%
  as_tibble() %>%
  group_by(Species) %>% 
  group_modify(~ add_row(.x,.before=0))
#> # A tibble: 153 x 5
#> # Groups:   Species [3]
#>    Species Sepal.Length Sepal.Width Petal.Length Petal.Width
#>    <fct>          <dbl>       <dbl>        <dbl>       <dbl>
#>  1 setosa          NA          NA           NA          NA  
#>  2 setosa           5.1         3.5          1.4         0.2
#>  3 setosa           4.9         3            1.4         0.2
Run Code Online (Sandbox Code Playgroud)

  • 在我发布问题几年后添加一条评论:截至 2022 年 5 月,group_modify 仍处于“实验”阶段。感谢您的回答亚历克斯洛克 (2认同)