Mat*_*ien 59 r dataframe dplyr
我对dplyr动词有点困惑mutate_each.
使用basic mutate将数据列转换为z-scores,并在data.frame中创建一个新列(此处带有名称z_score_data)非常简单:
newDF <- DF %>%
select(one_column) %>%
mutate(z_score_data = one_column - (mean(one_column) / sd(one_column))
Run Code Online (Sandbox Code Playgroud)
但是,由于我想要转换许多数据列,看起来我应该使用mutate_each动词.
newDF <- DF %>%
mutate_each(funs(scale))
Run Code Online (Sandbox Code Playgroud)
到现在为止还挺好.但到目前为止我还没弄清楚:
mutate?select在第一种情况下所做的那样?谢谢你的帮助.
tal*_*lat 103
在dplyr开发版本0.4.3.9000(撰写本文时),内部命名mutate_each并summarise_each已简化,如新闻中所述:
的命名行为
summarise_each(),并mutate_each()已经过调整,这样可以强制纳入该函数和变量名都的:summarise_each(mtcars, funs(mean = mean), everything())
如果您只想在mutate_each/中应用1个函数summarise_each并且想要为这些列赋予新名称,这一点非常重要.
为了显示差异,这里是使用新命名功能的dplyr 0.4.3.9000的输出,与下面的选项a.2相反:
library(dplyr) # >= 0.4.3.9000
iris %>% mutate_each(funs(mysum = sum(.)), -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species Sepal.Length_mysum Sepal.Width_mysum
#1 5.1 3.5 1.4 0.2 setosa 876.5 458.6
#2 4.9 3.0 1.4 0.2 setosa 876.5 458.6
#3 4.7 3.2 1.3 0.2 setosa 876.5 458.6
#4 4.6 3.1 1.5 0.2 setosa 876.5 458.6
#5 5.0 3.6 1.4 0.2 setosa 876.5 458.6
#6 5.4 3.9 1.7 0.4 setosa 876.5 458.6
# Petal.Length_mysum Petal.Width_mysum
#1 563.7 179.9
#2 563.7 179.9
#3 563.7 179.9
#4 563.7 179.9
#5 563.7 179.9
#6 563.7 179.9
Run Code Online (Sandbox Code Playgroud)
如果您不提供新名称并且您只提供1个函数,则dplyr将更改现有列(与以前版本中一样):
iris %>% mutate_each(funs(sum), -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
#1 876.5 458.6 563.7 179.9 setosa
#2 876.5 458.6 563.7 179.9 setosa
#3 876.5 458.6 563.7 179.9 setosa
#4 876.5 458.6 563.7 179.9 setosa
#5 876.5 458.6 563.7 179.9 setosa
#6 876.5 458.6 563.7 179.9 setosa
Run Code Online (Sandbox Code Playgroud)
我假设这个新功能将在下一个版本0.4.4中通过CRAN提供.
我怎样才能给这些新列提供适当的名称,就像我可以变异一样?
mutate_each/的1个功能summarise_each如果在mutate_eachor中只应用1个函数summarise_each,则现有列将被转换,名称将保持原样,除非你向mutate_each_/ 提供命名向量summarise_each_(参见选项a.4)
这里有些例子:
iris %>% mutate_each(funs(sum), -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
#1 876 459 564 180 setosa
#2 876 459 564 180 setosa
#3 876 459 564 180 setosa
#4 876 459 564 180 setosa
#5 876 459 564 180 setosa
#6 876 459 564 180 setosa
Run Code Online (Sandbox Code Playgroud)
iris %>% mutate_each(funs(mysum = sum(.)), -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
#1 876 459 564 180 setosa
#2 876 459 564 180 setosa
#3 876 459 564 180 setosa
#4 876 459 564 180 setosa
#5 876 459 564 180 setosa
#6 876 459 564 180 setosa
Run Code Online (Sandbox Code Playgroud)
iris %>% mutate_each(funs(sum), SLsum = Sepal.Length,SWsum = Sepal.Width, -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species SLsum SWsum
#1 5.1 3.5 1.4 0.2 setosa 876 459
#2 4.9 3.0 1.4 0.2 setosa 876 459
#3 4.7 3.2 1.3 0.2 setosa 876 459
#4 4.6 3.1 1.5 0.2 setosa 876 459
#5 5.0 3.6 1.4 0.2 setosa 876 459
#6 5.4 3.9 1.7 0.4 setosa 876 459
Run Code Online (Sandbox Code Playgroud)
案例1:保留原始列
与选项a.1,a.2和a.3相反,dplyr将保持现有列不变并在此方法中创建新列.新列的名称等于您事先创建的命名向量的名称(vars在本例中).
vars <- names(iris)[1:2] # choose which columns should be mutated
vars <- setNames(vars, paste0(vars, "_sum")) # create new column names
iris %>% mutate_each_(funs(sum), vars) %>% head
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species Sepal.Length_sum Sepal.Width_sum
#1 5.1 3.5 1.4 0.2 setosa 876.5 458.6
#2 4.9 3.0 1.4 0.2 setosa 876.5 458.6
#3 4.7 3.2 1.3 0.2 setosa 876.5 458.6
#4 4.6 3.1 1.5 0.2 setosa 876.5 458.6
#5 5.0 3.6 1.4 0.2 setosa 876.5 458.6
#6 5.4 3.9 1.7 0.4 setosa 876.5 458.6
Run Code Online (Sandbox Code Playgroud)
案例2:删除原始列
如您所见,此方法保持现有列不变,并添加具有指定名称的新列.如果您不想保留原始列,只想保留新创建的列(以及其他列),则可以在select以后添加语句:
iris %>% mutate_each_(funs(sum), vars) %>% select(-one_of(vars)) %>% head
# Petal.Length Petal.Width Species Sepal.Length_sum Sepal.Width_sum
#1 1.4 0.2 setosa 876.5 458.6
#2 1.4 0.2 setosa 876.5 458.6
#3 1.3 0.2 setosa 876.5 458.6
#4 1.5 0.2 setosa 876.5 458.6
#5 1.4 0.2 setosa 876.5 458.6
#6 1.7 0.4 setosa 876.5 458.6
Run Code Online (Sandbox Code Playgroud)
mutate_each/中应用了多于1个功能summarise_each如果你应用了多个函数,你可以让dplyr自己找出名字(它会保留现有的列):
iris %>% mutate_each(funs(sum, mean), -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species Sepal.Length_sum Sepal.Width_sum Petal.Length_sum
#1 5.1 3.5 1.4 0.2 setosa 876 459 564
#2 4.9 3.0 1.4 0.2 setosa 876 459 564
#3 4.7 3.2 1.3 0.2 setosa 876 459 564
#4 4.6 3.1 1.5 0.2 setosa 876 459 564
#5 5.0 3.6 1.4 0.2 setosa 876 459 564
#6 5.4 3.9 1.7 0.4 setosa 876 459 564
# Petal.Width_sum Sepal.Length_mean Sepal.Width_mean Petal.Length_mean Petal.Width_mean
#1 180 5.84 3.06 3.76 1.2
#2 180 5.84 3.06 3.76 1.2
#3 180 5.84 3.06 3.76 1.2
#4 180 5.84 3.06 3.76 1.2
#5 180 5.84 3.06 3.76 1.2
#6 180 5.84 3.06 3.76 1.2
Run Code Online (Sandbox Code Playgroud)
使用多个函数时,另一个选项是自己指定列名扩展名:
iris %>% mutate_each(funs(MySum = sum(.), MyMean = mean(.)), -Species) %>% head()
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species Sepal.Length_MySum Sepal.Width_MySum Petal.Length_MySum
#1 5.1 3.5 1.4 0.2 setosa 876 459 564
#2 4.9 3.0 1.4 0.2 setosa 876 459 564
#3 4.7 3.2 1.3 0.2 setosa 876 459 564
#4 4.6 3.1 1.5 0.2 setosa 876 459 564
#5 5.0 3.6 1.4 0.2 setosa 876 459 564
#6 5.4 3.9 1.7 0.4 setosa 876 459 564
# Petal.Width_MySum Sepal.Length_MyMean Sepal.Width_MyMean Petal.Length_MyMean Petal.Width_MyMean
#1 180 5.84 3.06 3.76 1.2
#2 180 5.84 3.06 3.76 1.2
#3 180 5.84 3.06 3.76 1.2
#4 180 5.84 3.06 3.76 1.2
#5 180 5.84 3.06 3.76 1.2
#6 180 5.84 3.06 3.76 1.2
Run Code Online (Sandbox Code Playgroud)
如何选择我想改变的某些列,就像我在第一种情况下选择select一样?
您可以通过引用要变异(或省略)的列来实现这一点,方法是在此处给出它们的名称(mutate Sepal.Length,但不是Species):
iris %>% mutate_each(funs(sum), Sepal.Length, -Species) %>% head()
Run Code Online (Sandbox Code Playgroud)
此外,您可以使用特殊函数来选择要变异的列,所有以特定单词开头或包含某些单词的列等,例如:
iris %>% mutate_each(funs(sum), contains("Sepal"), -Species) %>% head()
Run Code Online (Sandbox Code Playgroud)
有关这些功能的更多信息,请参阅?mutate_each和?select.
如果要使用标准评估,dplyr会提供大多数以附加"_"结尾的函数的SE版本.所以在这种情况下你会使用:
x <- c("Sepal.Width", "Sepal.Length") # vector of column names
iris %>% mutate_each_(funs(sum), x) %>% head()
Run Code Online (Sandbox Code Playgroud)
请注意mutate_each_我在这里使用的.
编辑2:使用选项a.4更新
pab*_*sci 13
mutate_each将被弃用,请考虑使用mutate_at.来自dplyr_0.5.0文档:
将来mutate_each()和summarise_each()将被弃用,以支持更多功能更强大的函数:mutate_all(),mutate_at(),mutate_if(),summarise_all(),summarise_at()和summarise_if().
Species:警告:不推荐使用'.cols'param,请参阅底部的注释!
iris %>% mutate_at(.cols=vars(-Species), .funs=funs(mysum = sum(.))) %>% head()
Sepal.Length Sepal.Width Petal.Length Petal.Width Species Sepal.Length_mysum Sepal.Width_mysum
1 5.1 3.5 1.4 0.2 setosa 876.5 458.6
2 4.9 3.0 1.4 0.2 setosa 876.5 458.6
3 4.7 3.2 1.3 0.2 setosa 876.5 458.6
4 4.6 3.1 1.5 0.2 setosa 876.5 458.6
5 5.0 3.6 1.4 0.2 setosa 876.5 458.6
6 5.4 3.9 1.7 0.4 setosa 876.5 458.6
Petal.Length_mysum Petal.Width_mysum
1 563.7 179.9
2 563.7 179.9
3 563.7 179.9
4 563.7 179.9
5 563.7 179.9
6 563.7 179.9
Run Code Online (Sandbox Code Playgroud)
vars_to_process=c("Petal.Length","Petal.Width")
iris %>% mutate_at(.cols=vars_to_process, .funs=funs(mysum = sum(.))) %>% head()
Sepal.Length Sepal.Width Petal.Length Petal.Width Species Petal.Length_mysum Petal.Width_mysum
1 5.1 3.5 1.4 0.2 setosa 563.7 179.9
2 4.9 3.0 1.4 0.2 setosa 563.7 179.9
3 4.7 3.2 1.3 0.2 setosa 563.7 179.9
4 4.6 3.1 1.5 0.2 setosa 563.7 179.9
5 5.0 3.6 1.4 0.2 setosa 563.7 179.9
6 5.4 3.9 1.7 0.4 setosa 563.7 179.9
Run Code Online (Sandbox Code Playgroud)
如果您看到以下消息:
.cols已重命名并已弃用,请使用.vars
然后换.cols用.vars.
iris %>% mutate_at(.vars=vars(-Species), .funs=funs(mysum = sum(.))) %>% head()
Run Code Online (Sandbox Code Playgroud)
另一个例子:
iris %>% mutate_at(.vars=vars(Sepal.Width), .funs=funs(mysum = sum(.))) %>% head()
Run Code Online (Sandbox Code Playgroud)
相当于:
iris %>% mutate_at(.vars=vars("Sepal.Width"), .funs=funs(mysum = sum(.))) %>% head()
Run Code Online (Sandbox Code Playgroud)
此外,在此版本中,mutate_each不推荐使用:
mutate_each()已弃用.使用mutate_all(),mutate_at()或mutate_if()改为.要映射funs选择的变量,请使用mutate_at()