尝试使用dplyr到group_by并应用scale()

Jos*_*erg 20 r dplyr

试图使用dplyr到group_by的stud_ID变量在以下的数据帧,如在此SO问题:

> str(df)
'data.frame':   4136 obs. of  4 variables:
 $ stud_ID         : chr  "ABB112292" "ABB112292" "ABB112292" "ABB112292" ...
 $ behavioral_scale: num  3.5 4 3.5 3 3.5 2 NA NA 1 2 ...
 $ cognitive_scale : num  3.5 3 3 3 3.5 2 NA NA 1 1 ...
 $ affective_scale : num  2.5 3.5 3 3 2.5 2 NA NA 1 1.5 ...
Run Code Online (Sandbox Code Playgroud)

我尝试了以下方法来获得学生的量表分数(而不是所有学生的观察量表得分):

scaled_data <- 
          df %>%
              group_by(stud_ID) %>%
                  mutate(behavioral_scale_ind = scale(behavioral_scale),
                         cognitive_scale_ind = scale(cognitive_scale),
                         affective_scale_ind = scale(affective_scale))
Run Code Online (Sandbox Code Playgroud)

结果如下:

> str(scaled_data)
Classes ‘grouped_df’, ‘tbl_df’, ‘tbl’ and 'data.frame': 4136 obs. of  7 variables:
 $ stud_ID             : chr  "ABB112292" "ABB112292" "ABB112292" "ABB112292" ...
 $ behavioral_scale    : num  3.5 4 3.5 3 3.5 2 NA NA 1 2 ...
 $ cognitive_scale     : num  3.5 3 3 3 3.5 2 NA NA 1 1 ...
 $ affective_scale     : num  2.5 3.5 3 3 2.5 2 NA NA 1 1.5 ...
 $ behavioral_scale_ind: num [1:12, 1] 0.64 1.174 0.64 0.107 0.64 ...
  ..- attr(*, "scaled:center")= num 2.9
  ..- attr(*, "scaled:scale")= num 0.937
 $ cognitive_scale_ind : num [1:12, 1] 1.17 0.64 0.64 0.64 1.17 ...
  ..- attr(*, "scaled:center")= num 2.4
  ..- attr(*, "scaled:scale")= num 0.937
 $ affective_scale_ind : num [1:12, 1] 0 1.28 0.64 0.64 0 ...
  ..- attr(*, "scaled:center")= num 2.5
  ..- attr(*, "scaled:scale")= num 0.782
Run Code Online (Sandbox Code Playgroud)

三个缩放变量(behavioral_scale,cognitive_scale和affective_scale)只有12个观察结果 - 第一个学生的观察数量相同,ABB112292.

这里发生了什么?我怎样才能获得个人的比例分数?

C8H*_*4O2 31

问题似乎在基scale()函数中,它需要一个矩阵.尝试自己写.

scale_this <- function(x){
  (x - mean(x, na.rm=TRUE)) / sd(x, na.rm=TRUE)
}
Run Code Online (Sandbox Code Playgroud)

然后这工作:

library("dplyr")

# reproducible sample data
set.seed(123)
n = 1000
df <- data.frame(stud_ID = sample(LETTERS, size=n, replace=TRUE),
                 behavioral_scale = runif(n, 0, 10),
                 cognitive_scale = runif(n, 1, 20),
                 affective_scale = runif(n, 0, 1) )
scaled_data <- 
  df %>%
  group_by(stud_ID) %>%
  mutate(behavioral_scale_ind = scale_this(behavioral_scale),
         cognitive_scale_ind = scale_this(cognitive_scale),
         affective_scale_ind = scale_this(affective_scale))
Run Code Online (Sandbox Code Playgroud)

或者,如果您对data.table解决方案持开放态度:

library("data.table")

setDT(df)

cols_to_scale <- c("behavioral_scale","cognitive_scale","affective_scale")

df[, lapply(.SD, scale_this), .SDcols = cols_to_scale, keyby = factor(stud_ID)] 
Run Code Online (Sandbox Code Playgroud)


krl*_*mlr 10

这是dplyr 中的一个已知问题,修复程序已合并到开发版本,您可以通过它安装

# install.packages("devtools")
devtools::install_github("hadley/dplyr")
Run Code Online (Sandbox Code Playgroud)

在稳定版本中,以下内容也应该有效:

scale_this <- function(x) as.vector(scale(x))
Run Code Online (Sandbox Code Playgroud)

  • 啊没关系。它可以工作,但是列名称随后显示为“[,1]”,这可能有点令人困惑。 (4认同)

Spa*_*ter 9

df <- df %>% mutate(across(is.numeric, ~ as.numeric(scale(.))))
Run Code Online (Sandbox Code Playgroud)

  • 虽然此代码可以回答问题,但提供有关如何和/或为何解决问题的附加上下文将提高​​答案的长期价值。您可以在帮助中心找到有关如何编写良好答案的更多信息:https://stackoverflow.com/help/how-to-answer。祝你好运 (5认同)