用R重塑表格 - 更好的方法?

use*_*931 7 r plyr reshape reshape2

我有一个被称为因子的数据框 questions

q1 q2 q3
A  A  B
C  A  A
A  B  C
Run Code Online (Sandbox Code Playgroud)

我想重塑一下

question answer freq
1        A      2
1        B      0
1        C      1
2        A      2
2        B      1
2        C      0
3        A      1
3        B      1
3        C      1
Run Code Online (Sandbox Code Playgroud)

我觉得应该有一种方法可以用reshape2或plyr,但我无法理解.

相反,我做了以下事情:

tbl <- data.frame()
for(i in 1:dim(questions)[2]){
    subtable <- cbind(question = rep(i, 3),
                      as.data.frame(table(questions[i])))
    tbl <- rbind(tbl, subtable)
}
Run Code Online (Sandbox Code Playgroud)

是否有更清洁的方法来重塑这张桌子?

A5C*_*2T1 5

这是一个基本R方法,其概念与@akrun发布的方法类似.我没有打扰清理,因为这主要是化妆品,与问题的概念无关.

一般方法是:

data.frame(table(stack(mydf))
Run Code Online (Sandbox Code Playgroud)

但是,stack不能使用factors,所以如果您的数据是factors而不是characters,则必须先使用转换as.character,如下所示:

data.frame(table(stack(lapply(mydf, as.character))))
#   values ind Freq
# 1      A  q1    2
# 2      B  q1    0
# 3      C  q1    1
# 4      A  q2    2
# 5      B  q2    1
# 6      C  q2    0
# 7      A  q3    1
# 8      B  q3    1
# 9      C  q3    1
Run Code Online (Sandbox Code Playgroud)

远离"plyr"和"reshape2"而不是"dplyr"和"tidyr",您可以尝试:

library(dplyr)
library(tidyr)

mydf %>% 
  gather(question, answer, everything()) %>%  ## Get the data into a long form
  group_by(question, answer) %>%              ## Group by both question and answer columns
  summarise(freq = n()) %>%                   ## Calculate the relevant frequency
  right_join(expand(., question, answer))     ## Merge with all combinations of Qs and As
# Joining by: c("question", "answer")
# Source: local data frame [9 x 3]
# Groups: question
# 
#   question answer freq
# 1       q1      A    2
# 2       q1      B   NA
# 3       q1      C    1
# 4       q2      A    2
# 5       q2      B    1
# 6       q2      C   NA
# 7       q3      A    1
# 8       q3      B    1
# 9       q3      C    1
Run Code Online (Sandbox Code Playgroud)