Ves*_*sna 3 r median ggplot2 boxplot dplyr
我正在根据大数据(2150000 例)绘制一个简单的箱线图,显示两组每年的体重。除去年的最后一组外,所有组的中位数均相同,但在箱线图上,它的绘制方式与其他组相同。
#boxplot
ggplot(dataset, aes(x=Year, y=SUM_MME_mg, fill=GenderPerson)) +
geom_boxplot(outlier.shape = NA)+
ylim(0,850)
#median by group
pivot <- dataset %>%
select(SUM_MME_mg,GenderPerson,Year )%>%
group_by(Year, GenderPerson) %>%
summarise(MedianValues = median(SUM_MME_mg,na.rm=TRUE))
Run Code Online (Sandbox Code Playgroud)
我无法弄清楚我做错了什么,或者在箱线图计算或中值函数中哪些数据更准确。R 不返回错误或警告。
#my data:
> dput(head(dataset[,c(1,7,10)]))
structure(list(GenderPerson = c(2L, 1L, 2L, 2L, 2L, 2L), Year = c("2015",
"2014", "2013", "2012", "2011", "2015"), SUM_MME_mg = c(416.16,
131.76, 790.56, 878.4, 878.4, 878.4)), row.names = c(NA, 6L), class = "data.frame")
Run Code Online (Sandbox Code Playgroud)
这种行为的原因与ylim()操作方式有关。 ylim()是一个方便的函数/包装器scale_y_continuous(limits=...。如果您查看这些scale_continuous函数的文档,您会发现设置限制不仅会放大某个区域,而且实际上还会删除该区域之外的所有数据点。这发生在计算/统计函数之前,因此这就是当您使用 时中位数不同的原因ylim()。您的“外部”计算ggplot()正在获取整个数据集,而使用 意味着ylim()在进行计算之前删除数据点。
幸运的是,有一个简单的解决方案,即使用coord_cartesian(ylim=...)代替ylim(), 因为coord_cartesian()它只会放大数据而不会删除数据点。看看这里的区别:
ggplot(dataset, aes(x=Year, y=SUM_MME_mg, fill=GenderPerson)) +
geom_boxplot(outlier.shape = NA) + ylim(0,850)
Run Code Online (Sandbox Code Playgroud)
ggplot(dataset, aes(x=Year, y=SUM_MME_mg, fill=GenderPerson)) +
geom_boxplot(outlier.shape = NA) + coord_cartesian(ylim=c(0,850))
Run Code Online (Sandbox Code Playgroud)
此行为的提示也应该很明显,因为使用的第一个代码块ylim()还应该给您一条警告消息:
Warning message:
Removed 3 rows containing non-finite values (stat_boxplot).
Run Code Online (Sandbox Code Playgroud)
而第二次使用则coord_cartesian(ylim=不然。