我正在尝试使用 Andy Field 的教科书Discovering Statistics Using R 中的代码进行引导回归模型。
我正在努力解释运行该boot()函数时收到的错误消息。通过阅读其他论坛帖子,我了解到它告诉我两个对象之间的项目数量不平衡,但我不明白这在我的上下文中意味着什么以及如何解决它。
您可以在此处下载我的数据(Airbnb 列表上的公开数据集)并在下面找到我的代码和完整的错误消息。我使用因子虚拟变量和连续变量的混合作为预测变量。在此先感谢您的帮助!
代码:
bootReg <- function (formula, data, i)
{
d <- data [i,]
fit <- lm(formula, data = d)
return(coef(fit))
}
bootResults <- boot(statistic = bootReg, formula = review_scores_rating ~ instant_bookable + cancellation_policy +
host_since_cat + host_location_cat + host_response_time +
host_is_superhost + host_listings_cat + property_type + room_type +
accommodates + bedrooms + beds + price + security_deposit +
cleaning_fee + extra_people + minimum_nights + amenityBreakfast +
amenityAC + amenityElevator + amenityKitchen + amenityHostGreeting +
amenitySmoking + amenityPets + amenityWifi + amenityTV,
data = listingsRating, R = 2000)
Run Code Online (Sandbox Code Playgroud)
错误:
Error in t.star[r, ] <- res[[r]] :
number of items to replace is not a multiple of replacement length
In addition: Warning message:
In doTryCatch(return(expr), name, parentenv, handler) :
restarting interrupted promise evaluation
Run Code Online (Sandbox Code Playgroud)
问题是你的因子变量。当您lm()对数据的一个子集进行操作时(在 中一遍又一遍地完成boot::boot()),您只能获得现有因子水平的系数。然后每个系数绘制可以是不同的长度。如果你这样做,这可以被复制
debug(boot)
set.seed(123)
bootResults <- boot(statistic = bootReg, formula = review_scores_rating ~ instant_bookable + cancellation_policy +
host_since_cat + host_location_cat + host_response_time +
host_is_superhost + host_listings_cat + property_type + room_type +
accommodates + bedrooms + beds + price + security_deposit +
cleaning_fee + extra_people + minimum_nights + amenityBreakfast +
amenityAC + amenityElevator + amenityKitchen + amenityHostGreeting +
amenitySmoking + amenityPets + amenityWifi + amenityTV,
data = listingsRating, R = 2)
Run Code Online (Sandbox Code Playgroud)
这将允许您一次一行地浏览函数调用。跑完这条线后
res <- if (ncpus > 1L && (have_mc || have_snow)) {
if (have_mc) {
parallel::mclapply(seq_len(RR), fn, mc.cores = ncpus)
}
else if (have_snow) {
list(...)
if (is.null(cl)) {
cl <- parallel::makePSOCKcluster(rep("localhost",
ncpus))
if (RNGkind()[1L] == "L'Ecuyer-CMRG")
parallel::clusterSetRNGStream(cl)
res <- parallel::parLapply(cl, seq_len(RR), fn)
parallel::stopCluster(cl)
res
}
else parallel::parLapply(cl, seq_len(RR), fn)
}
} else lapply(seq_len(RR), fn)
Run Code Online (Sandbox Code Playgroud)
然后试试
setdiff(names(res[[1]]), names(res[[2]]))
# [1] "property_typeBarn" "property_typeNature lodge"
Run Code Online (Sandbox Code Playgroud)
第一个子集中存在两个因子水平,第二个子集中不存在。这导致了您的问题。
用于model.matrix()事先扩展您的因素(在此 Stack Overflow 帖子之后):
df2 <- model.matrix( ~ review_scores_rating + instant_bookable + cancellation_policy +
host_since_cat + host_location_cat + host_response_time +
host_is_superhost + host_listings_cat + property_type + room_type +
accommodates + bedrooms + beds + price + security_deposit +
cleaning_fee + extra_people + minimum_nights + amenityBreakfast +
amenityAC + amenityElevator + amenityKitchen + amenityHostGreeting +
amenitySmoking + amenityPets + amenityWifi + amenityTV - 1, data = listingsRating)
undebug(boot)
set.seed(123)
bootResults <- boot(statistic = bootReg, formula = review_scores_rating ~ .,
data = as.data.frame(df2), R = 2)
Run Code Online (Sandbox Code Playgroud)
(请注意,R为了在调试期间更快的运行时间,我一直将其减少到 2)。
| 归档时间: |
|
| 查看次数: |
2108 次 |
| 最近记录: |