如何解决引导回归中的“要替换的项目数不是替换长度的倍数”错误?

Dav*_*alf 2 r

我正在尝试使用 Andy Field 的教科书Discovering Statistics Using R 中的代码进行引导回归模型。

我正在努力解释运行该boot()函数时收到的错误消息。通过阅读其他论坛帖子,我了解到它告诉我两个对象之间的项目数量不平衡,但我不明白这在我的上下文中意味着什么以及如何解决它。

您可以在此处下载我的数据(Airbnb 列表上的公开数据集)并在下面找到我的代码和完整的错误消息。我使用因子虚拟变量和连续变量的混合作为预测变量。在此先感谢您的帮助!

代码:

bootReg <- function (formula, data, i)
{
d <- data [i,]
fit <- lm(formula, data = d)
return(coef(fit))
}

bootResults <- boot(statistic = bootReg, formula = review_scores_rating ~ instant_bookable + cancellation_policy + 
                  host_since_cat + host_location_cat + host_response_time + 
                  host_is_superhost + host_listings_cat + property_type + room_type + 
                  accommodates + bedrooms + beds + price + security_deposit + 
                  cleaning_fee + extra_people + minimum_nights + amenityBreakfast + 
                  amenityAC + amenityElevator + amenityKitchen + amenityHostGreeting + 
                  amenitySmoking + amenityPets + amenityWifi + amenityTV,
                  data = listingsRating, R = 2000)
Run Code Online (Sandbox Code Playgroud)

错误:

Error in t.star[r, ] <- res[[r]] : 
number of items to replace is not a multiple of replacement length
In addition: Warning message:
In doTryCatch(return(expr), name, parentenv, handler) :
restarting interrupted promise evaluation
Run Code Online (Sandbox Code Playgroud)

duc*_*ayr 5

问题

问题是你的因子变量。当您lm()对数据的一个子集进行操作时(在 中一遍又一遍地完成boot::boot()),您只能获得现有因子水平的系数。然后每个系数绘制可以是不同的长度。如果你这样做,这可以被复制

debug(boot)
set.seed(123)
bootResults <- boot(statistic = bootReg, formula = review_scores_rating ~ instant_bookable + cancellation_policy + 
                        host_since_cat + host_location_cat + host_response_time + 
                        host_is_superhost + host_listings_cat + property_type + room_type + 
                        accommodates + bedrooms + beds + price + security_deposit + 
                        cleaning_fee + extra_people + minimum_nights + amenityBreakfast + 
                        amenityAC + amenityElevator + amenityKitchen + amenityHostGreeting + 
                        amenitySmoking + amenityPets + amenityWifi + amenityTV,
                    data = listingsRating, R = 2)
Run Code Online (Sandbox Code Playgroud)

这将允许您一次一行地浏览函数调用。跑完这条线后

res <- if (ncpus > 1L && (have_mc || have_snow)) {
    if (have_mc) {
        parallel::mclapply(seq_len(RR), fn, mc.cores = ncpus)
    }
    else if (have_snow) {
        list(...)
        if (is.null(cl)) {
            cl <- parallel::makePSOCKcluster(rep("localhost", 
                ncpus))
            if (RNGkind()[1L] == "L'Ecuyer-CMRG") 
                parallel::clusterSetRNGStream(cl)
            res <- parallel::parLapply(cl, seq_len(RR), fn)
            parallel::stopCluster(cl)
            res
        }
        else parallel::parLapply(cl, seq_len(RR), fn)
    }
} else lapply(seq_len(RR), fn)
Run Code Online (Sandbox Code Playgroud)

然后试试

setdiff(names(res[[1]]), names(res[[2]]))
# [1] "property_typeBarn"         "property_typeNature lodge"
Run Code Online (Sandbox Code Playgroud)

第一个子集中存在两个因子水平,第二个子集中不存在。这导致了您的问题。

解决方案

用于model.matrix()事先扩展您的因素(在此 Stack Overflow 帖子之后):

df2 <- model.matrix( ~ review_scores_rating + instant_bookable + cancellation_policy + 
                        host_since_cat + host_location_cat + host_response_time + 
                        host_is_superhost + host_listings_cat + property_type + room_type + 
                        accommodates + bedrooms + beds + price + security_deposit + 
                        cleaning_fee + extra_people + minimum_nights + amenityBreakfast + 
                        amenityAC + amenityElevator + amenityKitchen + amenityHostGreeting + 
                        amenitySmoking + amenityPets + amenityWifi + amenityTV - 1, data = listingsRating)
undebug(boot)

set.seed(123)
bootResults <- boot(statistic = bootReg, formula = review_scores_rating ~ .,
                    data = as.data.frame(df2), R = 2)
Run Code Online (Sandbox Code Playgroud)

(请注意,R为了在调试期间更快的运行时间,我一直将其减少到 2)。