xgboost软件包和随机森林回归

M.V*_*.V. 5 r random-forest xgboost

xgboost软件包允许构建随机森林(实际上,它选择列的随机子集来选择用于整棵树拆分的变量,而不是点头算法,这与传统算法的版本一样,但是它可以容忍)。但是似乎为了进行回归,只使用了森林中的一棵树(也许是最后一棵树)。

为了确保这一点,请考虑一个标准的玩具示例。

library(xgboost)
library(randomForest)
data(agaricus.train, package = 'xgboost')
    dtrain = xgb.DMatrix(agaricus.train$data,
 label = agaricus.train$label)
 bst = xgb.train(data = dtrain, 
                 nround = 1, 
                 subsample = 0.8, 
                 colsample_bytree = 0.5, 
                 num_parallel_tree = 100, 
                 verbose = 2, 
                 max_depth = 12)

answer1 = predict(bst, dtrain); 
(answer1 - agaricus.train$label) %*% (answer1 -  agaricus.train$label)

forest = randomForest(x = as.matrix(agaricus.train$data), y = agaricus.train$label, ntree = 50)

answer2 = predict(forest, as.matrix(agaricus.train$data))
(answer2 - agaricus.train$label) %*% (answer2 -  agaricus.train$label)
Run Code Online (Sandbox Code Playgroud)

是的,当然,默认版本的xgboost随机森林不使用基尼分数函数,而仅使用MSE;它可以轻松更改。同样,进行这样的验证也是不正确的,依此类推。它不会影响主要问题。不管尝试哪种参数集,与randomForest实现相比,结果都是令人惊讶的糟糕。这也适用于另一个数据集。

有人能暗示这种奇怪的行为吗?当涉及分类任务时,该算法可以按预期工作。

#

好吧,所有的树都长出来了,都用来做预测。您可以通过对“预测”功能使用参数“ ntreelimit”进行检查。

主要问题仍然存在:由xgbbost软件包产生的随机森林算法的特定形式有效吗?

交叉验证,参数调整和其他废话与之无关—每个人都可以对代码进行必要的更正,然后看看会发生什么。

您可以像这样指定“目标”选项:

mse = function(predict, dtrain)
{
  real = getinfo(dtrain, 'label')
  return(list(grad = 2 * (predict - real),
              hess = rep(2, length(real))))
}
Run Code Online (Sandbox Code Playgroud)

这提供了在选择拆分变量时使用MSE的功能。即使在那之后,与randomForest相比,结果仍然令人吃惊。

也许,该问题是学术性质的问题,涉及如何选择特征的随机子集进行分割的方式。经典实现为每个拆分分别选择功能的子集(对于randomForest包,其大小由'mtry'指定),而xgboost实现为树(由'colsample_bytree'指定)选择一个子集。

因此,至少对于某些类型的数据集而言,这种细微的差异显得非常重要。确实很有趣。

Sor*_*ing 5

xgboost(随机森林样式)确实使用了不止一棵树来进行预测。但是还有许多其他差异需要探索。

我本人是xgboost的新手,但很好奇。所以我写了下面的代码来可视化树木。您可以自己运行代码以验证或探索其他差异。

您选择的数据集是一个分类问题,因为标签为0或1。我想切换到一个简单的回归问题以可视化xgboost的作用。

真实模型:$ y = x_1 * x_2 $ +噪声

如果训练一棵树或多棵树,在下面的代码示例中,您会发现学习的模型结构确实包含更多的树。您不能从预测准确性中单独争论训练了多少棵树。

预测可能不同,因为实现方式不同。我所知道的〜5个RF实现中,没有一个是完全相似的,并且这个xgboost(rf样式)最接近一个遥远的“表兄弟”。

我观察到colsample_bytree不等于mtry,因为前者对整个树使用相同的变量/ 列子集。我的回归问题只是一个大的交互作用,如果树仅使用x1或x2则无法学习。因此,在这种情况下,必须将colsample_bytree设置为1才能在所有树中使用两个变量。常规的RF可以使用mtry = 1对这个问题进行建模,因为每个节点将使用X1或X2

我看到您的randomForest预测没有经过袋式交叉验证。如果对预测得出任何结论,则必须进行交叉验证,尤其是对于完全生长的树木。

NB You need to fix the function vec.plot as does not support xgboost out of the box, because xgboost out of some other box do not take data.frame as an valid input. The instruction in the code should be clear

library(xgboost)
library(rgl)
library(forestFloor)
Data = data.frame(replicate(2,rnorm(5000)))
Data$y = Data$X1*Data$X2 + rnorm(5000)*.5
gradientByTarget =fcol(Data,3)
plot3d(Data,col=gradientByTarget) #true data structure

fix(vec.plot) #change these two line in the function, as xgboost do not support data.frame
#16# yhat.vec = predict(model, as.matrix(Xtest.vec))
#21# yhat.obs = predict(model, as.matrix(Xtest.obs))

#1 single deep tree
xgb.model =  xgboost(data = as.matrix(Data[,1:2]),label=Data$y,
                     nrounds=1,params = list(max.depth=250))
vec.plot(xgb.model,as.matrix(Data[,1:2]),1:2,col=gradientByTarget,grid=200)
plot(Data$y,predict(xgb.model,as.matrix(Data[,1:2])),col=gradientByTarget)
#clearly just one tree

#100 trees (gbm boosting)
xgb.model =  xgboost(data = as.matrix(Data[,1:2]),label=Data$y,
                     nrounds=100,params = list(max.depth=16,eta=.5,subsample=.6))
vec.plot(xgb.model,as.matrix(Data[,1:2]),1:2,col=gradientByTarget) 
plot(Data$y,predict(xgb.model,as.matrix(Data[,1:2])),col=gradientByTarget) ##predictions are not OOB cross-validated!


#20 shallow trees (bagging)
xgb.model =  xgboost(data = as.matrix(Data[,1:2]),label=Data$y,
                     nrounds=1,params = list(max.depth=250,
                     num_parallel_tree=20,colsample_bytree = .5, subsample = .5))
vec.plot(xgb.model,as.matrix(Data[,1:2]),1:2,col=gradientByTarget) #bagged mix of trees
plot(Data$y,predict(xgb.model,as.matrix(Data[,1:2]))) #terrible fit!!
#problem, colsample_bytree is NOT mtry as columns are only sampled once
# (this could be raised as an issue on their github page, that this does not mimic RF)


#20 deep tree (bagging), no column limitation
xgb.model =  xgboost(data = as.matrix(Data[,1:2]),label=Data$y,
                     nrounds=1,params = list(max.depth=500,
                     num_parallel_tree=200,colsample_bytree = 1, subsample = .5))
vec.plot(xgb.model,as.matrix(Data[,1:2]),1:2,col=gradientByTarget) #boosted mix of trees
plot(Data$y,predict(xgb.model,as.matrix(Data[,1:2])))
#voila model can fit data
Run Code Online (Sandbox Code Playgroud)