Julia 中的随机森林和 ROC 曲线

lar*_*ara 5 machine-learning decision-tree random-forest roc julia

我正在使用DecisionTree.jl包的 ScikitLearn 风格为RDatasets数据集之一的二元分类问题创建随机森林模型(有关 ScikitLearn 风格的含义,请参阅 DecisionTree.jl 主页的底部)。我还使用MLBase包进行模型评估。

我已经为我的数据构建了一个随机森林模型,并想为这个模型创建一个 ROC 曲线。阅读可用的文档,我确实理解理论上的 ROC 曲线是什么。我只是不知道如何为特定模型创建一个。

从维基百科页面,我在下面用粗斜体标记的第一句话的最后一部分引起了我的困惑:“在统计学中,接受者操作特征 (ROC) 或 ROC 曲线是一个图形图,说明二元分类器系统的性能作为其区分阈值是变化的。” 整篇文章中有更多关于阈值的内容,但这仍然让我对二元分类问题感到困惑。什么是阈值以及如何改变它?

此外,在关于 ROC 曲线的MLBase 文档中,它说“根据给定的分数和阈值计算 ROC 实例或 ROC 曲线(ROC 实例的向量)。” 但实际上并没有在其他任何地方提到这个阈值。

下面给出了我的项目的示例代码。基本上,我想为随机森林创建一个 ROC 曲线,但我不确定如何或者它是否合适。

using DecisionTree
using RDatasets
using MLBase

quakes_data = dataset("datasets", "quakes");

# Add in a binary column as feature column for classification
quakes_data[:MagGT5] = convert(Array{Int32,1}, quakes_data[:Mag] .> 5.0)

# Getting features and labels where label = 1 is mag > 1 and label = 2 is mag <= 5
features = convert(Array, quakes_data[:, [1:3;5]]);
labels = convert(Array, quakes_data[:, 6]);
labels[labels.==0] = 2

# Create a random forest model with the tuning parameters I want
r_f_model = RandomForestClassifier(nsubfeatures = 3, ntrees = 50, partialsampling=0.7, maxdepth = 4)

# Train the model in-place on the dataset (there isn't a fit function without the in-place functionality)
DecisionTree.fit!(r_f_model, features, labels)

# Apply the trained model to the test features data set (here I haven't partitioned into training and test)
r_f_prediction = convert(Array{Int64,1}, DecisionTree.predict(r_f_model, features))

# Applying the model to the training set and looking at model stats
TrainingROC = roc(labels, r_f_prediction) #getting the stats around the model applied to the train set
#     p::T    # positive in ground-truth
#     n::T    # negative in ground-truth
#     tp::T   # correct positive prediction
#     tn::T   # correct negative prediction
#     fp::T   # (incorrect) positive prediction when ground-truth is negative
#     fn::T   # (incorrect) negative prediction when ground-truth is positive
Run Code Online (Sandbox Code Playgroud)

我也阅读了这个问题,并没有发现它真的有帮助。

Dan*_*etz 4

二元分类的任务是为新的、未标记的数据点提供0/ 1(或true/ false、red/ )标签。blue大多数分类算法旨在输出连续的实值。对于具有已知或预测标签的点,该值被优化为较高1,对于具有已知或预测标签的点,该值较低0。为了使用该值生成0/预测,需要使用1附加阈值。值高于阈值的点被预测为被标记1(对于低于阈值的点0被预测为标签)。

为什么这个设置有用?因为,有时错误预测 a0而不是 a 的1成本更高,然后您可以将阈值设置得较低,使算法输出预测1s 的频率更高。

在极端情况下,当预测0而不是1应用程序不需要任何成本时,您可以将阈值设置为无穷大,使其始终输出0(这显然是最好的解决方案,因为它不产生任何成本)。

阈值技巧无法消除分类器中的错误 - 现实世界问题中没有任何分类器是完美的或没有噪声的。它可以做的是改变最终分类的0“when-really-”1错误和1“-when-really-”错误之间的比率。0

随着阈值的增加,更多的点将被分类为标签0。考虑一个图表,其中 x 轴上分类的点的分数,以及y 轴上具有“真正时”错误0的点的分数。对于每个阈值,在此图表上为生成的分类器绘制一个点。为所有阈值绘制一个点,您将得到一条曲线。这是 ROC 曲线(的某种变体),它总结了分类器的能力。分类质量的常用指标是该图表的 AUC 或曲线下面积,但事实上,整个曲线在应用中可能会引起兴趣。01

像这样的总结出现在许多关于机器学习的文本中,只需谷歌查询即可。

希望这能澄清阈值的作用及其与 ROC 曲线的关系。