KFolds Cross Validation vs train_test_split

Dav*_*vid 6 python machine-learning scikit-learn cross-validation

I just built my first random forest classifier today and I am trying to improve its performance. I was reading about how cross-validation is important to avoid overfitting of data and hence obtain better results. I implemented StratifiedKFold using sklearn, however, surprisingly this approach resulted to be less accurate. I have read numerous posts suggesting that cross-validating is much more efficient than train_test_split.

Estimator:

rf = RandomForestClassifier(n_estimators=100, random_state=42)
Run Code Online (Sandbox Code Playgroud)

K-Fold:

ss = StratifiedKFold(n_splits=10, shuffle=True, random_state=42)
for train_index, test_index in ss.split(features, labels):
    train_features, test_features = features[train_index], features[test_index]
    train_labels, test_labels = labels[train_index], labels[test_index]
Run Code Online (Sandbox Code Playgroud)

TTS:

train_feature, test_feature, train_label, test_label = \
    train_test_split(features, labels, train_size=0.8, test_size=0.2, random_state=42)
Run Code Online (Sandbox Code Playgroud)

Below are results:

CV:

AUROC:  0.74
Accuracy Score:  74.74 %.
Specificity:  0.69
Precision:  0.75
Sensitivity:  0.79
Matthews correlation coefficient (MCC):  0.49
F1 Score:  0.77
Run Code Online (Sandbox Code Playgroud)

TTS:

AUROC:  0.76
Accuracy Score:  76.23 %.
Specificity:  0.77
Precision:  0.79
Sensitivity:  0.76
Matthews correlation coefficient (MCC):  0.52
F1 Score:  0.77
Run Code Online (Sandbox Code Playgroud)

Is this actually possible? Or have I wrongly set up my models?

Also, is this the correct way of using cross-validation?

Cle*_*ard 7

很高兴看到你记录了自己!

\n\n

造成这种差异的原因是 TTS 方法引入了偏差(因为您没有使用所有观察结果进行测试),这解释了这种差异。

\n\n
\n

在验证方法中,仅使用训练集而不是验证集\xe2\x80\x94 中包含的观测值\xe2\x80\x94 的子集\n 来拟合模型。由于统计方法在对较少观测值进行训练时往往表现较差,这表明验证集错误率可能会高估模型在整个数据集上的拟合测试错误率。

\n
\n\n

结果可能相差很大:

\n\n
\n

测试错误率的验证估计可能存在很大差异,具体取决于训练集中包含哪些观测值以及验证集中包含哪些观测值

\n
\n\n

交叉验证通过使用所有可用数据来解决这个问题,从而消除偏差。

\n\n

此处,TTS 方法的结果存在更多偏差,在分析结果时应牢记这一点。也许您在采样的测试/验证集上也很幸运

\n\n

再次,这里有一篇很棒的、适合初学者的文章,详细介绍了该主题:\n https://codesachin.wordpress.com/2015/08/30/cross-validation-and-the-bias-variance-tradeoff-for-dummies /

\n\n

更深入的来源,请参考“模型评估与选择”\n这里的章节(引用内容来源):

\n\n

https://web.stanford.edu/~hastie/Papers/ESLII.pdf

\n

  • 是的,您将数据分成 K 个等于集,然后在 K-1 集上进行训练并在其余集上进行测试。你这样做了 K 次,每次都改变测试集,这样最终每组都将成为测试集一次,并成为训练集 K-1 次。然后对 K 个结果进行平均以获得 K 重 CV 结果 (2认同)