使用不平衡数据构建 ML 分类器

Mat*_*ath 5 python machine-learning resampling scikit-learn cross-validation

我有一个包含 1400 个 obs 和 19 列的数据集。Target 变量的值为 1(我最感兴趣的值)和 0。类别的分布显示不平衡 (70:30)。

使用下面的代码我得到了奇怪的值(全 1)。我不知道这是由于数据过度拟合/不平衡问题还是由于特征选择问题(我使用皮尔逊相关性,因为所有值都是数字/布尔值)。我认为接下来的步骤是错误的。

import numpy as np
import math
import sklearn.metrics as metrics
from sklearn.metrics import f1_score

y = df['Label']
X = df.drop('Label',axis=1)

def create_cv(X,y):
    if type(X)!=np.ndarray:
        X=X.values
        y=y.values
 
    test_size=1/5
    proportion_of_true=y[y==1].shape[0]/y.shape[0]
    num_test_samples=math.ceil(y.shape[0]*test_size)
    num_test_true_labels=math.floor(num_test_samples*proportion_of_true)
    num_test_false_labels=math.floor(num_test_samples-num_test_true_labels)
    
    y_test=np.concatenate([y[y==0][:num_test_false_labels],y[y==1][:num_test_true_labels]])
    y_train=np.concatenate([y[y==0][num_test_false_labels:],y[y==1][num_test_true_labels:]])

    X_test=np.concatenate([X[y==0][:num_test_false_labels] ,X[y==1][:num_test_true_labels]],axis=0)
    X_train=np.concatenate([X[y==0][num_test_false_labels:],X[y==1][num_test_true_labels:]],axis=0)
    return X_train,X_test,y_train,y_test

X_train,X_test,y_train,y_test=create_cv(X,y)
X_train,X_crossv,y_train,y_crossv=create_cv(X_train,y_train)
    
tree = DecisionTreeClassifier(max_depth = 5)
tree.fit(X_train, y_train)       

y_predict_test = tree.predict(X_test)

print(classification_report(y_test, y_predict_test))
f1_score(y_test, y_predict_test)
Run Code Online (Sandbox Code Playgroud)

输出:

     precision    recall  f1-score   support

           0       1.00      1.00      1.00        24
           1       1.00      1.00      1.00        70

    accuracy                           1.00        94
   macro avg       1.00      1.00      1.00        94
weighted avg       1.00      1.00      1.00        94
Run Code Online (Sandbox Code Playgroud)

当数据不平衡时,使用 CV 和/或欠采样构建分类器时,是否有人遇到过类似的问题?很高兴分享整个数据集,以防您可能想要复制输出。我想请您提供一些明确的答案,以告诉我步骤以及我做错了什么。

我知道,为了减少过度拟合并处理平衡数据,有一些方法,例如随机采样(过度/不足)、SMOTE、CV。我的想法是

  • 考虑到不平衡,分割训练/测试数据
  • 在列车组上执行 CV
  • 仅在测试折叠上应用欠采样
  • 在 CV 的帮助下选择模型后,对训练集进行欠采样并训练分类器
  • 估计未受影响的测试集上的性能(f1-score)

正如这个问题中所概述的:CV 和测试折叠上的采样不足。

我认为上述步骤应该有意义,但很高兴收到您对此的任何反馈。

Jua*_*rez 3

当您的数据不平衡时,您必须执行分层。通常的方法是对具有较少值的类进行过采样。

另一种选择是用更少的数据来训练你的算法。如果你有一个好的数据集,那应该不是问题。在这种情况下,您首先从代表性较少的类中获取样本,然后使用集合的大小来计算从其他类中获取的样本数量:

此代码可以帮助您以这种方式分割数据集:

def split_dataset(dataset: pd.DataFrame, train_share=0.8):
    """Splits the dataset into training and test sets"""
    all_idx = range(len(dataset))
    train_count = int(len(all_idx) * train_share)

    train_idx = random.sample(all_idx, train_count)
    test_idx = list(set(all_idx).difference(set(train_idx)))

    train = dataset.iloc[train_idx]
    test = dataset.iloc[test_idx]

    return train, test

def split_dataset_stratified(dataset, target_attr, positive_class, train_share=0.8):
    """Splits the dataset as in `split_dataset` but with stratification"""

    data_pos = dataset[dataset[target_attr] == positive_class]
    data_neg = dataset[dataset[target_attr] != positive_class]

    if len(data_pos) < len(data_neg):
        train_pos, test_pos = split_dataset(data_pos, train_share)
        train_neg, test_neg = split_dataset(data_neg, len(train_pos)/len(data_neg))
        # set.difference makes the test set larger
        test_neg = test_neg.iloc[0:len(test_pos)]
    else:
        train_neg, test_neg = split_dataset(data_neg, train_share)
        train_pos, test_pos = split_dataset(data_pos, len(train_neg)/len(data_pos))
        # set.difference makes the test set larger
        test_pos = test_pos.iloc[0:len(test_neg)]

    return train_pos.append(train_neg).sample(frac = 1).reset_index(drop = True), \
           test_pos.append(test_neg).sample(frac = 1).reset_index(drop = True)
Run Code Online (Sandbox Code Playgroud)

用法:

train_ds, test_ds = split_dataset_stratified(data, target_attr, positive_class)
Run Code Online (Sandbox Code Playgroud)

您现在可以在 中执行交叉验证train_ds并评估您的模型test_ds。