相关疑难解决方法(0)

在Scikit学习中将Smote与Gridsearchcv一起使用

我正在处理不平衡的数据集,并希望使用scikit的gridsearchcv进行网格搜索以调整模型的参数。为了对数据进行过采样,我想使用SMOTE,我知道我可以将其作为管道的一个阶段,并将其传递给gridsearchcv。我担心的是,我认为训练和验证折纸都将使用击打,这不是您应该做的。验证集不应过采样。我是否正确,整个管道将应用于两个数据集拆分?如果是的话,我该如何扭转呢?提前谢谢

python machine-learning scikit-learn grid-search oversampling

9
推荐指数
1
解决办法
3548
查看次数

使用自定义函数在 sklearn 中创建管道?

如何使用自定义函数创建 sklearn 管道?\n我有两个函数,一个用于清理数据,第二个用于构建模型。

\n\n
def preprocess(df):\n   \xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6.\n   # clean data\n   return df_clean\n\ndef model(df_clean):\n   \xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\xe2\x80\xa6\n   #split data train and test and build randomForest Model\n   return model\n
Run Code Online (Sandbox Code Playgroud)\n\n

所以我使用 FunctionTransformer 并创建了管道

\n\n
from sklearn.pipeline import Pipeline, make_pipeline\nfrom sklearn.preprocessing import FunctionTransformer\n\npipe = Pipeline([("preprocess", FunctionTransformer(preprocess)),("model",FunctionTransformer(model))])\n\npred = pipe.predict_proba(new_test_data)\nprint(pred)\n
Run Code Online (Sandbox Code Playgroud)\n\n

我知道上面是错误的,不知道如何处理,在管道中我需要先传递训练数据,然后我必须传递 new_test_data ?

\n

python pipeline scikit-learn

5
推荐指数
1
解决办法
4278
查看次数