ncg*_*ody 3 python machine-learning scikit-learn
我正在使用scikit-learn. 我想设计两个二元特征:“旧金山 10 公里内”和“洛杉矶 10 公里内”。我创建了一个自定义转换器,它本身可以正常工作,但是TypeError当我将它放入ColumnTransformer. 这是代码:
from math import radians
from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.compose import ColumnTransformer
from sklearn.metrics.pairwise import haversine_distances
from sklearn.datasets import fetch_california_housing
import numpy as np
import pandas as pd
# Import data into DataFrame
data = fetch_california_housing()
X = pd.DataFrame(data['data'], columns=data['feature_names'])
y = data['target']
# Custom transformer for 'Latitude' and 'Longitude' cols
class NearCity(BaseEstimator, TransformerMixin):
def __init__(self, distance=10):
self.la = (34.05, -118.24)
self.sf = (37.77, -122.41)
self.dis = distance
def calc_dist(self, coords_1, coords_2):
coords_1 = [radians(_) for _ in coords_1]
coords_2 = [radians(_) for _ in coords_2]
result = haversine_distances([coords_1, coords_2])[0,-1]
return result * 6_371
def fit(self, X, y=None):
return self
def transform(self, X):
dist_to_sf = np.apply_along_axis(self.calc_dist, 1, X, coords_2=self.sf)
dist_to_sf = (dist_to_sf < self.dis).astype(int)
dist_to_la = np.apply_along_axis(self.calc_dist, 1, X, coords_2=self.la)
dist_to_la = (dist_to_la < self.dis).astype(int)
X_trans = np.column_stack((X, dist_to_sf, dist_to_la))
return X_trans
ct = ColumnTransformer([('near_city', NearCity(), ['Latitude', 'Longitude'])],
remainder='passthrough')
ct.fit_transform(X)
#> /Users/.../anaconda3/envs/data3/lib/python3.7/site-packages/sklearn/base.py:197: FutureWarning: From version 0.24, get_params will raise an AttributeError if a parameter cannot be retrieved as an instance attribute. Previously it would return None.
#> FutureWarning)
#> Traceback (most recent call last):
#> <ipython-input-13-603f6cd4afd3> in transform(self, X)
#> 17 def transform(self, X):
#> 18 dist_to_sf = np.apply_along_axis(self.calc_dist, 1, X, coords_2=self.sf)
#> ---> 19 dist_to_sf = (dist_to_sf < self.dis).astype(int)
#> 20
#> 21 dist_to_la = np.apply_along_axis(self.calc_dist, 1, X, coords_2=self.la)
#> TypeError: '<' not supported between instances of 'float' and 'NoneType'
Run Code Online (Sandbox Code Playgroud)
由reprexpy 包于 2020 年 4 月 23 日创建
问题是该self.dis属性不持久。如果我本身实例变压器,没问题:self.dis = distance = 10。但在 中ColumnTransformer,它以NoneType. 奇怪的是,如果我只是硬编码self.dis = 10,它就可以工作。
人们认为这是怎么回事?
Session info --------------------------------------------------------------------
Platform: Darwin-18.7.0-x86_64-i386-64bit (64-bit)
Python: 3.7
Date: 2020-04-23
Packages ------------------------------------------------------------------------
numpy==1.18.1
pandas==1.0.1
reprexpy==0.3.0
scikit-learn==0.22.1
Run Code Online (Sandbox Code Playgroud)
原来问题出在sklearn.base.
deep_items = value.get_params().items()
Run Code Online (Sandbox Code Playgroud)
该get_params()函数查看init参数以找出类参数是什么,然后假定它们与内部变量名称相同。
所以我可以通过改变我的init方法来解决这个问题:
def __init__(self, distance=10):
self.la = (34.05, -118.24)
self.sf = (37.77, -122.41)
self.distance = distance # <-- give same name
Run Code Online (Sandbox Code Playgroud)
非常感谢我的一位同事解决了这个问题!
| 归档时间: |
|
| 查看次数: |
1012 次 |
| 最近记录: |