为什么我无法使用coefficients_sgd方法获得sklearn LogisticRegression得到的结果?

6 python iteration gradient-descent scikit-learn

from math import exp
import numpy as np
from sklearn.linear_model import LogisticRegression
Run Code Online (Sandbox Code Playgroud)

我使用了下面的代码来自 How To Implement Logistic Regression From Scratch in Python

def predict(row, coefficients):
    yhat = coefficients[0]
    for i in range(len(row)-1):
        yhat += coefficients[i + 1] * row[i]
    return 1.0 / (1.0 + exp(-yhat))

def coefficients_sgd(train, l_rate, n_epoch):
    coef = [0.0 for i in range(len(train[0]))]
    for epoch in range(n_epoch):
        sum_error = 0
        for row in train:
            yhat = predict(row, coef)
            error = row[-1] - yhat
            sum_error += error**2
            coef[0] = coef[0] + l_rate * error * yhat * (1.0 - yhat)
            for i in range(len(row)-1):
                coef[i + 1] = coef[i + 1] + l_rate * error * yhat * (1.0 - yhat) * row[i]
    return coef

dataset = [[2.7810836,2.550537003,0],
[1.465489372,2.362125076,0],
[3.396561688,4.400293529,0],
[1.38807019,1.850220317,0],
[3.06407232,3.005305973,0],
[7.627531214,2.759262235,1],
[5.332441248,2.088626775,1],
[6.922596716,1.77106367,1],
[8.675418651,-0.242068655,1],
[7.673756466,3.508563011,1]]

l_rate = 0.3
n_epoch = 100
coef = coefficients_sgd(dataset, l_rate, n_epoch)
print(coef)
Run Code Online (Sandbox Code Playgroud)

[-0.39233141593823756, 1.4791536027917747, -2.316697087065274]

x = np.array(dataset)[:,:2]
y = np.array(dataset)[:,2]
model = LogisticRegression(penalty="none")
model.fit(x,y)
print(model.intercept_.tolist() + model.coef_.ravel().tolist())
Run Code Online (Sandbox Code Playgroud)

[-3.233238244349982, 6.374828107647225, -9.631487530388092]

我应该改变什么才能获得相同或更接近的系数?如何建立初始系数、学习率、n_epoch?

小智 7

嗯,这里有很多细微差别

首先,回想一下,可以使用各种优化方法(包括您实现的 SGD)来估计具有(负)对数似然的逻辑回归系数,但没有精确的封闭式解决方案。因此,即使您实现了 scikit-learn 的精确副本LogisticRegression,您也需要设置相同的超参数(时期数、学习率等)和随机状态以获得相同的系数。

其次,LogisticRegression提供了五种不同的优化方法(solver参数)。您LogisticRegression(penalty="none")使用其默认参数运行,默认值为solver,'lbfgs'而不是 SGD;因此,根据您的数据和超参数,您可能会得到明显不同的结果。

我应该改变什么才能获得相同或更接近的系数?

我建议将您的实现与SGDClassifier(loss='log')第一个进行比较,因为LogisticRegression不提供 SGD 求解器。尽管请记住 scikit-learn 的实现更加复杂,特别是具有更多用于提前停止的超参数,例如tol.

如何建立初始系数、学习率、n_epoch?

uniform(-1/(2n), 1/(2n))通常,SGD 的系数是使用一些数据统计数据(例如,dot(y, w)/(dot(w, w)对于每个系数w)或使用预先训练的模型参数来随机初始化的(例如, ) 。相反,学习率或轮数没有黄金法则。通常,我们设置大量的 epoch 和其他一些停止标准(例如,当前系数和先前系数之间的范数是否小于某个较小的tol)、适度的学习率,并且每次迭代我们都会遵循某种规则降低学习率(参见learning_rate参数或用户指南SGDClassifier)并检查停止标准。