我正在使用Apache Spark来使用MLib提供的LogisticRegressionWithLBFGS()类来构建LRM.构建模型后,我们可以使用提供的预测函数,它只给出二进制标签作为输出.我也希望计算相同的概率.
有一个相同的实现
override protected def predictPoint(
dataMatrix: Vector,
weightMatrix: Vector,
intercept: Double) = {
require(dataMatrix.size == numFeatures)
// If dataMatrix and weightMatrix have the same dimension, it's binary logistic regression.
if (numClasses == 2) {
val margin = dot(weightMatrix, dataMatrix) + intercept
val score = 1.0 / (1.0 + math.exp(-margin))
threshold match {
case Some(t) => if (score > t) 1.0 else 0.0
case None => score
}
}
Run Code Online (Sandbox Code Playgroud)
此方法未公开,并且概率不可用.我可以知道如何使用此函数来获取概率.在上述功能中使用的点方法也没有暴露,它存在于BLAS包中但不公开.
我正在使用Spark/Scala,我想在我的DataFrame中使用基于列类型的默认值填充空值.
ie String Columns - >"string",Numeric Columns - > 111,Boolean Columns - > False等.
目前DF.na.functions API提供na.fill之
fill(valueMap: Map[String, Any])类的
df.na.fill(Map(
"A" -> "unknown",
"B" -> 1.0
))
Run Code Online (Sandbox Code Playgroud)
这需要知道列名称以及列的类型.
要么
fill(value: String, cols: Seq[String])
Run Code Online (Sandbox Code Playgroud)
这只是String/Double类型,甚至不是布尔值.
有一种聪明的方法吗?