Tensorflow中的多类分类的类别精确度和召回率?

pra*_*wal 11 python classification machine-learning tensorflow

有没有办法在使用张量流进行多类分类时获得每类精度或召回.

例如,如果我有每个批次的y_true和y_pred,如果我有超过2个类,是否有一种功能性的方法来获得精度或每个类的回忆.

LeC*_*Gui 6

2个事实:

  1. 正如其他答案中所述,Tensorflow 内置指标精度和召回率 不支持多类(文档说will be cast to bool)

  2. 有多种方法可以通过使用precision_at_k指定 来获得一对多分数,或者通过简单地将和以正确的方式转换为 。class_idlabelspredictionstf.bool

因为这并不令人满意且不完整,所以我编写了一个用于多类指标tf_metrics的简单包,您可以在github上找到它。它支持多种平均方法,例如.scikit-learn

例子

import tensorflow as tf
import tf_metrics

y_true = [0, 1, 0, 0, 0, 2, 3, 0, 0, 1]
y_pred = [0, 1, 0, 0, 1, 2, 0, 3, 3, 1]
pos_indices = [1]        # Metrics for class 1 -- or
pos_indices = [1, 2, 3]  # Average metrics, 0 is the 'negative' class
num_classes = 4
average = 'micro'

# Tuple of (value, update_op)
precision = tf_metrics.precision(
    y_true, y_pred, num_classes, pos_indices, average=average)
recall = tf_metrics.recall(
    y_true, y_pred, num_classes, pos_indices, average=average)
f2 = tf_metrics.fbeta(
    y_true, y_pred, num_classes, pos_indices, average=average, beta=2)
f1 = tf_metrics.f1(
    y_true, y_pred, num_classes, pos_indices, average=average)
Run Code Online (Sandbox Code Playgroud)


Avi*_*Avi 5

这是一个解决方案,适用于n = 6类的问题.如果你有更多的类,这个解决方案可能很慢,你应该使用某种映射而不是循环.

假设在张量的行中有一个热编码的类标签,在张量中有一个labelslogits(或后验)labels.然后,如果n是类的数量,请尝试这样:

y_true = tf.argmax(labels, 1)
y_pred = tf.argmax(logits, 1)

recall = [0] * n
update_op_rec = [[]] * n

for k in range(n):
    recall[k], update_op_rec[k] = tf.metrics.recall(
        labels=tf.equal(y_true, k),
        predictions=tf.equal(y_pred, k)
    )
Run Code Online (Sandbox Code Playgroud)

请注意,在内部tf.metrics.recall,变量labels和predictions设置为布尔向量,如2变量情况,允许使用该函数.


小智 1

我相信TF还没有提供这样的功能。根据文档(https://www.tensorflow.org/api_docs/python/tf/metrics/ precision),它表示标签和预测都将转换为 bool,因此它仅与二元分类相关。也许可以对示例进行一次性编码并且它会起作用?但对此并不确定。