在Keras中实现批次依赖性损失

dun*_*r94 6 python keras tensorflow loss-function

我在Keras中设置了自动编码器。我希望能够根据预定的“精度”向量加权输入向量的特征。此连续值向量的长度与输入的长度相同,并且每个元素都位于范围内[0, 1],与相应输入元素的置信度相对应,其中1是完全置信度,0是无置信度。

对于每个示例,我都有一个精确向量。

我定义了一种损失,其中考虑了该精度向量。在这里,低置信度特征的重构权重降低。

def MAEpw_wrapper(y_prec):
    def MAEpw(y_true, y_pred):
        return K.mean(K.square(y_prec * (y_pred - y_true)))
    return MAEpw
Run Code Online (Sandbox Code Playgroud)

我的问题是精度张量y_prec取决于批次。我希望能够y_prec根据当前批次进行更新,以使每个精度向量都与其观察值正确关联。

我做了以下工作:

global y_prec
y_prec = K.variable(P[:32])
Run Code Online (Sandbox Code Playgroud)

这P是一个包含所有精度向量的numpy数组,其索引与示例相对应。我初始化y_prec为批次大小为32的正确形状。然后定义以下内容DataGenerator:

class DataGenerator(Sequence):

    def __init__(self, batch_size, y, shuffle=True):

        self.batch_size = batch_size

        self.y = y

        self.shuffle = shuffle
        self.on_epoch_end()

    def on_epoch_end(self):
        self.indexes = np.arange(len(self.y))

        if self.shuffle == True:
            np.random.shuffle(self.indexes)

    def __len__(self):
        return int(np.floor(len(self.y) / self.batch_size))

    def __getitem__(self, index):
        indexes = self.indexes[index * self.batch_size: (index+1) * self.batch_size]

        # Set precision vector.
        global y_prec
        new_y_prec = K.variable(P[indexes])
        y_prec = K.update(y_prec, new_y_prec)

        # Get training examples.
        y = self.y[indexes]

        return y, y
Run Code Online (Sandbox Code Playgroud)

在这里,我旨在更新y_prec生成批处理的功能。这似乎正在y_prec按预期更新。然后定义我的模型架构:

dims = [40, 20, 2]

model2 = Sequential()
model2.add(Dense(dims[0], input_dim=64, activation='relu'))
model2.add(Dense(dims[1], input_dim=dims[0], activation='relu'))
model2.add(Dense(dims[2], input_dim=dims[1], activation='relu', name='bottleneck'))
model2.add(Dense(dims[1], input_dim=dims[2], activation='relu'))
model2.add(Dense(dims[0], input_dim=dims[1], activation='relu'))
model2.add(Dense(64, input_dim=dims[0], activation='linear'))
Run Code Online (Sandbox Code Playgroud)

最后,我编译并运行:

model2.compile(optimizer='adam', loss=MAEpw_wrapper(y_prec))
model2.fit_generator(DataGenerator(32, digits.data), epochs=100)
Run Code Online (Sandbox Code Playgroud)

digits.datanumpy个观察值数组在哪里。

但是,这最终定义了单独的图:

StopIteration: Tensor("Variable:0", shape=(32, 64), dtype=float32_ref) must be from the same graph as Tensor("Variable_4:0", shape=(32, 64), dtype=float32_ref).
Run Code Online (Sandbox Code Playgroud)

我一直在寻找SO解决我的问题的方法,但没有发现任何可行的方法。感谢您提供任何有关如何正确执行此操作的帮助。

rvi*_*nas 1

该自动编码器可以使用Keras 功能 API轻松实现。这将允许有一个额外的输入占位符y_prec_input,它将被输入“精度”向量。完整的源代码可以在这里找到。


数据生成器

首先,让我们重新实现数据生成器,如下所示:

class DataGenerator(Sequence):
    def __init__(self, batch_size, y, prec, shuffle=True):
        self.batch_size = batch_size
        self.y = y
        self.shuffle = shuffle
        self.prec = prec
        self.on_epoch_end()

    def on_epoch_end(self):
        self.indexes = np.arange(len(self.y))
        if self.shuffle:
            np.random.shuffle(self.indexes)

    def __len__(self):
        return int(np.floor(len(self.y) / self.batch_size))

    def __getitem__(self, index):
        indexes = self.indexes[index * self.batch_size: (index + 1) * self.batch_size]
        y = self.y[indexes]
        y_prec = self.prec[indexes]
        return [y, y_prec], y
Run Code Online (Sandbox Code Playgroud)

请注意,我删除了全局变量。现在,精度向量P作为输入参数 ( prec) 提供,并且生成器生成一个附加输入,该输入将馈送到精度占位符y_prec_input(请参阅模型定义)。


模型

最后,您的模型可以按如下方式定义和训练:

y_input = Input(shape=(input_dim,))
y_prec_input = Input(shape=(1,))
h_enc = Dense(dims[0], activation='relu')(y_input)
h_enc = Dense(dims[1], activation='relu')(h_enc)
h_enc = Dense(dims[2], activation='relu', name='bottleneck')(h_enc)
h_dec = Dense(dims[1], activation='relu')(h_enc)
h_dec = Dense(input_dim, activation='relu')(h_dec)
model2 = Model(inputs=[y_input, y_prec_input], outputs=h_dec)
model2.compile(optimizer='adam', loss=MAEpw_wrapper(y_prec_input))

# Train model
model2.fit_generator(DataGenerator(32, digits.data, P), epochs=100)
Run Code Online (Sandbox Code Playgroud)

在哪里input_dim = digits.data.shape[1]。请注意,我还将解码器的输出维度更改为input_dim,因为它必须与输入维度匹配。