对此问题的更准确描述是,当is_training未明确设置为true时,MobileNet会表现不佳.我指的是TensorFlow在其模型库https://github.com/tensorflow/models/blob/master/slim/nets/mobilenet_v1.py中提供的MobileNet .
这就是我创建网络的方式(phase_train = True):
with slim.arg_scope(mobilenet_v1.mobilenet_v1_arg_scope(is_training=phase_train)):
features, endpoints = mobilenet_v1.mobilenet_v1(
inputs=images_placeholder, features_layer_size=features_layer_size, dropout_keep_prob=dropout_keep_prob,
is_training=phase_train)
Run Code Online (Sandbox Code Playgroud)
我正在训练一个识别网络,在训练时我在LFW上进行测试.我在培训期间获得的结果随着时间的推移而得到改善并获得了良好的准确性.
在部署之前,我冻结了图表.如果我使用is_training = True冻结图形,我在LFW上获得的结果与训练期间相同.但是,如果我设置is_training = False,我得到的结果就像网络根本没有训练过......
这种行为实际上发生在像Inception这样的其他网络上.
我倾向于认为我错过了一些非常基本的东西,这不是TensorFlow中的错误...
任何帮助,将不胜感激.
添加更多代码......
这就是我准备培训的方式:
images_placeholder = tf.placeholder(tf.float32, shape=(None, image_size, image_size, 1), name='input')
labels_placeholder = tf.placeholder(tf.int32, shape=(None))
dropout_placeholder = tf.placeholder_with_default(1.0, shape=(), name='dropout_keep_prob')
phase_train_placeholder = tf.Variable(True, name='phase_train')
global_step = tf.Variable(0, name='global_step', trainable=False)
# build graph
with slim.arg_scope(mobilenet_v1.mobilenet_v1_arg_scope(is_training=phase_train_placeholder)):
features, endpoints = mobilenet_v1.mobilenet_v1(
inputs=images_placeholder, features_layer_size=512, dropout_keep_prob=1.0,
is_training=phase_train_placeholder)
# loss
logits = slim.fully_connected(inputs=features, num_outputs=train_data.get_class_count(), activation_fn=None,
weights_initializer=tf.truncated_normal_initializer(stddev=0.1),
weights_regularizer=slim.l2_regularizer(scale=0.00005),
scope='Logits', reuse=False)
tf.losses.sparse_softmax_cross_entropy(labels=labels_placeholder, logits=logits, …Run Code Online (Sandbox Code Playgroud) 我正在使用 BatchNorm 层。use_global_stats我知道通常false为训练和测试/部署而设置的设置的含义true。这是我在测试阶段的设置。
layer {
name: "bnorm1"
type: "BatchNorm"
bottom: "conv1"
top: "bnorm1"
batch_norm_param {
use_global_stats: true
}
}
layer {
name: "scale1"
type: "Scale"
bottom: "bnorm1"
top: "bnorm1"
bias_term: true
scale_param {
filler {
value: 1
}
bias_filler {
value: 0.0
}
}
}
Run Code Online (Sandbox Code Playgroud)
在solver.prototxt中,我使用了Adam方法。我发现我的案例中出现了一个有趣的问题。如果我选择,那么当我在测试阶段base_lr: 1e-3设置时,我会得到很好的性能。use_global_stats: false然而,如果我选择,那么当我在测试阶段base_lr: 1e-4设置时,我会得到很好的性能。use_global_stats: true它证明了base_lr对批规范设置的影响(即使我使用了 Adam 方法)?你能提出任何理由吗?谢谢大家