我试图复制使用Tensorflow 在优秀文章http://karpathy.github.io/2015/05/21/rnn-effectiveness/中演示的字符级语言建模.
到目前为止,我的尝试失败了.我的网络通常在处理800个左右的字符后输出单个字符.我相信我已经从根本上误解了张量流实现LSTM的方式,也许还有rnns.我发现难以遵循的文档.
这是我的代码的本质:
图形定义
idata = tf.placeholder(tf.int32,[None,1]) #input byte, use value 256 for start and end of file
odata = tf.placeholder(tf.int32,[None,1]) #target output byte, ie, next byte in sequence..
source = tf.to_float(tf.one_hot(idata,257)) #input byte as 1-hot float
target = tf.to_float(tf.one_hot(odata,257)) #target output as 1-hot float
with tf.variable_scope("lstm01"):
cell1 = tf.nn.rnn_cell.BasicLSTMCell(257)
val1, state1 = tf.nn.dynamic_rnn(cell1, source, dtype=tf.float32)
output = val1
Run Code Online (Sandbox Code Playgroud)
损失计算
cross_entropy = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(output, target))
train_step = tf.train.AdamOptimizer(1e-4).minimize(cross_entropy)
output_am = tf.argmax(output,2)
target_am = tf.argmax(target,2)
correct_prediction = tf.equal(output_am, target_am) …Run Code Online (Sandbox Code Playgroud)