小编cur*_*r23的帖子

Tensorflow LSTM字符按字符序列预测

我试图复制使用Tensorflow 在优秀文章http://karpathy.github.io/2015/05/21/rnn-effectiveness/中演示的字符级语言建模.

到目前为止,我的尝试失败了.我的网络通常在处理800个左右的字符后输出单个字符.我相信我已经从根本上误解了张量流实现LSTM的方式,也许还有rnns.我发现难以遵循的文档.

这是我的代码的本质:

图形定义

idata = tf.placeholder(tf.int32,[None,1])   #input byte, use value 256 for start and end of file
odata = tf.placeholder(tf.int32,[None,1])    #target output byte, ie, next byte in sequence..
source =  tf.to_float(tf.one_hot(idata,257)) #input byte as 1-hot float
target = tf.to_float(tf.one_hot(odata,257))  #target output as 1-hot float

with tf.variable_scope("lstm01"):
    cell1 = tf.nn.rnn_cell.BasicLSTMCell(257)
    val1, state1 = tf.nn.dynamic_rnn(cell1, source, dtype=tf.float32)

output = val1
Run Code Online (Sandbox Code Playgroud)

损失计算

cross_entropy = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(output, target))
train_step = tf.train.AdamOptimizer(1e-4).minimize(cross_entropy)  
output_am = tf.argmax(output,2)
target_am = tf.argmax(target,2)
correct_prediction = tf.equal(output_am, target_am) …
Run Code Online (Sandbox Code Playgroud)

sequences lstm tensorflow

8
推荐指数
1
解决办法
1746
查看次数

标签 统计

lstm ×1

sequences ×1

tensorflow ×1