我为文本摘要创建了一个 Seq2Seq 模型。我有两种模型,一种有注意力,一种没有。没有注意力的人能够产生预测,但我不能为有注意力的人做预测,即使它成功地拟合。
这是我的模型:
latent_dim = 300
embedding_dim = 200
clear_session()
# Encoder
encoder_inputs = Input(shape=(max_text_len, ))
# Embedding layer
enc_emb = Embedding(x_voc, embedding_dim,
trainable=True)(encoder_inputs)
# Encoder LSTM 1
encoder_lstm1 = Bidirectional(LSTM(latent_dim, return_sequences=True,
return_state=True, dropout=0.4,
recurrent_dropout=0.4))
(encoder_output1, forward_h1, forward_c1, backward_h1, backward_c1) = encoder_lstm1(enc_emb)
# Encoder LSTM 2
encoder_lstm2 = Bidirectional(LSTM(latent_dim, return_sequences=True,
return_state=True, dropout=0.4,
recurrent_dropout=0.4))
(encoder_output2, forward_h2, forward_c2, backward_h2, backward_c2) = encoder_lstm2(encoder_output1)
# Encoder LSTM 3
encoder_lstm3 = Bidirectional(LSTM(latent_dim, return_state=True,
return_sequences=True, dropout=0.4,
recurrent_dropout=0.4))
(encoder_outputs, forward_h, forward_c, backward_h, backward_c) = encoder_lstm3(encoder_output2)
state_h …Run Code Online (Sandbox Code Playgroud) 我已经使用 keras 来使用预训练的词嵌入,但我不太确定如何在 scikit-learn 模型上做到这一点。
我也需要在 sklearn 中执行此操作,因为我正在使用vecstack集成 keras 顺序模型和 sklearn 模型。
这就是我为 keras 模型所做的:
glove_dir = '/home/Documents/Glove'
embeddings_index = {}
f = open(os.path.join(glove_dir, 'glove.6B.200d.txt'), 'r', encoding='utf-8')
for line in f:
values = line.split()
word = values[0]
coefs = np.asarray(values[1:], dtype='float32')
embeddings_index[word] = coefs
f.close()
embedding_dim = 200
embedding_matrix = np.zeros((max_words, embedding_dim))
for word, i in word_index.items():
if i < max_words:
embedding_vector = embeddings_index.get(word)
if embedding_vector is not None:
embedding_matrix[i] = embedding_vector
model = Sequential()
model.add(Embedding(max_words, embedding_dim, input_length=maxlen)) …Run Code Online (Sandbox Code Playgroud) 我正在做情绪分析,我想使用预训练的 fasttext 嵌入,但是文件非常大(6.7 GB)并且程序需要很长时间才能编译。
fasttext_dir = '/Fasttext'
embeddings_index = {}
f = open(os.path.join(fasttext_dir, 'wiki.en.vec'), 'r', encoding='utf-8')
for line in tqdm(f):
values = line.rstrip().rsplit(' ')
word = values[0]
coefs = np.asarray(values[1:], dtype='float32')
embeddings_index[word] = coefs
f.close()
print('found %s word vectors' % len(embeddings_index))
embedding_dim = 300
embedding_matrix = np.zeros((max_words, embedding_dim))
for word, i in word_index.items():
if i < max_words:
embedding_vector = embeddings_index.get(word)
if embedding_vector is not None:
embedding_matrix[i] = embedding_vector
Run Code Online (Sandbox Code Playgroud)
有什么办法可以加快这个过程吗?
我想在散景中绘制多个图?例如,这段代码以非常低效的方式完成了我想要的事情。我想要同样的东西,但可能带有loop? 或任何我不知道的散景功能。
p = figure(...)
p1 = figure(...)
p2 = figure(...)
y = [[1,2,3,4,5,6],[7,8,9,10,11,12],[3,1,4,3,2,5]]
x = [[2,3,4,5,6,7],[8,9,10,11,12,13],[1,4,3,2,5,6]]
plots = []
p.line(x=np.arange(6), y=y[0], color='#CE1141', legend='Prediction')
p.line(x=np.arange(6,12), y=x[0], color='#006BB6', legend='Prediction')
plots.append(p)
p1.line(x=np.arange(6), y=y[1], color='#CE1141', legend='Prediction')
p1.line(x=np.arange(6,12), y=x[1], color='#006BB6', legend='Prediction')
plots.append(p1)
p2.line(x=np.arange(6), y=y[2], color='#CE1141', legend='Prediction')
p2.line(x=np.arange(6,12), y=x[2], color='#006BB6', legend='Prediction')
plots.append(p2)
show(column(*plots))
Run Code Online (Sandbox Code Playgroud)
基本上,我有两个二维数组,我想将它们绘制在一个图中。我在 a 中尝试过,for loop但它绘制了图中的所有内容并多次显示相同的图,我尝试了类似的操作:
for i in range(3):
p.line(x=np.arange(6), y=y[i], color='#CE1141', legend='Prediction')
p.line(x=np.arange(6,12), y=x[i], color='#006BB6', legend='Prediction')
p.show()
Run Code Online (Sandbox Code Playgroud)
我可以在这里看到问题,所有内容都绘制在 中p,所以最后我得到了一个所有内容都被绘制的图。
我还尝试创建一个数组并在其中附加绘图,如上面工作效率低下的示例所示,但它显示第一列的空白图,并最后绘制包含其中所有内容的单个图。
已经有很多关于删除异常值的问题,但我无法用它们解决我的问题。
我想从 中删除带有异常值的行dataframe。
说,我有以下内容dataframe:
0 1 2 3 4 5 6 7
a 1 2 3 4 100 2 1 3
b 2 1 3 4 1 2 300 123
c 100 200 300 400 200 500 200 400
Run Code Online (Sandbox Code Playgroud)
对于 row,a我们可以假设 100 是异常值,所以我想删除a.
尽管 Row 中的所有值c都很高,但它们并不是该行本身的异常值,因此,我想保留它。
所以,基本上我想删除所有带有异常值的行。
我尝试调换 DF 并做了类似的事情
df = df[(np.abs(stats.zscore(df)) < 2).all(axis=1)],但没有成功
我在看iriskedro 提供的项目示例。除了记录准确性之外,我还想将predictions和保存test_y为 csv。
这是 kedro 提供的示例节点。
def report_accuracy(predictions: np.ndarray, test_y: pd.DataFrame) -> None:
"""Node for reporting the accuracy of the predictions performed by the
previous node. Notice that this function has no outputs, except logging.
"""
# Get true class index
target = np.argmax(test_y.to_numpy(), axis=1)
# Calculate accuracy of predictions
accuracy = np.sum(predictions == target) / target.shape[0]
# Log the accuracy of the model
log = logging.getLogger(__name__)
log.info("Model accuracy on test set: %0.2f%%", accuracy …Run Code Online (Sandbox Code Playgroud)