这是我试图打开的文件
https://drive.google.com/file/d/1K2kDBTNXS2ikx9xKmi2Fy0Wsc5u_Lls0/view
它描述here
https://github.com/armancohan/long-summarization
将文件添加到我的谷歌驱动器后,这是我试图用来打开它的代码。
from google.colab import drive
drive.mount('/content/gdrive')
import zipfile
zip_ref = zipfile.ZipFile('/content/gdrive/My Drive/arxiv-release.zip', 'r')
zip_ref.extractall('arxiv-release')
zip_ref.close()
Run Code Online (Sandbox Code Playgroud)
这是引发的错误
---------------------------------------------------------------------------
BadZipFile Traceback (most recent call last)
<ipython-input-9-9965160388a1> in <module>()
1
----> 2 zip_ref.extractall('arxiv-release')
3 zip_ref.close()
5 frames
/usr/lib/python3.6/zipfile.py in extractall(self, path, members, pwd)
1522
1523 for zipinfo in members:
-> 1524 self._extract_member(zipinfo, path, pwd)
1525
1526 @classmethod
/usr/lib/python3.6/zipfile.py in _extract_member(self, member, targetpath, pwd)
1577 with self.open(member, pwd=pwd) as source, \
1578 open(targetpath, "wb") as target:
-> 1579 shutil.copyfileobj(source, target) …Run Code Online (Sandbox Code Playgroud) 每当我使用 wget 时,下载的输出和进度都会显示在下面。
我今天刚尝试使用它,它只说'将输出重定向到'wget-log.4'。
有没有办法让它恢复到原来的样子?
这是我在 python 3 中运行它时的示例
!pip install wget
!wget -i https://s3-us-west-2.amazonaws.com/ai2-s2-research-public/open-corpus/corpus-2018-05-03/s2-corpus-01.gz
Run Code Online (Sandbox Code Playgroud)
将输出重定向到“wget-log.4”。
在 tf.nn.sampled_softmax_loss 中,可选输入之一是放置您自己的样本值。我想提供我自己的样本值,以便我可以使用 float16(半精度)变量。如果sampled_values留空,Tensorflow 将使用log_uniform_candidate_sampler获取值,该值只能返回 float32。
这里是所有的输入。
tf.nn.sampled_softmax_loss(
weights,
biases,
labels,
inputs,
num_sampled,
num_classes,
num_true=1,
sampled_values=None,
remove_accidental_hits=True,
partition_strategy='mod',
name='sampled_softmax_loss',
seed=None
)
Run Code Online (Sandbox Code Playgroud)
https://www.tensorflow.org/api_docs/python/tf/nn/sampled_softmax_loss
这是他们为 sampled_values arg 提供的信息:
sampled_values:*_candidate_sampler 函数返回的 (sampled_candidates, true_expected_count, sampled_expected_count) 元组。(如果没有,我们默认为 log_uniform_candidate_sampler)
我想弄清楚如何提供这个元组。sampled_candidates, true_expected_count,究竟是什么sampled_expected_count?
我知道它正在对权重和相应的偏差进行采样,所以我是否将它们放在它自己的元组中sampled_candidates?另外,我是将 int 放在矩阵中的权重位置,还是将整个嵌入本身放入?
我还查看了 Tensorflow 对负采样的数学补充,但找不到有关我的问题的任何信息https://www.tensorflow.org/extras/candidate_sampling.pdf
在我的搜索中,我在谷歌论坛上发现了这个非常相似的问题
https://groups.google.com/a/tensorflow.org/forum/#!topic/discuss/6IDJ-XAIb9M
给出的答案是
sampled_values是我们的 *candidate_sampler 类返回的元组。这些类实现了根据一些分布 Q 对对比标签(未观察到,但在训练期间使用)进行采样的方法,以用于近似训练方法,如噪声对比估计 (NCE) 和采样 Softmax。一个例子是 log_uniform_candidate_sampler,它根据对数均匀分布对标签进行采样。您几乎不需要自己提供这些。您只需将调用结果传递给 tf.nn 模块中的 *candidate_sampler 函数(其中 * 可以是“uniform”、“log_uniform”、“zipfian_binned”等),例如
sampled_values = tf.nn.zipfian_binned_candidate_sampler(...)
如果您只是想让它工作,只需将其保留为 None,它将默认为 log_uniform_candidate_sampler(通常是一个不错的选择)。
如果您对此背后的数学感兴趣,请参阅此文档: …
对此问题的跟进:
在使用TPU模式时如何从Google Colaboratory保存Tensorflow Checkpoint文件?
使用Tensorflow TPU时保存检查点的官方方法是使用Google Cloud Service.
如果对于那些不希望使用GCS的人有解决方法,我正在工作.也许对于每个变量,执行.eval(),保存变量.然后将save变量设置为每个变量的'init'值.
我预见的一个主要问题是保存和加载优化器的参数.
对于Keras来说,权重似乎从TPU保存到本地
INFO:tensorflow:将TPU权重复制到CPU
所以我想也有一个普遍的解决方法,不使用keras.
我试图将Tensorflow的官方基本word2vec实现转换为使用tf.Estimator.问题是当使用Tensorflow Estimators时,损失函数(sampled_softmax_loss或nce_loss)会出错.它在原始实现中完美地运行.
这是Tensorflow的官方基本word2vec实现:
以下是我实施此代码的Google Colab笔记本,该代码正常运行.
https://colab.research.google.com/drive/1nTX77dRBHmXx6PEF5pmYpkIVxj_TqT5I
这是Google Colab笔记本,我在其中更改了代码,因此它使用Tensorflow Estimator,它不起作用.
https://colab.research.google.com/drive/1IVDqGwMx6BK5-Bgrw190jqHU6tt3ZR3e
为方便起见,这里是我定义的Estimator版本的精确代码 model_fn
batch_size = 128
embedding_size = 128 # Dimension of the embedding vector.
skip_window = 1 # How many words to consider left and right.
num_skips = 2 # How many times to reuse an input to generate a label.
num_sampled = 64 # Number of negative examples to sample.
def my_model( features, labels, mode, params):
with tf.name_scope('inputs'):
train_inputs = features
train_labels = labels …Run Code Online (Sandbox Code Playgroud) 如果您不指定 apadding_values那么padded_batch将自动填充 0。但是,如果您想要不同的值,例如 -1,则不能只设置padded_batch = -1。您需要为需要填充的每个插槽输入一个序列。
但是,我正在使用一个具有随机数组长度值的数据集,所以我不能真正做到这一点,因为我不知道需要填充多少个数字。
由于padding_values会自动用 0 填充其余的值,我希望有某种方法可以使用不同的值(例如“-1”)来做到这一点。
这是一个最小的例子
import math
import numpy as np
import tensorflow as tf
cells = np.array([[0,1,2,3], [2,3,4], [3,6,5,4,3], [3,9]])
mells = np.array([[0], [2], [3], [9]])
print(cells)
writer = tf.python_io.TFRecordWriter('test.tfrecords')
for index in range(mells.shape[0]):
example = tf.train.Example(features=tf.train.Features(feature={
'num_value':tf.train.Feature(int64_list=tf.train.Int64List(value=mells[index])),
'list_value':tf.train.Feature(int64_list=tf.train.Int64List(value=cells[index]))
}))
writer.write(example.SerializeToString())
writer.close()
#Generate Samples with batch size of 2
filenames = ["test.tfrecords"]
dataset = tf.data.TFRecordDataset(filenames)
def _parse_function(example_proto):
keys_to_features = {'num_value':tf.VarLenFeature(tf.int64),
'list_value':tf.VarLenFeature(tf.int64)}
parsed_features = tf.parse_single_example(example_proto, keys_to_features) …Run Code Online (Sandbox Code Playgroud) 从 Tensorflow 数据集指南它说
为元素的每个组件命名通常很方便,例如,如果它们代表训练示例的不同特征。除了元组,您还可以使用 collections.namedtuple 或将字符串映射到张量的字典来表示数据集的单个元素。
dataset = tf.data.Dataset.from_tensor_slices(
{"a": tf.random_uniform([4]),
"b": tf.random_uniform([4, 100], maxval=100, dtype=tf.int32)})
print(dataset.output_types) # ==> "{'a': tf.float32, 'b': tf.int32}"
print(dataset.output_shapes) # ==> "{'a': (), 'b': (100,)}"
Run Code Online (Sandbox Code Playgroud)
https://www.tensorflow.org/guide/datasets
这在 Keras 中非常有用。如果将数据集对象传递给model.fit,则组件的名称可用于匹配 Keras 模型的输入。例子:
image_input = keras.Input(shape=(32, 32, 3), name='img_input')
timeseries_input = keras.Input(shape=(None, 10), name='ts_input')
x1 = layers.Conv2D(3, 3)(image_input)
x1 = layers.GlobalMaxPooling2D()(x1)
x2 = layers.Conv1D(3, 3)(timeseries_input)
x2 = layers.GlobalMaxPooling1D()(x2)
x = layers.concatenate([x1, x2])
score_output = layers.Dense(1, name='score_output')(x)
class_output = layers.Dense(5, activation='softmax', name='class_output')(x)
model = keras.Model(inputs=[image_input, timeseries_input], …Run Code Online (Sandbox Code Playgroud) 我在 Pandas 中创建了一个大型数据库,大约有 600 万行文本数据。我想将其保存为 SQL 数据库文件,但是当我尝试保存它时,出现内存不足的 RAM 错误。我什至将卡盘尺寸减小到 100,但它仍然崩溃。
但是,如果我只有具有 100,000 行的该数据帧的较小版本,并将其保存到未指定chucksize 的数据库中,则保存数据帧没有问题。
这是我的代码
from sqlalchemy import create_engine
engine = sqlalchemy.create_engine("sqlite:///databasefile.db")
dataframe.to_sql("CS_table", engine, chunksize = 100)
Run Code Online (Sandbox Code Playgroud)
我的理解是,由于它一次仅处理 100 行,因此 RAM 使用量应反映保存 100 行的情况。幕后还有其他事情发生吗?也许多线程?
在运行此代码之前,我使用的是 4.8 GB RAM,而 Google Colab 中提供了 12.8 GB RAM。运行上述代码会耗尽所有 RAM,直到环境崩溃。
我希望能够将我的 Pandas 数据帧保存到 SQL 文件中,而我的环境不会崩溃。我所在的环境是 Google Colab。Pandas 数据名是 2 列,约 600 万行。每个单元格包含大约这么多文本:
在八个 GPU 上训练 3.5 天后,我们的模型建立了一个新的单模型最先进的 BLEU 分数 41.8,这是文献中最佳模型训练成本的一小部分。我们表明,通过将 Transformer 成功应用于具有大量和有限训练数据的英语选区解析,可以很好地推广到其他任务。”
编辑:
我在不同阶段做了键盘中断。这是在RAM中第一次跳转后键盘中断的结果
---------------------------------------------------------------------------
KeyboardInterrupt Traceback (most recent call last)
<ipython-input-22-51b6e444f80d> in <module>()
----> …Run Code Online (Sandbox Code Playgroud) 看起来使用分布策略不支持梯度裁剪
(“使用分布策略时,当前不支持优化器中的梯度裁剪”(通过设置clipnorm或clipvalue)。”)
这有什么原因吗?我很想定义一个def _minimize(strategy, tape, optimizer, loss, trainable_variables):直接剪切渐变的自定义。
假设我有三个示例字符串
text1 = "Patient has checked in for abdominal pain which started 3 days ago. Patient was prescribed idx 20 mg every 4 hours."
text2 = "The time of discomfort was 3 days ago."
text3 = "John was given a prescription of idx, 20mg to be given every four hours"
Run Code Online (Sandbox Code Playgroud)
如果我得到 text2 和 text3 与 text1 的所有匹配子字符串,我会得到
text1_text2_common = [
'3 days ago.',
]
text2_text3_common = [
'of',
]
text1_text3_common = [
'was',
'idx'
'every'
'hours'
]
Run Code Online (Sandbox Code Playgroud)
我正在寻找的是模糊匹配,使用诸如Levenshtein distance …