小编San*_*ta7的帖子

zipfile extractall 引发“BadZipFile: Bad CRC-32 for file”错误

这是我试图打开的文件

https://drive.google.com/file/d/1K2kDBTNXS2ikx9xKmi2Fy0Wsc5u_Lls0/view

它描述here

https://github.com/armancohan/long-summarization

将文件添加到我的谷歌驱动器后,这是我试图用来打开它的代码。

from google.colab import drive
drive.mount('/content/gdrive')
import zipfile

zip_ref = zipfile.ZipFile('/content/gdrive/My Drive/arxiv-release.zip', 'r')
zip_ref.extractall('arxiv-release')
zip_ref.close()
Run Code Online (Sandbox Code Playgroud)

这是引发的错误

---------------------------------------------------------------------------
BadZipFile                                Traceback (most recent call last)
<ipython-input-9-9965160388a1> in <module>()
      1 
----> 2 zip_ref.extractall('arxiv-release')
      3 zip_ref.close()

5 frames
/usr/lib/python3.6/zipfile.py in extractall(self, path, members, pwd)
   1522 
   1523         for zipinfo in members:
-> 1524             self._extract_member(zipinfo, path, pwd)
   1525 
   1526     @classmethod

/usr/lib/python3.6/zipfile.py in _extract_member(self, member, targetpath, pwd)
   1577         with self.open(member, pwd=pwd) as source, \
   1578              open(targetpath, "wb") as target:
-> 1579             shutil.copyfileobj(source, target) …
Run Code Online (Sandbox Code Playgroud)

python zip

6
推荐指数
0
解决办法
563
查看次数

wget 现在自动将输出重定向到日志文件,如何返回到将输出放在下面

每当我使用 wget 时,下载的输出和进度都会显示在下面。

我今天刚尝试使用它,它只说'将输出重定向到'wget-log.4'。

有没有办法让它恢复到原来的样子?

这是我在 python 3 中运行它时的示例

!pip install wget
!wget -i https://s3-us-west-2.amazonaws.com/ai2-s2-research-public/open-corpus/corpus-2018-05-03/s2-corpus-01.gz
Run Code Online (Sandbox Code Playgroud)

将输出重定向到“wget-log.4”。

linux wget

5
推荐指数
1
解决办法
6860
查看次数

Tesnorflow:如何为 tf.nn.sampled_softmax_loss 提供您自己的 `sampled_values`?

在 tf.nn.sampled_softmax_loss 中,可选输入之一是放置您自己的样本值。我想提供我自己的样本值,以便我可以使用 float16(半精度)变量。如果sampled_values留空,Tensorflow 将使用log_uniform_candidate_sampler获取值,该值只能返回 float32。

这里是所有的输入。

tf.nn.sampled_softmax_loss(
    weights,
    biases,
    labels,
    inputs,
    num_sampled,
    num_classes,
    num_true=1,
    sampled_values=None,
    remove_accidental_hits=True,
    partition_strategy='mod',
    name='sampled_softmax_loss',
    seed=None
)
Run Code Online (Sandbox Code Playgroud)

https://www.tensorflow.org/api_docs/python/tf/nn/sampled_softmax_loss

这是他们为 sampled_values arg 提供的信息:

sampled_values:*_candidate_sampler 函数返回的 (sampled_candidates, true_expected_count, sampled_expected_count) 元组。(如果没有,我们默认为 log_uniform_candidate_sampler)

我想弄清楚如何提供这个元组。sampled_candidates, true_expected_count,究竟是什么sampled_expected_count?

我知道它正在对权重和相应的偏差进行采样,所以我是否将它们放在它自己的元组中sampled_candidates?另外,我是将 int 放在矩阵中的权重位置,还是将整个嵌入本身放入?

我还查看了 Tensorflow 对负采样的数学补充,但找不到有关我的问题的任何信息https://www.tensorflow.org/extras/candidate_sampling.pdf

在我的搜索中,我在谷歌论坛上发现了这个非常相似的问题

https://groups.google.com/a/tensorflow.org/forum/#!topic/discuss/6IDJ-XAIb9M

给出的答案是

sampled_values是我们的 *candidate_sampler 类返回的元组。这些类实现了根据一些分布 Q 对对比标签(未观察到,但在训练期间使用)进行采样的方法,以用于近似训练方法,如噪声对比估计 (NCE) 和采样 Softmax。一个例子是 log_uniform_candidate_sampler,它根据对数均匀分布对标签进行采样。

您几乎不需要自己提供这些。您只需将调用结果传递给 tf.nn 模块中的 *candidate_sampler 函数(其中 * 可以是“uniform”、“log_uniform”、“zipfian_binned”等),例如

sampled_values = tf.nn.zipfian_binned_candidate_sampler(...)

如果您只是想让它工作,只需将其保留为 None,它将默认为 log_uniform_candidate_sampler(通常是一个不错的选择)。

如果您对此背后的数学感兴趣,请参阅此文档: …

python tensorflow

5
推荐指数
0
解决办法
379
查看次数

在Tensorflow中使用TPU时,是否有一个可靠的解决方法来保存本地驱动器中的检查点?

对此问题的跟进:

在使用TPU模式时如何从Google Colaboratory保存Tensorflow Checkpoint文件?

使用Tensorflow TPU时保存检查点的官方方法是使用Google Cloud Service.

如果对于那些不希望使用GCS的人有解决方法,我正在工作.也许对于每个变量,执行.eval(),保存变量.然后将save变量设置为每个变量的'init'值.

我预见的一个主要问题是保存和加载优化器的参数.

对于Keras来说,权重似乎从TPU保存到本地

https://colab.research.google.com/github/tensorflow/tpu/blob/master/tools/colab/shakespeare_with_tpu_and_keras.ipynb

INFO:tensorflow:将TPU权重复制到CPU

所以我想也有一个普遍的解决方法,不使用keras.

python tensorflow google-colaboratory google-cloud-tpu

5
推荐指数
1
解决办法
437
查看次数

将Tensorflow图转换为使用Estimator,使用`sampled_softmax_loss`或`nce_loss`在损失函数中获取'TypeError:数据类型不被理解'

我试图将Tensorflow的官方基本word2vec实现转换为使用tf.Estimator.问题是当使用Tensorflow Estimators时,损失函数(sampled_softmax_loss或nce_loss)会出错.它在原始实现中完美地运行.

这是Tensorflow的官方基本word2vec实现:

https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/tutorials/word2vec/word2vec_basic.py

以下是我实施此代码的Google Colab笔记本,该代码正常运行.

https://colab.research.google.com/drive/1nTX77dRBHmXx6PEF5pmYpkIVxj_TqT5I

这是Google Colab笔记本,我在其中更改了代码,因此它使用Tensorflow Estimator,它不起作用.

https://colab.research.google.com/drive/1IVDqGwMx6BK5-Bgrw190jqHU6tt3ZR3e

为方便起见,这里是我定义的Estimator版本的精确代码 model_fn

batch_size = 128
embedding_size = 128  # Dimension of the embedding vector.
skip_window = 1  # How many words to consider left and right.
num_skips = 2  # How many times to reuse an input to generate a label.
num_sampled = 64  # Number of negative examples to sample.

def my_model( features, labels, mode, params):

    with tf.name_scope('inputs'):
        train_inputs = features
        train_labels = labels …
Run Code Online (Sandbox Code Playgroud)

python tensorflow tensorflow-estimator

5
推荐指数
1
解决办法
331
查看次数

在 Tensorflow 数据集 api 中:如何使用 padded_batch 以便在不指定 pads 数量的情况下填充具有特定值的 pads

如果您不指定 apadding_values那么padded_batch将自动填充 0。但是,如果您想要不同的值,例如 -1,则不能只设置padded_batch = -1。您需要为需要填充的每个插槽输入一个序列。

但是,我正在使用一个具有随机数组长度值的数据集,所以我不能真正做到这一点,因为我不知道需要填充多少个数字。

由于padding_values会自动用 0 填充其余的值,我希望有某种方法可以使用不同的值(例如“-1”)来做到这一点。

这是一个最小的例子

import math
import numpy as np
import tensorflow as tf

cells = np.array([[0,1,2,3], [2,3,4], [3,6,5,4,3], [3,9]])
mells = np.array([[0], [2], [3], [9]])
print(cells)

writer = tf.python_io.TFRecordWriter('test.tfrecords')
for index in range(mells.shape[0]):
    example = tf.train.Example(features=tf.train.Features(feature={
        'num_value':tf.train.Feature(int64_list=tf.train.Int64List(value=mells[index])),
        'list_value':tf.train.Feature(int64_list=tf.train.Int64List(value=cells[index]))
    }))
    writer.write(example.SerializeToString())
writer.close()

#Generate Samples with batch size of 2

filenames = ["test.tfrecords"]
dataset = tf.data.TFRecordDataset(filenames)
def _parse_function(example_proto):
    keys_to_features = {'num_value':tf.VarLenFeature(tf.int64),
                        'list_value':tf.VarLenFeature(tf.int64)}
    parsed_features = tf.parse_single_example(example_proto, keys_to_features) …
Run Code Online (Sandbox Code Playgroud)

tensorflow tensorflow-datasets

5
推荐指数
1
解决办法
5666
查看次数

如何将组件的名称添加/更改到现有的 Tensorflow 数据集对象?

从 Tensorflow 数据集指南它说

为元素的每个组件命名通常很方便,例如,如果它们代表训练示例的不同特征。除了元组,您还可以使用 collections.namedtuple 或将字符串映射到张量的字典来表示数据集的单个元素。

dataset = tf.data.Dataset.from_tensor_slices(
   {"a": tf.random_uniform([4]),
    "b": tf.random_uniform([4, 100], maxval=100, dtype=tf.int32)})
print(dataset.output_types)  # ==> "{'a': tf.float32, 'b': tf.int32}"
print(dataset.output_shapes)  # ==> "{'a': (), 'b': (100,)}"
Run Code Online (Sandbox Code Playgroud)

https://www.tensorflow.org/guide/datasets

这在 Keras 中非常有用。如果将数据集对象传递给model.fit,则组件的名称可用于匹配 Keras 模型的输入。例子:

image_input = keras.Input(shape=(32, 32, 3), name='img_input')
timeseries_input = keras.Input(shape=(None, 10), name='ts_input')

x1 = layers.Conv2D(3, 3)(image_input)
x1 = layers.GlobalMaxPooling2D()(x1)

x2 = layers.Conv1D(3, 3)(timeseries_input)
x2 = layers.GlobalMaxPooling1D()(x2)

x = layers.concatenate([x1, x2])

score_output = layers.Dense(1, name='score_output')(x)
class_output = layers.Dense(5, activation='softmax', name='class_output')(x)

model = keras.Model(inputs=[image_input, timeseries_input], …
Run Code Online (Sandbox Code Playgroud)

python tensorflow tensorflow-datasets

5
推荐指数
2
解决办法
1298
查看次数

大(600 万行)pandas df 当 chunksize =100 时会导致内存错误,使用 `to_sql`,但可以轻松保存 100,000 的文件而没有 chunksize

我在 Pandas 中创建了一个大型数据库,大约有 600 万行文本数据。我想将其保存为 SQL 数据库文件,但是当我尝试保存它时,出现内存不足的 RAM 错误。我什至将卡盘尺寸减小到 100,但它仍然崩溃。

但是,如果我只有具有 100,000 行的该数据帧的较小版本,并将其保存到未指定chucksize 的数据库中,则保存数据帧没有问题。

这是我的代码

from sqlalchemy import create_engine
engine = sqlalchemy.create_engine("sqlite:///databasefile.db")
dataframe.to_sql("CS_table", engine, chunksize = 100)
Run Code Online (Sandbox Code Playgroud)

我的理解是,由于它一次仅处理 100 行,因此 RAM 使用量应反映保存 100 行的情况。幕后还有其他事情发生吗?也许多线程?

在运行此代码之前,我使用的是 4.8 GB RAM,而 Google Colab 中提供了 12.8 GB RAM。运行上述代码会耗尽所有 RAM,直到环境崩溃。

我希望能够将我的 Pandas 数据帧保存到 SQL 文件中,而我的环境不会崩溃。我所在的环境是 Google Colab。Pandas 数据名是 2 列,约 600 万行。每个单元格包含大约这么多文本:

在八个 GPU 上训练 3.5 天后,我们的模型建立了一个新的单模型最先进的 BLEU 分数 41.8,这是文献中最佳模型训练成本的一小部分。我们表明,通过将 Transformer 成功应用于具有大量和有限训练数据的英语选区解析,可以很好地推广到其他任务。”

编辑:

我在不同阶段做了键盘中断。这是在RAM中第一次跳转后键盘中断的结果

---------------------------------------------------------------------------
KeyboardInterrupt                         Traceback (most recent call last)
<ipython-input-22-51b6e444f80d> in <module>()
----> …
Run Code Online (Sandbox Code Playgroud)

python sql pandas

5
推荐指数
2
解决办法
5304
查看次数

为什么 Tensorflow 中的分布策略不支持梯度裁剪?

看起来使用分布策略不支持梯度裁剪

https://github.com/tensorflow/tensorflow/blob/f9f6b4cec2a1bdc5781e4896d80cee1336a2fbab/tensorflow/python/keras/optimizer_v2/optimizer_v2.py#L383

(“使用分布策略时,当前不支持优化器中的梯度裁剪”(通过设置clipnorm或clipvalue)。”)

这有什么原因吗?我很想定义一个def _minimize(strategy, tape, optimizer, loss, trainable_variables):直接剪切渐变的自定义。

tensorflow

5
推荐指数
1
解决办法
828
查看次数

如何在python中获取两个字符串之间的所有模糊匹配子串?

假设我有三个示例字符串

text1 = "Patient has checked in for abdominal pain which started 3 days ago. Patient was prescribed idx 20 mg every 4 hours."
text2 = "The time of discomfort was 3 days ago."
text3 = "John was given a prescription of idx, 20mg to be given every four hours"
Run Code Online (Sandbox Code Playgroud)

如果我得到 text2 和 text3 与 text1 的所有匹配子字符串,我会得到

text1_text2_common = [
    '3 days ago.',
]

text2_text3_common = [
    'of',
]

text1_text3_common = [
    'was',
    'idx'
    'every'
    'hours'
]
Run Code Online (Sandbox Code Playgroud)

我正在寻找的是模糊匹配,使用诸如Levenshtein distance …

python string fuzzy-search

5
推荐指数
1
解决办法
1979
查看次数