How to define max_queue_size, workers and use_multiprocessing in keras fit_generator()?

Sop*_*nck 21 python gpu machine-learning keras tensorflow

I am applying transfer-learning on a pre-trained network using the GPU version of keras. I don't understand how to define the parameters max_queue_size, workers, and use_multiprocessing. If I change these parameters (primarily to speed-up learning), I am unsure whether all data is still seen per epoch.

max_queue_size:

  • maximum size of the internal training queue which is used to "precache" samples from the generator

  • Question: Does this refer to how many batches are prepared on CPU? How is it related to workers? How to define it optimally?

workers:

  • number of threads generating batches in parallel. Batches are computed in parallel on the CPU and passed on the fly onto the GPU for neural network computations

  • Question: How do I find out how many batches my CPU can/should generate in parallel?

use_multiprocessing:

  • whether to use process-based threading

  • Question: Do I have to set this parameter to true if I change workers? Does it relate to CPU usage?

Related questions can be found here:

I am using fit_generator() as follows:

    history = model.fit_generator(generator=trainGenerator,
                                  steps_per_epoch=trainGenerator.samples//nBatches,     # total number of steps (batches of samples)
                                  epochs=nEpochs,                   # number of epochs to train the model
                                  verbose=2,                        # verbosity mode. 0 = silent, 1 = progress bar, 2 = one line per epoch
                                  callbacks=callback,               # keras.callbacks.Callback instances to apply during training
                                  validation_data=valGenerator,     # generator or tuple on which to evaluate the loss and any model metrics at the end of each epoch
                                  validation_steps=
                                  valGenerator.samples//nBatches,   # number of steps (batches of samples) to yield from validation_data generator before stopping at the end of every epoch
                                  class_weight=classWeights,                # optional dictionary mapping class indices (integers) to a weight (float) value, used for weighting the loss function
                                  max_queue_size=10,                # maximum size for the generator queue
                                  workers=1,                        # maximum number of processes to spin up when using process-based threading
                                  use_multiprocessing=False,        # whether to use process-based threading
                                  shuffle=True,                     # whether to shuffle the order of the batches at the beginning of each epoch
                                  initial_epoch=0)   
Run Code Online (Sandbox Code Playgroud)

The specs of my machine are:

CPU : 2xXeon E5-2260 2.6 GHz
Cores: 10
Graphic card: Titan X, Maxwell, GM200
RAM: 128 GB
HDD: 4TB
SSD: 512 GB
Run Code Online (Sandbox Code Playgroud)

a-d*_*a-d 14

Q_0:

问题:这是否指的是在CPU上准备多少批次?它与工人有什么关系?如何最佳定义?

从发布的链接中,您可以了解到CPU一直在创建批处理,直到队列达到最大队列大小或到达停止为止。您需要准备好批处理以供GPU“使用”,以使GPU不必等待CPU。队列大小的理想值是使其足够大,以使您的GPU始终在接近最大值的情况下运行,而不必等待CPU准备新批处理。

Q_1:

问题:如何找出我的CPU可以/应该并行生成多少个批次?

如果您发现GPU处于空闲状态并正在等待批处理,请尝试增加工作程序的数量,也许还增加队列的大小。

Q_2:

如果更改工作人员,是否必须将此参数设置为true?它与CPU使用率有关吗?

是将其设置为True或时发生的情况的实用分析False这里是一个建议,将其设置为False防止冻结(在我的设置True工作正常不结冰)。也许其他人可以增进我们对该主题的理解。

综上所述:

尝试不进行顺序设置,尝试使CPU为GPU提供足够的数据。

另外:您可以(应该?)在下一次提出几个问题,以便于回答。

  • 很有帮助,但我不同意问题张贴者应分别提出这些问题。这些问题是相关的,例如,您在文稿末尾做了一句话摘要。 (6认同)