MuZero伪代码中的奖励值是否错位?

Pro*_*o Q 4 python algorithm artificial-intelligence structure machine-learning

MuZero是一种深度强化学习技术,刚刚发布,我一直在尝试通过查看其伪代码和 Medium 上的这篇有用教程来实现它。

然而,我对伪代码训练期间如何处理奖励感到困惑,如果有人能够验证我是否正确地阅读了代码,那就太好了,如果我是,请解释为什么这个训练算法有效。

这是训练函数(来自伪代码):

def update_weights(optimizer: tf.train.Optimizer, network: Network, batch,
                   weight_decay: float):
  loss = 0
  for image, actions, targets in batch:
    # Initial step, from the real observation.
    value, reward, policy_logits, hidden_state = network.initial_inference(
        image)
    predictions = [(1.0, value, reward, policy_logits)]

    # Recurrent steps, from action and previous hidden state.
    for action in actions:
      value, reward, policy_logits, hidden_state = network.recurrent_inference(
          hidden_state, action)
      predictions.append((1.0 / len(actions), value, reward, policy_logits))

      hidden_state = tf.scale_gradient(hidden_state, 0.5)

    for prediction, target in zip(predictions, targets):
      gradient_scale, value, reward, policy_logits = prediction
      target_value, target_reward, target_policy = target

      l = (
          scalar_loss(value, target_value) +
          scalar_loss(reward, target_reward) +
          tf.nn.softmax_cross_entropy_with_logits(
              logits=policy_logits, labels=target_policy))

      loss += tf.scale_gradient(l, gradient_scale)

  for weights in network.get_weights():
    loss += weight_decay * tf.nn.l2_loss(weights)

  optimizer.minimize(loss)
Run Code Online (Sandbox Code Playgroud)

reward我对损失特别感兴趣。请注意,损失从 中获取其所有值predictions。第一个reward添加的predictions是来自network.initial_inference函数的。之后还有len(actions)更多的奖励加入predictions,都是来自于network.recurrent_inference功能。

根据教程initial_inference,recurrent_inference函数由 3 个不同的函数组成:

  1. 预测输入:内部游戏状态。输出:策略、价值(未来可能的最佳奖励的预测总和)
  2. 动态输入:游戏的内部状态、动作。输出:采取该行动的奖励,游戏的新内部状态。
  3. 表示输入:游戏的外部状态。输出:游戏的内部状态

该initial_inference函数接受外部游戏状态,使用该representation函数将其转换为内部状态,然后prediction对该内部游戏状态使用该函数。它输出内部状态、策略和值。

该recurrent_inference函数接受内部游戏状态和动作。它使用该dynamics函数从旧的游戏状态和动作中获取新的内部游戏状态和奖励。然后,它将prediction函数应用于新的内部游戏状态,以获得该新内部状态的策略和值。因此,最终的输出是一个新的内部状态、奖励、策略和值。

然而,在伪代码中,该initial_inference函数还返回奖励。

我的主要问题:这个奖励代表什么?

在教程中,他们只是隐含地假设initial_inference函数的奖励为 0。(请参阅教程中的这张图片。)那么这是怎么回事呢?难道真的没有奖励,所以initial_inference奖励总是返回0吗?

让我们假设情况确实如此。

在此假设下,列表中的第一个奖励将是函数将返回的奖励predictions0 。initial_inference然后,在损失中,这个0将与列表的第一个元素进行比较target。

target创建方式如下:

  def make_target(self, state_index: int, num_unroll_steps: int, td_steps: int,
                  to_play: Player):
    # The value target is the discounted root value of the search tree N steps
    # into the future, plus the discounted sum of all rewards until then.
    targets = []
    for current_index in range(state_index, state_index + num_unroll_steps + 1):
      bootstrap_index = current_index + td_steps
      if bootstrap_index < len(self.root_values):
        value = self.root_values[bootstrap_index] * self.discount**td_steps
      else:
        value = 0

      for i, reward in enumerate(self.rewards[current_index:bootstrap_index]):
        value += reward * self.discount**i  # pytype: disable=unsupported-operands

      if current_index < len(self.root_values):
        targets.append((value, self.rewards[current_index],
                        self.child_visits[current_index]))
      else:
        # States past the end of games are treated as absorbing states.
        targets.append((0, 0, []))
    return targets
Run Code Online (Sandbox Code Playgroud)

targets该函数返回的结果成为函数target中的列表update_weights。所以第一个值targets是self.rewards[current_index]。这self.rewards是玩游戏时收到的所有奖励的列表。唯一一次编辑是在此函数内apply:

  def apply(self, action: Action):
    reward = self.environment.step(action)
    self.rewards.append(reward)
    self.history.append(action)
Run Code Online (Sandbox Code Playgroud)

该apply函数仅在此处被调用:

# Each game is produced by starting at the initial board position, then
# repeatedly executing a Monte Carlo Tree Search to generate moves until the end
# of the game is reached.
def play_game(config: MuZeroConfig, network: Network) -> Game:
  game = config.new_game()

  while not game.terminal() and len(game.history) < config.max_moves:
    # At the root of the search tree we use the representation function to
    # obtain a hidden state given the current observation.
    root = Node(0)
    current_observation = game.make_image(-1)
    expand_node(root, game.to_play(), game.legal_actions(),
                network.initial_inference(current_observation))
    add_exploration_noise(config, root)

    # We then run a Monte Carlo Tree Search using only action sequences and the
    # model learned by the network.
    run_mcts(config, root, game.action_history(), network)
    action = select_action(config, len(game.history), root, network)
    game.apply(action)
    game.store_search_statistics(root)
  return game
Run Code Online (Sandbox Code Playgroud)

对我来说,看起来每次采取行动都会产生奖励。因此列表中的第一个奖励self.rewards应该是在游戏中采取第一个行动的奖励。

current_index = 0如果在 中,问题就很清楚了self.rewards[current_index]。在这种情况下,predictions列表中的第一个奖励将为 0,因为它总是如此。然而,targets列表中,将有完成第一个动作的奖励。

所以,对我来说,奖励似乎是错位的。

如果我们继续,列表中的第二个奖励将是完成第一个predictions操作的奖励。然而,列表中的第二个奖励将是游戏中存储的完成第二个动作的奖励。recurrent_inferencetargets

因此,总的来说,我有三个相互关联的问题:

  1. 获得的奖励代表什么initial_inference?(它是什么?)
  2. 如果它是 0,并且它应该代表奖励,那么 和 之间的奖励是否predictions未targets对齐?(即,中的第二个奖励predictions实际上应该与中的第一个奖励相匹配吗targets?)
  3. 如果它们不一致,网络是否仍能正确训练和工作?

(另一个需要注意的好奇心是,尽管存在这种错位(假设存在错位),但 和predictionslength的长度确实相同。目标长度由上面函数中的targets行定义。上面,我们还计算了 的长度为. And是由函数中定义的(参见伪代码)。因此,两个列表的长度相同。)for current_index in range(state_index, state_index + num_unroll_steps + 1)make_targetpredictionslen(actions) + 1len(actions)g.history[i:i + num_unroll_steps]sample_batch

这是怎么回事?

Mon*_*ofu 5

作者在此。

初始推理的奖励代表什么?

最初的推论“预测”了最后观察到的奖励。这实际上没有任何用途,但使我们的代码更简单:预测头总是可以简单地预测紧邻的前一个奖励。对于动态网络,这将是在应用作为动态网络输入给出的动作后观察到的奖励。

游戏开始时没有最后观察到的奖励,因此我们将其设置为 0。

伪代码中的奖励目标计算确实是错位的;我刚刚将新版本上传到 arXiv。

以前常说的地方

      if current_index < len(self.root_values):
        targets.append((value, self.rewards[current_index],
                        self.child_visits[current_index]))
      else:
        # States past the end of games are treated as absorbing states.
        targets.append((0, 0, []))
Run Code Online (Sandbox Code Playgroud)

它应该是:

      # For simplicity the network always predicts the most recently received
      # reward, even for the initial representation network where we already
      # know this reward.
      if current_index > 0 and current_index <= len(self.rewards):
        last_reward = self.rewards[current_index - 1]
      else:
        last_reward = 0

      if current_index < len(self.root_values):
        targets.append((value, last_reward, self.child_visits[current_index]))
      else:
        # States past the end of games are treated as absorbing states.
        targets.append((0, last_reward, []))
Run Code Online (Sandbox Code Playgroud)

希望有帮助!

  • 最初的推理中的这个预测是如何运作的?伪代码目前说它只是“表示+预测函数”,但论文说“表示”函数仅获取新的内部*状态*,而“预测”函数仅获取*策略*和*值*。那么最初推论中的“奖励”从何而来?表示函数也应该“预测”奖励吗? (3认同)