how to understand the code for calculating rewards

Author: lyzKFCreated Aug 22, 2023Updated Aug 28, 2025

https://github.com/microsoft/DeepSpeedExamples/blob/master/applications/DeepSpeed-Chat/training/step3_rlhf_finetuning/ppo_trainer.py#L148

  def compute_rewards(self, prompts, log_probs, ref_log_probs, reward_score,
                      action_mask):
      kl_divergence_estimate = -self.kl_ctl * (log_probs - ref_log_probs)
      rewards = kl_divergence_estimate
      start = prompts.shape[1] - 1
      ends = start + action_mask[:, start:].sum(1) + 1
      reward_clip = torch.clamp(reward_score, -self.clip_reward_value,  self.clip_reward_value)
      batch_size = log_probs.shape[0]
      for j in range(batch_size):
          rewards[j, start:ends[j]][-1] += reward_clip[j]

      return rewards

why should we calculate rewards like this? is there a theory to back it up?

Source: deepspeedai/DeepSpeedExamples