https://arxiv.org/pdf/2604.24532

Multi-Objective RL & Reward-Free RL:

  • For reward function obtained from the environment \(r'_1,\dots,r'_t\in R^d\) and a preference vector \(\lambda \in \Lambda\), we compute the true reward by \(r_t=\lambda^{\top} r_t'\).
  • We hope to get a model that, given any preference, directly outputs a policy to fit that preference.
  • Preference-conditioned MORL solves this problem by directly training on sampled preferences.
  • Reward Free RL tries to explore and learn all kinds of reward functions such that it could fit any unseen reward functions in evaluation. MORL is sort of a subset of RFRL.
  • However, RFRL is just too board for a small set of objectives.
  • The objective of this paper is to find a policy that shares weight among \(\lambda\) such that \(\pi_{\lambda}\) reaches rather good reward for preference \(\lambda\).

Methodology

Successor Measure

  • Define successor measure \(M^{\pi}(s,a,s',a')\) as the expected discounted number of times that we will be in state \((s',a')\) in the future when currently in state \((s,a)\) for policy \(\pi\).

  • The successor measure breaks the \(Q\) function into \(\sum {s',a'}M^{\pi}(s,a,s',a')R(s',a')\).

    This separates the dynamic estimation and value estimation.

Forward-Backward for reward-free RL

  • We use low-rank decomposition to approximate \(M^{\pi_z}(s,a,s',a')=F_{\theta}(s,a,z)^{\top}B_{\omega}(s',a')\), where \(z\) is a vector featuring the preference for the future action. Both vectors are of width \(d_z\).

    In this method, \(B\) acts as a \(d_z\)-dimension basis, representing some valuable latent features for future states & actions, while \(F\) weights these features (or, how many times we will encounter them).

  • The \(z\) vector could be obtained from reward function \(R\) by \(z_r=E_{(s,a)\sim D}[B_\omega(s,a)R(s,a)]\), with Q-value being \(Q(s,a,z_r)=F_{\theta}(s,a,z_r)^{\top}z_r\) and actor \(\pi_{s,z_r}=\arg\max Q(s,a,z_r)\).

Preference-Guided exploration

  • Vanilla reward-free RL samples \(z\) from gaussians, which is too board for multi-objective RL learning.
  • The paper solves this by:
    • Warm-up the replay buffer by \(z\sim N(0,I)\) and repeat the following process.
    • Sample a mini batch from replay buffer.
    • Calculate \(z_{\lambda}\) based on the mean \(B_{\omega}(s,a)r\lambda\).
    • Train the actor critic network.
  • The \(z\) is sampled from true reward and preference, which is more task-specific. On the other hand, the noise introduced by mini batch sampling increases variety.

Loss Design

  • Measure Loss

    First, we need to make sure the forward-backward measure is consistent with Bellman function. Thus we use the standard Q-loss to optimize \(F^{\top}B\).

    Second, we optimize the dynamic prediction by loss \(-2E[F(s_t,a_t,z)^{\top} B_{\omega}(s_{t+1},a_{t+1})]\).

  • Q Loss

    Pseudo reward is used in RFRL due to the blindness to reward. However, we could just use the standard Q-loss using the real \(\lambda^{\top}r\) in MORL setting to optimize \(F\).

  • Orthonormality Loss

    As B should act as a basis, we need a regulation loss to prevent it from collapse: \(E[(B^{\top}(s,a)B(s',a'))^2-||B(s,a)||^2-||B(s',a')||^2]\)

  • Vanilla Policy Loss for the actor