2026-07-20
https://arxiv.org/pdf/2604.24532
Multi-Objective RL & Reward-Free RL:
- For reward function obtained from the environment \(r'_1,\dots,r'_t\in R^d\) and a preference vector \(\lambda \in \Lambda\), we compute the true reward by \(r_t=\lambda^{\top} r_t'\).
- We hope to get a model that, given any preference, directly outputs a policy to fit that preference.
- Preference-conditioned MORL solves this problem by directly training on sampled preferences.
- Reward Free RL tries to explore and learn all kinds of reward functions such that it could fit any unseen reward functions in evaluation. MORL is sort of a subset of RFRL.
- However, RFRL is just too board for a small set of objectives.
- The objective of this paper is to find a policy that shares weight among \(\lambda\) such that \(\pi_{\lambda}\) reaches rather good reward for preference \(\lambda\).
Methodology
Successor Measure
Define successor measure \(M^{\pi}(s,a,s',a')\) as the expected discounted number of times that we will be in state \((s',a')\) in the future when currently in state \((s,a)\) for policy \(\pi\).
The successor measure breaks the \(Q\) function into \(\sum {s',a'}M^{\pi}(s,a,s',a')R(s',a')\).
This separates the dynamic estimation and value estimation.
Forward-Backward for reward-free RL
We use low-rank decomposition to approximate \(M^{\pi_z}(s,a,s',a')=F_{\theta}(s,a,z)^{\top}B_{\omega}(s',a')\), where \(z\) is a vector featuring the preference for the future action. Both vectors are of width \(d_z\).
In this method, \(B\) acts as a \(d_z\)-dimension basis, representing some valuable latent features for future states & actions, while \(F\) weights these features (or, how many times we will encounter them).
The \(z\) vector could be obtained from reward function \(R\) by \(z_r=E_{(s,a)\sim D}[B_\omega(s,a)R(s,a)]\), with Q-value being \(Q(s,a,z_r)=F_{\theta}(s,a,z_r)^{\top}z_r\) and actor \(\pi_{s,z_r}=\arg\max Q(s,a,z_r)\).
Preference-Guided exploration
- Vanilla reward-free RL samples \(z\) from gaussians, which is too board for multi-objective RL learning.
- The paper solves this by:
- Warm-up the replay buffer by \(z\sim N(0,I)\) and repeat the following process.
- Sample a mini batch from replay buffer.
- Calculate \(z_{\lambda}\) based on the mean \(B_{\omega}(s,a)r\lambda\).
- Train the actor critic network.
- The \(z\) is sampled from true reward and preference, which is more task-specific. On the other hand, the noise introduced by mini batch sampling increases variety.
Loss Design
Measure Loss
First, we need to make sure the forward-backward measure is consistent with Bellman function. Thus we use the standard Q-loss to optimize \(F^{\top}B\).
Second, we optimize the dynamic prediction by loss \(-2E[F(s_t,a_t,z)^{\top} B_{\omega}(s_{t+1},a_{t+1})]\).
Q Loss
Pseudo reward is used in RFRL due to the blindness to reward. However, we could just use the standard Q-loss using the real \(\lambda^{\top}r\) in MORL setting to optimize \(F\).
Orthonormality Loss
As B should act as a basis, we need a regulation loss to prevent it from collapse: \(E[(B^{\top}(s,a)B(s',a'))^2-||B(s,a)||^2-||B(s',a')||^2]\)
Vanilla Policy Loss for the actor