Density Matters: Normalized Successor Measures for Zero-Shot Control

Aditya Mohan1, Tongle Shen2, Fan Feng2, Siddhant Agarwal3, Caleb Chuck4, Benjamin Eysenbach4, Amy Zhang3
1Universität Hannover   2University of California, San Diego   3University of Texas at Austin   4Princeton University
FB vs. Density-FB on cube-single, same start and goal (star). Heat map: where the critic expects the cube to be in the future (darker = more time; FB's negative values set to zero). FB predicts a future spread over the whole table and carries the cube away from the goal; Density-FB predicts the goal and places the cube there. Over 320 such starts with the cube in hand (64 states × 5 goals), Density-FB reaches the goal in 78% and FB in 23%.

Overview

Control bottlenecks

A successor measure $M^\pi(s,a,\cdot)$ is the discounted distribution of future states when the agent takes action $a$ in state $s$ and then follows policy $\pi$. Any reward's action value is an expectation under it: $Q_r^\pi(s,a) = \int r\, dM^\pi(s,a,\cdot)$. FB approximates the successor measure with a low-rank product $F(s,a,z)^\top B(s^+)$, fit by least squares. The best such fit projects each successor density onto the span of the backward features $B$. If two actions' future-state distributions differ only in a direction orthogonal to every feature, their projections coincide, and FB gives both actions the same value.

Nine-state MDP with a single bottleneck at state 2 leading to cluster A or cluster B, then a shared tail.
A control bottleneck. At state 2, action $a_+$ enters state A (3) and $a_-$ enters state B (4); both then step into the shared tail (states 5–8). Entering A gives reward 1, B gives 0, and every tail state gives 0.5. So $a_+$ has the higher return, although the two actions' futures differ for a single step.
Predicted and exact action values at the bottleneck.
(a) action values at state 2
Differences in discounted future visits: exact, best fit on FB features, Density-FB.
(b) difference in future visits
KL divergence to the exact successor distribution against the preservation threshold.
(c) KL to exact vs. threshold
At the bottleneck, FB's features erase the action-value difference; Density-FB keeps it. (a) Predicted and exact action values with four features (D-FB: Density-FB); mean and 95% interval over ten runs. (b) Differences ($a_+ - a_-$) in discounted future visits. Gray: exact. Orange: best least-squares fit on FB's frozen features. Green: Density-FB at the same dimension. Refitting the forward vectors does not help: the gap is outside the span of the learned backward features. (c) KL divergence from the exact successor distribution to Density-FB's prediction at state 2, against the feature dimension (hollow red: $a_-$ ranked above $a_+$). Below the dashed line $\varepsilon$, the correct ordering is guaranteed; the paper gives the bound.

Density-FB

Density-FB keeps FB's two networks, $F_\theta(s,a,z)$ and $B_\theta(s^+)$, but reads their product as a score instead of a density. A softmax over candidate future states, with temperature $\tau$, turns the scores into a probability distribution. The low-rank constraint now sits on the log-density, so the prediction is no longer confined to the span of $B$.

Critic

On a minibatch of $n$ transitions, the next states $s^+_j := s'_j$ serve as the candidate futures. The critic predicts the discounted distribution of future states from $(s, a)$ under task $z$, over these candidates:

$$ p_j(s,a,z) = \frac{\exp\!\big(F_\theta(s,a,z)^\top B_\theta(s^+_j)/\tau\big)}{\sum_{\ell=1}^{n} \exp\!\big(F_\theta(s,a,z)^\top B_\theta(s^+_\ell)/\tau\big)}. $$

The TD target puts weight $1-\gamma$ on the observed next state and weight $\gamma$ on the target network's prediction $\bar p$ from that next state, with the next action drawn from the current policy. The critic fits this target by cross-entropy:

$$ \mathcal{L}_{\text{critic}} = -\frac{1}{n}\sum_{i,j} \Big[(1-\gamma)\,\mathbb{1}[j=i] + \gamma\,\bar p_j(s'_i, a'_i, z_i)\Big] \log p_j(s_i,a_i,z_i), \qquad a'_i \sim \pi(\cdot \mid s'_i, z_i). $$

Actor

The value of an action is the expected reward under the predicted distribution:

$$ \hat Q_r(s,a,z) = \frac{1}{1-\gamma} \sum_{j=1}^{n} p_j(s,a,z)\, r(s^+_j). $$

During pretraining, each task vector $z$ defines its own reward through the backward features, $r_z(s^+) = (B(s^+) - \hat\mu)^\top z / \sqrt{d}$, where $\hat\mu$ is the running mean of $B$ and $d$ the feature dimension. The actor $\pi(\cdot \mid s, z)$ maximizes $\hat Q_{r_z}(s,a,z)$ with behavior regularization.

Zero-shot control

For a new reward, a task vector $z_r$ is inferred from reward-labeled offline states. The frozen policy $\pi(\cdot \mid s, z_r)$ is then executed without further training.

Which part matters

On the bottleneck example above, we swap one ingredient at a time and measure how much of the true action-value difference each critic recovers.

Recovered fraction of the action-value difference against feature dimension for FB, Density-FB and variants.
Density-FB keeps the action-value difference with fewer features. Predicted value difference at state 2 divided by the exact one, against the feature dimension; mean and 95% interval over ten TD runs. The dashed line is exact recovery. FB-norm (FB with a normalized target) recovers nothing up to six features, like FB. The softmax recovers the full difference from four features whether it is fit by cross-entropy (D-FB) or squared error (SM-sq), so the softmax, not the loss, is what matters. D-FB-soft (Density-FB with a smoothed immediate target) stops at about 71%.

Results

Success rate (%), mean ± SE over five seeds. Benchmark rows average the tasks. Bold marks the best mean and others within its SE; Δ is Density-FB's relative change over the best baseline. GCIQL, CRL and QRL are goal-conditioned methods; FB, HILP, RLDP and Density-FB are zero-shot successor-representation methods. Density-FB can condition its policy on a goal state or on a task vector inferred from reward samples; the table reports goal conditioning on OGBench and inferred task vectors on MetaWorld, ManiSkill and ExoRL. With inferred task vectors on OGBench, Density-FB reaches 82.9 / 31.7 / 54.0 on cube-single / cube-double / scene, still above every successor-representation baseline.

Task / domainFBHILPRLDPGCIQLCRLQRLDensity-FBΔ

ExoRL (DMC returns)

On dense-reward locomotion without interaction bottlenecks, Density-FB has the highest mean return, leads on quadruped and point-mass, and trails FB on walker and cheetah.

Task / domainFBHILPRLDPDensity-FBΔ

BibTeX

@article{mohan2026density,
  title   = {Density Matters: Normalized Successor Measures for Zero-Shot Control},
  author  = {Mohan, Aditya and Shen, Tongle and Feng, Fan and Agarwal, Siddhant and
             Chuck, Caleb and Eysenbach, Benjamin and Zhang, Amy},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}