Role-Based Reward Decomposition for Legged Locomotion Reinforcement Learning
Changda Tian , Panos Trahanias
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) — Accepted · 2026
Partitioning locomotion rewards by stance and swing roles stabilizes multi-objective reinforcement learning — validated on a real Unitree Go2 and extended to the G1 biped. Accepted at IROS 2026.
Legged Robots Reinforcement Learning Sim-to-Real
@inproceedings{tian2026rolebased,
title = {Role-Based Reward Decomposition for Legged Locomotion Reinforcement Learning},
author = {Tian, Changda and Trahanias, Panos},
booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
note = {Accepted},
year = {2026},
abstract = {Legged locomotion requires optimizing several partially conflicting objectives, including velocity tracking, contact stability, foothold selection, and swing-foot clearance. Standard reinforcement learning aggregates them into a single scalar reward, so one value function must approximate returns that come from heterogeneous hybrid dynamics. This coupling can introduce value interference and destabilize learning. We investigate a role-based reward decomposition framework that partitions locomotion rewards according to stance and swing phases, and we compare single-critic, multi-critic, and role-specialized multi-agent realizations under a protocol varying only the value and actor architecture. When objectives are complementary, organizing the reward by roles is what helps, and the realizations are largely interchangeable. When objectives genuinely conflict, the realizations separate sharply. We induce conflict with a stepping-stones suite whose foothold objectives concentrate reward in the swing role, together with probes on ice, sprinting, and degraded observability. Scalarized training then exploits a dominant minority objective while behavior collapses, per-stream multi-critic normalization preserves the minority objective, and hard actor routing pays a coordination price that a cooperative gated blend avoids. A hybrid probe that grafts a footstep planner onto the swing role reproduces the same behavior at the control level. A learned proprioceptive gate enables sensor-free deployment, which we validate in sim-to-sim transfer and on a real Unitree Go2; a Unitree G1 study extends the conclusions to bipeds. Structuring rewards by physically meaningful roles thus gives an architecture-flexible way to stabilize multi-objective locomotion learning, and preserving coordination is the rule that separates decompositions that help from those that harm.},
}
% TODO: add DOI and page numbers once IEEE Xplore indexes the AIM 2026 proceedings.