Robotics Digest
Robotics Paper Digest — 2026-08-14
🤖 Scanned 282 new arXiv papers (cs.RO / eess.SY / cs.LG, last 48 h), picked 5 for modular & legged robotics — summarized by DeepSeek.
🤖 扫描了近 48 小时 arXiv(cs.RO / eess.SY / cs.LG)的 282 篇新论文,围绕模块化与足式机器人精选 5 篇 — 由 DeepSeek 生成双语摘要。
1. HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments
HumanoidVLN:面向多种人形形态的物理仿真视觉-语言导航模拟器与基准
Authors / 作者: Quan-Dung Pham, Anh Dao, The-Anh Nguyen, Minh Nguyen-Dinh, Phuong Nam Dang, Tri Pham et al.
arXiv: 2608.12860 · PDF
Presents HumanoidVLN, a physics-grounded Isaac Sim benchmark for vision-language navigation across four humanoid robots (Unitree G1/H1 and two internal platforms, 10–12 lower-body DoF, heights 1.17–1.80 m). A hierarchical control stack combines a reinforcement learning locomotion policy with interchangeable PD or MPC path trackers, and the platform integrates with several VLN models. The benchmark provides 933 collision-aware episodes with fine-grained and stylistic instruction variants.
中文摘要: 提出了 HumanoidVLN,一个基于 NVIDIA Isaac Sim 的物理仿真视觉-语言导航(VLN)基准,支持多种人形机器人形态(Unitree G1/H1 及两个内部平台,高度 1.17–1.80 米,下肢 10–12 自由度)。系统采用分层控制栈:强化学习步态策略与可互换的 PD 或 MPC 路径跟踪器相结合,并且兼容 NaVILA、DualVLN、StreamVLN、JanusVLN 等主流 VLN 模型。环境包含艺术家设计的场景和 3D 高斯泼溅重建,筛选出超 100 平方米的可通行区域;通过生成-评审-改写多智能体管线并加入人工验证,得到 933 个碰撞感知参考轨迹,每个轨迹配有一条细粒度指令和三种语体(正式、自然、随意)变体。实验在四模型四机器人上验证了平台的可扩展性和物理一致性,为跨形态人形机器人导航研究提供了统一的评测工具,也为足式机器人的视觉-语言导航和 sim-to-real 研究提供了参考。
💬 Strong integration of RL locomotion policies with PD/MPC trackers in a physics simulator, directly relevant to legged robot control stacks.
💬 将强化学习步态策略与 PD/MPC 跟踪器在物理仿真中深度集成,对足式机器人控制栈具有直接参考价值。
Why read it / 推荐理由: It shows a practical hierarchical legged-locomotion control design (RL + MPC/PD) evaluated across multiple humanoid embodiments in a physics simulator. 它展示了在物理仿真中跨多种人形形态验证的 RL+MPC/PD 分层足式运动控制设计。
2. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
HumanTracker:迈向全面且符合人类感知的运动跟踪基准

Authors / 作者: Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang et al.
arXiv: 2608.13555 · PDF
Introduces HumanTracker, a large humanoid motion tracking benchmark with ~153 hours of optical motion trajectories from professional performers across four motion families, plus HumanScore, a preference-aligned evaluation metric trained on 12K motion pairs (24K motions). HumanScore predicts human preferences better than kinematic errors and reveals contact/stability failures such as foot skating and mistimed touchdowns.
中文摘要: HumanTracker 旨在让人形机器人运动跟踪评估既符合人类感知又可扩展。基准包含约 153 小时来自多位专业表演者的光学运动数据,按四个动作族组织并带有文本标签,用于细粒度诊断;同时提出 HumanScore 指标,在 12K 个运动对(含 24K 个运动片段)上训练,对齐人类偏好。与常见的平均逐帧位姿误差相比,HumanScore 能更好预测人类主观评判,并暴露运动学指标容易忽略的接触与稳定性伪影,如脚底滑动、触地时机错误等,为全身控制/遥操作跟踪器的评估提供更可靠的度量。
💬 A perception-aligned evaluation metric focused on foot contact and stability, highly relevant for whole-body control and teleoperation of legged robots.
💬 提出面向脚部接触与稳定性的感知对齐评估指标,对足式机器人全身控制与遥操作非常相关。
Why read it / 推荐理由: It addresses the common mismatch between kinematic tracking metrics and physical quality in legged motion, providing a benchmark and metric that better reflect real control performance. 它解决了足式运动评估中运动学指标与物理质量脱节的常见问题,提供了更能反映真实控制性能的基准与度量。
3. Weight Certificates for Convex Multi-Objective MPC: Geometric Characterization, $\ell^1$ Construction, and $\ell^2$ Foreclosure
凸多目标MPC的权重证书:几何刻画、ℓ1构造与ℓ2不可行性

Authors / 作者: Hadi Hajieghrary, Benedikt Walter, Chaitanya Shinde, Miguel Hurtadoand Jerry Lopez
arXiv: 2608.12520 · PDF
Shows when a weighted-sum MPC exactly reproduces lexicographic multi-objective behavior, providing a geometric characterization via normal-cone slices and a linear program to compute certified ℓ1 weights; for squared-hinge penalties, finite weights are provably inexact when limiting multipliers are non-zero. Closed-loop nuPlan experiments on a 25-rule rulebook show improved legal-tier event precision and near-equal tier components.
中文摘要: 本文研究凸多目标模型预测控制(MPC)中权重和与字典序优化的等价条件。作者证明,当且仅当增广单位性能系数的权重向量位于成就映射上图像的 outward normal cone 切片时,加权和才能精确复现字典序最优解;在铰链损失下,该切片为多面体,可通过线性规划求得带认证裕度的内部权重(ℓ1 构造),而在平方铰链损失下,若极限乘子非零,则不存在有限精确权重,违反程度沿局部极小值分支以 O(1/w) 衰减。基于 25 条规则的闭环 nuPlan 实验表明,与启发式分离权重相比,认证权重在 11 类场景中的 9 类中获得近似相等的权级分量,并将法律级事件精确率提高约一倍。该结果对使用优先级或加权组合的多目标运动控制/MPC 设计具有直接指导意义。
💬 Provides a rigorous method to certify MPC weights against a lexicographic task hierarchy, a key concern in whole-body control of legged robots.
💬 为 MPC 权重相对于字典序任务层次提供严格的认证方法,这是足式机器人全身控制中的关键问题。
Why read it / 推荐理由: It offers a theoretically grounded way to choose or validate cost weights in multi-objective MPC, which is directly applicable to legged locomotion controllers with prioritized tasks. 它为多目标 MPC 中代价权重的选取与验证提供了理论依据,直接适用于带优先级任务的足式运动控制器。
4. Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning
Temporal GRPO:超越轨迹级信用分配的视觉-语言-动作强化学习

Authors / 作者: Yao Zhou, Hang Gao, Fengge Wu, Changwen Zheng, Wenwen Qiang
arXiv: 2608.13026 · PDF
Identifies trajectory-level credit aliasing in outcome-driven RL, where a single rollout-level advantage penalizes all actions including those enabling early progress. Temporal GRPO constructs detectable task stages, aligns rollouts by stage, and computes stage-specific advantages applied only to the corresponding action intervals; experiments on RoboTwin 2.0 show improved success and sample efficiency.
中文摘要: Temporal GRPO 指出基于结果奖励的强化学习(如 GRPO 微调 VLA 策略)中常见的轨迹级信用分配混淆问题:一个 rollout 级 advantage 被应用于整条轨迹的所有动作,导致已完成多个有效阶段但最终失败的策略被过度惩罚。为此,该方法将任务划分为可检测的阶段,将不同 rollout 按进入的阶段对齐,仅对同一阶段的 rollout 计算阶段优势,并把优势只作用于对应阶段的动作区间。在 RoboTwin 2.0 上,Temporal GRPO 提高了任务成功率和样本效率,并在不同任务时域上保持一致提升;在 LIBERO-Long 上的受控更新保留了共享的先决阶段,改进集中于 rollout 结果最早出现分歧的阶段。该方法为稀疏奖励下机器人策略的信用分配提供了更细粒度的方案,可迁移到足式机器人运动控制的强化学习训练中。
💬 A generic RL credit-assignment improvement that can be applied to legged locomotion RL training with sparse or staged rewards.
💬 一种通用的强化学习信用分配改进方法,可应用于具有稀疏或分阶段奖励的足式运动 RL 训练。
Why read it / 推荐理由: It solves a subtle but common problem in outcome-driven RL—penalizing earlier useful actions due to later failure—which is highly relevant to training legged locomotion policies from sparse task success. 它解决了结果驱动 RL 中一个常见但细微的问题——因最终失败而错误惩罚早期有效动作,这对从稀疏任务成功信号训练足式运动策略非常有价值。
5. Genetic Fuzzy System-Based Multi-Robot Coordination for Planetary Missions
基于遗传模糊系统的行星任务多机器人协调

Authors / 作者: Daegyun Choi, Donghoon Kim
arXiv: 2608.12755 · PDF
Proposes a decentralized multi-robot coordination method using genetic fuzzy systems for collaborative object transport in unstructured terrain. The approach converts an elevation map into a 2D traversability map via slope analysis, then GA optimizes fuzzy inference systems that generate robot velocities under scenarios including local minima and cluttered environments; trained FIS models are validated in testing environments.
中文摘要: 本文提出一种基于遗传模糊系统的分散式多机器人协同方法,用于非结构化环境下协作搬运物体并最小化总路径长度。首先根据高程图进行地形可通行性分析,将坡度不可通行的区域视为障碍并转换为二维可通行图;然后利用遗传算法优化多个模糊推理系统(FIS),以生成机器人搬运物体到目标位置的速度指令,训练场景包括局部极小、目标靠近障碍物和杂乱环境等。训练后的 FIS 模型被应用到转换后的可通行图测试环境中,在多个场景中验证了方法的有效性。虽然面向通用多机器人系统,但其地形可通行性建模与避障/路径规划思想可直接迁移到多足模块化机器人的 traverse-capability-aware 路径规划中,为行星探测等复杂地形下的多足机器人协同提供了借鉴。
💬 Directly addresses terrain traversability modeling and decentralized multi-robot coordination, useful for modular legged platforms operating on rough terrain.
💬 直接涉及地形可通行性建模与分散式多机器人协调,对在粗糙地形上运行的模块化足式平台具有借鉴价值。
Why read it / 推荐理由: It provides a traversability-map-based planning and decentralized coordination framework that can be transferred to legged robot path planning over uneven terrain. 它提供了基于可通行性地图的规划与分散式协调框架,可迁移到足式机器人在不平地形上的路径规划。