Skip to content
Changda Tian

Robotics Digest

Robotics Paper Digest — 2026-08-21

5 papers

🤖 Scanned 279 new arXiv papers (cs.RO / eess.SY / cs.LG, last 48 h), picked 5 for modular & legged robotics — summarized by DeepSeek.
🤖 扫描了近 48 小时 arXiv(cs.RO / eess.SY / cs.LG)的 279 篇新论文,围绕模块化与足式机器人精选 5 篇 — 由 DeepSeek 生成双语摘要。

1. Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space

学习受限空间中的四足跳跃技能转换

Figure from 2608.19977

Authors / 作者: Zeren Luo, Jiahui Zhang, Yimin Han, Ji Ma, Minghao Lu, Ioannis Havoutis et al.
arXiv: 2608.19977 · PDF

Proposes a hierarchical reinforcement learning pipeline for quadruped jumping through narrow gates. A low-level policy is trained by imitation learning to produce diverse animal-like dynamic skills, while a high-level controller uses vision-based gate detection and capability awareness to select collision-free maneuvers. The framework generalizes to other highly dynamic tasks and is among the first to enable autonomous agile aerial-gate traversal on ground-walking robots.

中文摘要: 提出了一种分层强化学习流水线,使四足机器人能够在狭窄障碍物(如窄门)中完成爆发式运动。底层策略通过模仿学习训练,模仿真实动物行为,形成多种动态技能;高层控制器具备底层技能能力感知,并通过视觉检测获取门的信息,选择合适的无碰撞轨迹进行动态穿越。该框架还可扩展至其他高动态任务。这是地面行走机器人上实现自主、敏捷穿越空中门任务的早期工作之一,使四足机器人展现出接近生物体的敏捷性。

💬 Strong hierarchical RL design pairing animal motion priors with vision-based skill selection for aggressive quadruped locomotion.
💬 分层RL结合动物运动先验与视觉技能选择,为四足激进运动提供了有力范式。

Why read it / 推荐理由: Directly addresses dynamic quadruped locomotion with vision-gated skill transitions, a core challenge in agile legged control. 直接解决四足动态运动中的视觉门控技能切换问题,是敏捷腿足控制的核心挑战。


2. MILD: Tractable Terrain Modeling for Learning Improved Bipedal Locomotion on Deformable Surfaces

MILD:用于学习可变形地面上双足运动的高效地形建模

Figure from 2608.19955

Authors / 作者: Zeren Luo, Jiahui Zhang, Zhe Xu, Wanyue Li, Xinqi Li, Xuechao Chen et al.
arXiv: 2608.19955 · PDF

MILD introduces a physics-grounded discrete-element contact solver that captures spatially varying foot-terrain interactions, overcoming simulator limitations for deformable surfaces. A terrain-aware locomotion controller is trained via deep RL with latent modulation and proprioceptive estimation. Hardware experiments show online terrain identification and adaptation across a range of surface stiffnesses.

中文摘要: MILD提出了一种基于物理的离散元接触求解器,能够准确模拟空间变化的脚-地形交互,弥补了现有模拟器无法表现可变形地面时空异质性的不足。利用该模型,通过深度强化学习并配合潜在调制与本体感觉估计,训练出地形感知的运动控制器。与最先进方法相比,该方法在训练中生成更多样、更真实的接触场景,使控制器在真实可变形地面上展现自然适应能力。硬件实验验证了系统在多种表面刚度范围内的在线地形识别与适应能力。

💬 Tractable deformable-terrain simulation plus terrain-aware RL enables robust adaptation on yielding surfaces, with hardware validation.
💬 可计算的可变形地形仿真结合地形感知RL,使机器人在柔性表面上具备鲁棒适应能力,并经硬件验证。

Why read it / 推荐理由: Provides a simulation and learning recipe for legged locomotion on terrain that most simulators cannot model faithfully. 为大多数仿真器无法忠实建模的柔性地面上的腿足运动提供了仿真与学习方案。


3. Hybrid Feedback Sampling for Sample-Efficient Model Predictive Control

用于高样本效率模型预测控制的混合反馈采样

Figure from 2608.19443

Authors / 作者: Chaoyi Pan, Zeji Yi, John Zhang, Zachary Manchester, Guannan Qu, Guanya Shi
arXiv: 2608.19443 · PDF

FS-MPC improves sampling-based MPC by sampling with an optimized feedback policy, making the proposal distribution optimal and addressing exponential sample growth for unstable dynamics. A hybrid sampling design balances local and global search based on system stability and computation budget. It outperforms MPPI and standard feedback sampling in contact-rich tasks, including real-world humanoid locomotion and manipulation.

中文摘要: 针对采样型MPC在高维、开环不稳定系统中采样效率差的问题,本文提出反馈采样MPC(FS-MPC)。该方法用优化后的反馈策略进行采样,使采样提议分布达到最优,避免随horizon指数增长的采样需求。FS-MPC采用混合采样设计,根据系统稳定性和计算预算在局部搜索与全局搜索之间取得平衡。理论分析表明混合采样比标准MPPI收敛更快、比标准反馈采样最优性更好。在类人机器人移动操作、灵巧操作等多个接触丰富任务以及真实世界的人形机器人运动与操作上,FS-MPC均优于现有方法。

💬 Puts sampling-based MPC on a more solid theoretical footing and demonstrates clear gains on real legged systems.
💬 为采样型MPC奠定了更坚实的理论基础,并在真实腿足系统上展示了明显优势。

Why read it / 推荐理由: MPC is central to legged control; this method makes sampling-based MPC viable for unstable, contact-rich legged platforms. MPC是腿足控制的核心,该方法使采样型MPC可用于不稳定的接触丰富腿足平台。


4. DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

DECOWAM:用于腿足移动操作的解耦全身世界-动作模型

Figure from 2608.20114

Authors / 作者: Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei et al.
arXiv: 2608.20114 · PDF

DECOWAM is a whole-body world-action model that explicitly separates camera ego-motion from base and arm actions via conditional interfaces and adversarially separated latents. It freezes an adapted FastWAM backbone and trains residual adapters with base-velocity conditioning, improving future-video and action prediction. On a real-robot dataset and 79 closed-loop trials per method, it achieves better whole-body coordination and base-displacement robustness than baselines.

中文摘要: 移动操作需要机器人预测腿部运动和手臂运动共同如何改变未来观测与控制。DECOWAM通过专用条件接口将相机自运动与底盘、手臂动作分离,构建全身世界-动作模型。模型冻结适配后的FastWAM主干,训练残差适配器、从特权观测中蒸馏的动作等价未来瓶颈、对抗分离的底盘/手臂潜在变量,并以底盘速度条件化视频预测。在提出的ARMDOG真实机器人数据集和固定重放协议下,相比FastWAM将动作MSE降低21.7%,仅用25.95M可训练参数。在每种方法79次闭环试验中,DECOWAM在全身协调性和底盘位移鲁棒性上均优于对比系统。

💬 Embodiment-aware factorization of whole-body world models yields parameter-efficient joint visual prediction and legged mobile manipulation control.
💬 全身世界模型的身形感知分解,以参数高效方式实现了视觉预测与腿足移动操作控制的联合学习。

Why read it / 推荐理由: Directly targets whole-body control for legged mobile manipulators, bridging visual prediction and action optimization under moving viewpoints. 直接针对腿足移动操作平台的全身控制,在移动视角下衔接视觉预测与动作优化。


5. Video2DoorTraversal: Push Door Traversal via Simulated Door Twins

Video2DoorTraversal:通过仿真门孪生实现推门穿越

Figure from 2608.20251

Authors / 作者: Xincheng Tang, Yiji Chen, Youhan Xie, Wanyu Li, Zhengjie Shu, Lai Jiang et al.
arXiv: 2608.20251 · PDF

Video2DoorTraversal is a single-video real-to-sim-to-real framework for wheel-legged mobile manipulators to perform push-door traversal. DoorTwin reconstructs instance-aligned articulated door twins from one RGB video, and a simulation-in-the-loop agent converts articulation into a parameterized skill program and iteratively refines rollouts. The ArticuACT dual-depth policy achieves 96.57% success across five real doors and 80.95% zero-shot on unseen similar doors, completing traversal in about 13s onboard.

中文摘要: 开门穿越是一项需要精确手柄交互和协调底盘-手臂控制的长时程移动操作任务。本文提出Video2DoorTraversal,一种用于轮腿移动操作器的单视频实到仿到实框架。给定真实门的单段RGB视频,DoorTwin重建实例对齐的、带关节的、可直接仿真的门孪生,包含真实几何和外观。仿真内循环智能体将关节信息转换为参数化技能程序,并通过迭代细化失败轨迹生成可物理执行的演示。这些演示训练ArticuACT双深度策略,以机器人中心相机条件预测底盘、手臂和夹爪命令。系统在五个真实门上平均成功率达96.57%,对结构相似未见门零样本成功率80.95%,平均约13秒完成接近、开门和穿越全流程。

💬 A practical real-to-sim-to-real pipeline for long-horizon loco-manipulation on wheel-legged robots, with strong real-world success rates.
💬 面向轮腿机器人长时程移动操作的实用实到仿到实流程,真实世界成功率很高。

Why read it / 推荐理由: Demonstrates how simulated articulation twins plus onboard dual-depth policy enable robust base-arm coordination for modular legged manipulators. 展示了仿真关节孪生与机载双深度策略如何实现模块化腿足操作器鲁棒的底盘-手臂协调。


← All digests

Comments