Robotics Digest
Robotics Paper Digest — 2026-07-30
🤖 Scanned 254 new arXiv papers (cs.RO / cs.SY / cs.LG, last 48 h), picked 5 for modular & legged robotics — summarized by DeepSeek.
🤖 扫描了近 48 小时 arXiv(cs.RO / cs.SY / cs.LG)的 254 篇新论文,围绕模块化与足式机器人精选 5 篇 — 由 DeepSeek 生成双语摘要。
1. Reinforcement Learning on Cost-Constrained Quadrupedal Hardware
成本受限四足硬件上的强化学习

Authors / 作者: Javier C. Weddington, Bence P. Ölveczky, Stephen A. Baccus
arXiv: 2607.26434 · PDF
This paper addresses the sim-to-real gap for low-cost quadrupedal robots (e.g., Mini Pupper 2) by using a time-aware neural network that incorporates a forward model of actuator delay. The approach learns a central pattern generator (CPG) that produces robust locomotion, even under latency perturbations. Real-robot experiments demonstrate the effectiveness of the method.
中文摘要: 本文针对低成本四足机器人平台(如Mini Pupper 2)因执行器传输延迟和电机反馈噪声导致的仿真到现实迁移难题,提出了一种生物启发的强化学习方法。该方法利用执行器平均延迟的前向模型,并采用时间感知神经网络来处理延迟和噪声,将标准马尔可夫决策过程转化为部分可观测过程。实验表明,该网络学习到了中央模式发生器(CPG),即一种自维持的节律性步态,对高达+320ms的延迟扰动具有鲁棒性。在真实低成本硬件上的测试验证了该方法的有效性,显著缩小了仿真与现实差距。结论指出,时间自组织可能是成本受限运动控制的通用策略。
💬 Directly addresses sim-to-real for quadruped locomotion on cost-constrained hardware with a biologically inspired CPG approach.
💬 直接针对低成本硬件上的四足运动仿真到现实迁移问题,采用了受生物启发的CPG方法。
Why read it / 推荐理由: Presents a practical RL solution for low-cost quadruped robots with real-robot validation and a novel time-aware network. 展示了一种实用的低成本四足机器人强化学习解决方案,包含真实机器人验证和新型时间感知网络。
2. P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning
P3:用于稳定VAE机器人学习的概率策略传播

Authors / 作者: Liyun Yan, Jianming Ma, Yang Zhang, Shengcheng Fu, Zhanxiang Cao, Keqi Zhu et al.
arXiv: 2607.25541 · PDF
The paper identifies and fixes a theoretical issue in VAE-based PPO policies for robot learning, where single-sample approximations cause variance. They propose Probabilistic Policy Propagation (P3), a distribution-aware optimization framework that couples moment-based and sampling-based methods. Experiments on humanoid parkour tasks show improved data efficiency and convergence.
中文摘要: 本文识别并修复了基于变分自编码器(VAE)的PPO策略在机器人学习中的根本性理论问题:单样本近似导致目标函数方差和偏差。为此,提出了概率策略传播(P3)框架,这是一种分布感知的优化方法,结合了基于矩的概率方法和基于采样的校准,以在潜在不确定性下实现稳定高效的学习。在人形机器人跑酷任务上的实验表明,P3将数据效率从64.6%提升至96%以上,并将收敛步数减少20%以上。该方法为VAE-PPO策略提供了坚实基础,对于腿部机器人运动学习的稳定性和效率具有重要价值。
💬 Addresses a fundamental issue in VAE-based policy learning for legged robots, with demonstration on humanoid parkour.
💬 解决了基于VAE的腿部机器人策略学习中的基本问题,并在人形机器人跑酷上进行了演示。
Why read it / 推荐理由: Provides a theoretically grounded improvement for VAE-based RL policies that can be applied to legged locomotion. 为可应用于腿部运动的基于VAE的强化学习策略提供了理论基础的改进。
3. Self-Adaptive Online Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret
用于跟踪未知动力学的自适应在线学习与模型预测控制

Authors / 作者: Atharva Navsalkar, Hongyu Zhou, Vasileios Tzoumas
arXiv: 2607.26370 · PDF
This paper proposes a self-adaptive online learning method for tracking unknown target dynamics, which can be applied to robots tracking moving objects. It learns multiple predictors adaptively and selects the best one, with finite-time near-optimality guarantees. The method is suitable for applications like pursuit-evasion and dynamic mapping.
中文摘要: 本文提出了一种自适应在线学习与模型预测控制(MPC)相结合的方法,用于跟踪未知且可能切换的目标动力学。该方法同时从零开始在线学习多个预测器,并通过自监督方式自适应选择最佳预测器以匹配目标行为,无需事先训练。理论上,该方法在期望意义下具有有限时间近最优性保证:当目标动力学无切换且学习误差为零时,渐近地匹配已知最优的非因果控制策略(即无遗憾);当存在学习误差或切换时,性能优雅退化。该方法适用于动态映射、交通控制、追逃等场景,对于腿部机器人在地形变化环境中自适应步态生成具有潜在应用价值。
💬 Presents a theoretically-grounded adaptive MPC approach that could be applied to gait adaptation in legged robots.
💬 提出了一种理论上严谨的自适应MPC方法,可应用于腿部机器人的步态适应。
Why read it / 推荐理由: Offers a no-regret online learning control method that could enhance adaptive locomotion in changing environments. 提供了一种无遗憾在线学习控制方法,可增强在变化环境中的自适应运动能力。
4. RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning
RLMM-Flow:基于流的移动操作框架与潜在空间强化学习

Authors / 作者: Shuhang Wang, Ziming Li, Hui Cheng
arXiv: 2607.26460 · PDF
This paper introduces a framework that combines flow-based generative policies with latent-space RL for mobile manipulation. It pre-trains a flow policy on demonstrations and then fine-tunes with RL in latent space. Experiments show improved task success and trajectory quality over imitation-only methods.
中文摘要: 本文提出RLMM-Flow框架,用于移动操作中的全身动作生成。该框架首先通过专家演示预训练一个基于流的生成策略,以捕获多模态的全身运动先验;然后冻结该策略,仅在一个低维潜在空间中训练一个强化学习“导向网络”来优化动作。为了稳定高维潜在优化,采用动作空间批评家预热和从粗到细的潜在导向策略。在移动操作运动规划基准上的实验表明,RLMM-Flow在任务成功率、避障和轨迹质量上显著优于仅模仿学习的流策略。该方法为多腿移动操作器的全身控制提供了可借鉴的范式。
💬 Demonstrates latent-space RL for whole-body control, which is relevant for multi-legged mobile manipulators.
💬 展示了用于全身控制的潜在空间强化学习,与多腿移动操作器相关。
Why read it / 推荐理由: Provides a method for learning whole-body control policies that could be extended to legged mobile manipulation. 提供了一种学习全身控制策略的方法,可扩展到腿部移动操作。
5. Two2Four: Generative Quadruped Puppeteering from Human Motion
Two2Four:从人类运动生成四足木偶操控

Authors / 作者: Fatemeh Zargarbashi, Zehong Qiu, Dhruv Agrawal, Stelian Coros, Robert W. Sumner, Martin Guay et al.
arXiv: 2607.26108 · PDF
This paper presents a generative diffusion model that retargets human motion to plausible quadruped motions. It supports various actions and provides fine-grained control like head movement and limb puppeteering. The method improves motion realism and controllability for animation.
中文摘要: 本文提出了一种两阶段生成式扩散模型,用于从普通人类运动数据自动生成逼真且可控的四足动物运动。该方法仅使用四足运动数据进行训练,通过结构化条件注入和修补策略支持行走、奔跑、跳跃、坐、躺等多种动作,并允许对四足动物的头部运动和单腿进行精细操控。与现有重定向方法相比,实验结果表明该方法在运动真实感和可控性上均有提升。虽然主要面向动画和虚拟制作,但生成的自然四足运动可作为腿式机器人运动规划和参考轨迹生成的丰富数据源,尤其适用于需要多样化步态的场景。
💬 While animation-focused, the method generates natural quadruped motions that could serve as reference trajectories for robot control.
💬 虽然侧重于动画,但该方法生成的自然四足运动可作为机器人控制的参考轨迹。
Why read it / 推荐理由: Offers a source of realistic quadruped motion data that may inspire gait generation for legged robots. 提供了逼真的四足运动数据来源,可能启发腿部机器人的步态生成。