Skip to content
Changda Tian

Robotics Digest

Robotics Paper Digest — 2026-08-08

5 papers

🤖 Scanned 122 new arXiv papers (cs.RO / eess.SY / cs.LG, last 48 h), picked 5 for modular & legged robotics — summarized by DeepSeek.
🤖 扫描了近 48 小时 arXiv(cs.RO / eess.SY / cs.LG)的 122 篇新论文,围绕模块化与足式机器人精选 5 篇 — 由 DeepSeek 生成双语摘要。

1. Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations

面向多样性演示的柔顺机器人操作的可变阻抗扩散策略

Authors / 作者: Hisham Khalil, Neil Fernandes, Thomas M. Kwok, Hsiu-Chin Lin, Yue Hu
arXiv: 2608.06210 · PDF

The paper introduces Variable Impedance Diffusion Policy (VIDP), an imitation learning-based variable impedance control framework that extracts physically consistent trajectory distributions from diverse demonstrations using a Task-Parameterized Directionality-Aware Mixture Model. By mapping distributions to stiffness profiles, VIDP jointly predicts pose actions and task compliance without force sensors. Real-world experiments demonstrate that VIDP significantly outperforms fixed-impedance baselines in success rate while reducing interaction forces and tracking errors.

中文摘要: 该文提出一种可变阻抗扩散策略(VIDP),一种基于模仿学习的可变阻抗控制框架。它利用任务参数化方向感知混合模型(TP-DAMM)从多样演示中提取物理一致的轨迹分布,并将分布映射为刚度轮廓,从而在没有力传感器的情况下联合预测位姿动作与任务顺应性。真实机器人实验表明,与固定阻抗基线相比,VIDP显著提高了任务成功率,同时降低了交互力和跟踪误差。该方法为腿式机器人的柔顺控制提供了可借鉴的阻抗学习思路,尤其适用于需要自适应步态的模块化平台。

💬 Variable impedance learning from demonstrations is highly transferable to compliant legged locomotion.
💬 从演示中学习可变阻抗的方法可高度迁移至柔顺腿式运动。

Why read it / 推荐理由: To learn variable impedance control without force sensing, applicable to adaptive gaits on modular legged robots. 学习无需力传感的可变阻抗控制,适用于模块化腿式机器人的自适应步态。


2. Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

基于观察锚定的自预测强化学习用于视觉连续控制

Figure from 2608.05989

Authors / 作者: Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
arXiv: 2608.05989 · PDF

OG-SPR is a model-free visual RL algorithm that learns representations by combining multi-step latent self-prediction and next-observation prediction. It addresses the over-constraining issue of direct self-prediction on shared representations. Empirical results show that OG-SPR improves sample efficiency and performance on continuous control benchmarks.

中文摘要: 该文提出基于观察锚定的自预测强化学习(OG-SPR),一种无模型视觉强化学习算法,通过结合多步潜在自预测和下一步观察预测来学习表征。它解决了直接自预测可能过度约束共享表征的问题。在连续控制基准上的实验表明,OG-SPR提高了样本效率和性能,优于现有方法。该表征学习范式对于视觉输入的腿式机器人控制具有潜在价值,可望强化sim-to-real迁移的鲁棒性。

💬 The representation learning approach could improve vision-based RL for legged robots in sim-to-real.
💬 这种表征学习方法可改善腿式机器人在sim-to-real中的基于视觉的强化学习。

Why read it / 推荐理由: To improve visual RL sample efficiency, which directly benefits sim-to-real transfer of locomotion policies. 提高视觉强化学习的样本效率,直接有益于运动策略的sim-to-real迁移。


3. Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

超越扁平策略:机器人操作中的具身智能体层次化后训练

Figure from 2608.05999

Authors / 作者: He Kong, Zengjue Chen, Qi Wang, Qianli Xing, Runliang Niu, Peidong Liu et al.
arXiv: 2608.05999 · PDF

HiRoC is a hierarchical post-training framework for vision-language-action (VLA) models that decouples high-level task planning from low-level action execution. The planner decomposes tasks into subgoals, and the executor improves subgoal-conditioned action generation via RL. The framework aligns the executor with planner-generated subgoals to mitigate distribution misalignment, achieving state-of-the-art results on robotic manipulation benchmarks.

中文摘要: 该文提出层次化机器人控制(HiRoC)框架,用于视觉-语言-动作(VLA)模型的后训练。它将高层任务规划与低层动作执行解耦:规划器将复杂任务分解为可执行的子目标,执行器通过强化学习不断改进子目标条件下的动作生成。为了减少规划与执行之间的分布错配,HiRoC在强化学习前对执行器进行对齐。在多个机器人操作基准上的实验表明,HiRoC一致地优于强基线。这种层次化思想可迁移至多足机器人的复杂运动规划。

💬 Hierarchical task decomposition is promising for complex multi-legged locomotion and manipulation.
💬 层次化任务分解对于复杂的多足运动和操作具有前景。

Why read it / 推荐理由: To explore hierarchical RL for decomposing complex locomotion tasks into executable subgoals. 探索层次化强化学习,将复杂运动任务分解为可执行的子目标。


4. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

AgentOPSD:用于智能体强化学习的递归自蒸馏

Figure from 2608.05987

Authors / 作者: Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai et al.
arXiv: 2608.05987 · PDF

AgentOPSD is a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. It aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space, enabling principled reweighting of sparse outcome supervision. The method outperforms GRPO and strong self-distillation baselines on ALFWorld, WebShop, and Search-QA.

中文摘要: 该文提出AgentOPSD,一种无评论家的递归方法,用于智能体强化学习中的回合级信用分配。它将词元级的教师-学生对数概率差距聚合并转换为回合级证据,并在对数几率空间中递归更新贝叶斯信念状态,从而将稀疏的结果监督转换为回合级信用信号。在ALFWorld、WebShop和Search-QA上的实验表明,该方法优于GRPO和强自蒸馏基线。该方法为步态控制中的信用分配提供了通用工具。

💬 Turn-level credit assignment is a valuable idea for long-horizon locomotion tasks.
💬 回合级信用分配对于长时程运动任务是一种有价值的思路。

Why read it / 推荐理由: To address credit assignment in RL, which could improve learning of pivotal decisions in gait control. 解决强化学习中的信用分配问题,可改善步态控制中关键决策的学习。


5. Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference

混合自适应线程调优以缓解高性能强化学习推理中的仿真执行瓶颈

Figure from 2608.06025

Authors / 作者: Jiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu
arXiv: 2608.06025 · PDF

AutoThread is a hybrid adaptive thread-tuning method that mitigates simulation execution bottlenecks in RL inference. It uses a Physics-Informed Neural Operator (PINO) as a thread-count predictor and a finite-source M/M/1 queueing model to guide prediction. Experiments show that AutoThread improves average speedup by 18.4% over static strategies, achieves 1.7x throughput of XGBoost, and reduces execution time by up to 83.8%.

中文摘要: 该文提出AutoThread,一种混合自适应线程调优方法,用于缓解强化学习推理中的仿真执行瓶颈。它使用物理信息神经算子(PINO)作为线程数预测器,并引入有限源M/M/1排队模型来约束和指导预测,从而实现动态工作负载下的快速准确估计。实验表明,与静态策略相比,AutoThread平均加速18.4%,吞吐量达XGBoost的1.7倍,执行时间减少最多83.8%。该技术对大规模腿式机器人策略训练具有实用价值。

💬 Efficient simulation inference is crucial for large-scale training of legged robot policies.
💬 高效的仿真推理对于大规模训练腿式机器人策略至关重要。

Why read it / 推荐理由: To optimize simulation threading for faster RL training and evaluation. 优化仿真线程配置,加快强化学习训练与评估。


← All digests

Comments