Skip to content
Changda Tian

机器人日报

机器人论文日报 — 2026-08-29

5 篇论文

🤖 Scanned 85 new arXiv papers (cs.RO / eess.SY / cs.LG, last 48 h), picked 5 for modular & legged robotics — summarized by DeepSeek.
🤖 扫描了近 48 小时 arXiv(cs.RO / eess.SY / cs.LG)的 85 篇新论文,围绕模块化与足式机器人精选 5 篇 — 由 DeepSeek 生成双语摘要。

1. Task-space model-based control of pneumatic soft actuators

气动软体驱动器的任务空间模型预测控制

Figure from 2608.27186

Authors / 作者: Nithin S. Kumar, Joshua Gaston, D. Caleb Rucker, Eric J. Barth
arXiv: 2608.27186 · PDF

This paper presents a real-time task-space feedback and estimation framework for soft actuators using a non-minimal coordinate discrete elastic rod model in absolute coordinates with holonomic constraints. A quasi-static feedforward inverse model is combined with a task-space PI controller and a dynamic observer that fuses measurement residuals as virtual forces, enabling full-state estimation from sparse sensing. Experiments on three planar pneumatic soft actuators show 1.5–2.3 mm RMSE for precision motions and 5.5–12.4 mm RMSE at 1–2 Hz, demonstrating real-time high-precision moderate-bandwidth control.

中文摘要: 针对软体驱动器强非线性、分布变形与动力学不确定带来的闭环任务空间控制难题,本文提出基于非最小坐标离散弹性杆模型的实时任务空间反馈与估计框架。模型采用绝对坐标与完整约束,既保留分布力学特性,又借助稀疏系统矩阵实现高达10根离散杆的实时计算。控制结构由准静态前馈逆模型、任务空间PI控制器和动态观测器组成;观测器将测量残差融合为虚拟力,从而仅凭稀疏传感即可完成全状态估计。在三种平面气动软体驱动器上验证了五个任务,包括绘制数字、周期性轨迹跟踪、跨平台泛化、稀疏感知和实时用户指定参考。结果表明,在1.5–2.3 mm(精细运动)与5.5–12.4 mm(1–2 Hz)的均方根误差下,该方法可实现实时、高精度、中带宽控制,证明了结构化非最小动态模型对柔性机器人控制的实用价值。

💬 Though demonstrated on soft actuators, the task-space dynamic-model-based control architecture and observer design are directly transferable to whole-body control of modular legged robots.
💬 尽管在软体驱动器上验证,其基于任务空间动态模型的控制架构与观测器设计可直接迁移到模块化腿足机器人的全身控制。

Why read it / 推荐理由: Offers a practical template for combining simplified dynamic models with task-space feedback/observation for real-time legged robot control. 为腿足机器人实时控制提供了将简化动力学模型与任务空间反馈/观测相结合的实用范例。


2. Riemann-1.0: An Embodied World Action Model for Physical AI

Riemann-1.0:面向物理AI的具身世界动作模型

Figure from 2608.27033

Authors / 作者: Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu, Yaokun Li, Boyi Jiang et al.
arXiv: 2608.27033 · PDF

Riemann-1.0 is a fully causal autoregressive World Action Model that jointly models multi-view visual observations, robot states, and embodiment-specific actions in a unified sequence. It unifies robot policy execution and action-conditioned world simulation in one model, and uses a progressive embodied pretraining framework to learn from egocentric human videos, handheld-gripper demonstrations, and heterogeneous robot trajectories. It achieves state-of-the-art success rates on RoboTwin2.0, LIBERO, and RoboCasa-365, demonstrating its dual role as a policy and a world simulator.

中文摘要: Riemann-1.0提出了一种全因果自回归世界动作模型(WAM),将多视角视觉观测、机器人状态与具身相关动作统一建模为因果序列,以状态转移形式表示动作与世界演化。与先前的联合生成、视频优先或解耦建模不同,该模型在单一网络中同时支持在线机器人策略执行与动作条件世界仿真,可同时充当可执行策略与多具身视觉世界模拟器。为实现跨异构数据源的规模化学习,作者设计了渐进式具身预训练框架,在共享的世界动作建模目标下统一学习自我中心人类视频、手持夹爪示教和异构机器人轨迹,并基于20万小时以上交互数据逐步迁移具身经验。在RoboTwin2.0、LIBERO与长程组合基准RoboCasa-365上分别达到94.3%、99.0%与62.6%的成功率,超越此前专用模型。该范式为腿足机器人提供了“策略+仿真器”统一建模的新方向。

💬 A unified world-action model that can simultaneously act as a legged locomotion policy and a learned simulator is attractive for modular legged systems with heterogeneous embodiments.
💬 统一世界动作模型可同时充当腿足运动策略与学习仿真器,对具身异构的模块化腿足系统尤为有吸引力。

Why read it / 推荐理由: Demonstrates how to jointly use heterogeneous embodied data to train a model that is both a policy and a simulator, applicable to modular legged robots. 展示了如何联合异构具身数据训练一个兼具策略与仿真器功能的模型,适用于模块化腿足机器人。


3. CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

CLAP:跨具身视频世界模型作为零样本物理模拟器

Figure from 2608.27406

Authors / 作者: Kechen Liu, Ola Shorinwa
arXiv: 2608.27406 · PDF

CLAP is a framework for cross-embodiment action-conditioned video generation trained on diverse internet-scale videos of human and robotic agents. It reconciles disparate action spaces using end-effector poses, language instructions, and latent actions, then uses curriculum-based learning to first learn physical priors from unlabeled video and later ground them for zero-shot deployment. It approaches or surpasses single-embodiment video models in DROID and provides a foundation for using world models as physical simulators.

中文摘要: CLAP提出了一种跨具身动作条件视频生成框架,可在包含人类与多种机器人智能体的互联网规模异构视频上训练。其核心创新包括:(1) 用末端执行器位姿、语言指令和潜动作统一不同机器人平台的动作表示;(2) 通过课程式跨具身学习,先从无标注视频中学习通用物理先验,再将这些先验锚定到具体末端执行器动作空间,实现零样本部署到真实任务。在DROID等困难环境中,CLAP接近或超越了单具身视频模型的性能。该工作表明,通用物理规律可跨越不同执行器进行迁移,为以视频世界模型作为机器人零样本物理模拟器提供了新思路,对腿足机器人基于学习的动力学建模与sim-to-real具有借鉴意义。

💬 A promising route to learning embodiment-agnostic physical priors from video, which could help sim-to-real transfer for legged robots without platform-specific dynamics models.
💬 该工作提供了一种从视频学习与具身无关的物理先验的路径,有望在没有平台特定动力学模型的情况下助力腿足机器人的sim-to-real迁移。

Why read it / 推荐理由: Relevant for using large-scale heterogeneous video data to build reusable world models for legged locomotion simulation and policy learning. 适用于利用大规模异构视频数据构建可复用的腿足运动仿真与策略学习世界模型。


4. GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

GRAFT:面向细粒度机器人操作的基于锚点的高效在线强化适应

Figure from 2608.27079

Authors / 作者: Yibo Qiu, Haoliang Ye, Shu’ang Sun, Zan Huang, Ronald X Xu, Mingzhai Sun
arXiv: 2608.27079 · PDF

GRAFT is a framework for efficient online adaptation of pretrained vision-language-action (VLA) policies to fine-grained tasks. It learns view-specific visual anchors from region-level supervision to focus perception on task-relevant local cues, and combines single-step action generation with cached visual-language prefix reuse to reduce computational overhead. Across four biomedical manipulation tasks, GRAFT improves success by 25 percentage points under matched adaptation budgets while cutting online update cost.

中文摘要: GRAFT面向预训练视觉-语言-动作(VLA)策略在细粒度任务上的高效在线适配问题。由于任务成功往往依赖细微且与视角相关的视觉线索,而任务级奖励难以提供区域级指导,本文引入区域级监督学习视图特定的视觉锚点,使感知聚焦于关键局部线索,且部署时无需区域提议。同时,GRAFT采用单步动作生成与缓存视觉-语言前缀复用,显著降低在线学习与推理的计算开销。在四个生物医学操作任务上,匹配适应预算下成功率提升25个百分点,并降低在线策略更新的计算负担。该框架的核心思想—利用结构化感知加速在线强化适应—对腿足机器人从仿真到真实环境的快速自适应控制具有直接参考价值。

💬 The grounded perception and cached-prefix techniques provide an efficient online RL adaptation recipe that can speed up sim-to-real fine-tuning of legged locomotion policies.
💬 其基于锚点感知与缓存前缀的技术为腿足运动策略的sim-to-real快速微调提供了一种高效的在线强化适应方案。

Why read it / 推荐理由: Directly useful for adapting pretrained policies to new legged platforms or terrains with limited real-world interaction. 对在有限真实交互下将预训练策略适配到新腿足平台或地形具有直接实用价值。


5. FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference

FlashVLA:用于快速异步VLA推理的流式动作解码

Figure from 2608.27384

Authors / 作者: Zekai Li, Jiaming Tang, Zhijian Liu
arXiv: 2608.27384 · PDF

FlashVLA is a streaming action decoding framework that enables fast and asynchronous inference for flow-matching-based vision-language-action (VLA) models. It keeps a streaming action buffer with multiple chunks at different noise levels and decodes them with chunk-wise causal attention, producing one executable action chunk per inference step while preserving action continuity. Experiments show it can reach ≥30 Hz control frequency on a single GPU with smooth asynchronous real-world operation.

中文摘要: FlashVLA针对流匹配类视觉-语言-动作(VLA)模型推理延迟高、异步执行不稳定等问题,提出流式动作解码框架。该方法维护一个包含不同噪声水平多个动作块的流式缓冲区,并采用块级因果注意力进行解码,每个推理步产出一个可执行动作块,同时通过块级自回归隐式保持动作连续性,无需额外未来状态条件即可实现平滑异步执行。在大量仿真与真实实验中,FlashVLA显著提升推理速度,并在单GPU上达到≥30 Hz控制频率,同时保持任务性能。对于腿足机器人,若采用VLA或类似条件动作生成模型,该低延迟流式解码机制有助于提高控制频率与实时性。

💬 The streaming decoding design directly addresses the inference-latency bottleneck of learned action models, enabling high-frequency control for legged robots.
💬 其流式解码设计直接解决学习型动作模型的推理延迟瓶颈,有助于实现腿足机器人的高频控制。

Why read it / 推荐理由: Valuable for deploying large learned policies on legged robots where control frequency and latency are critical. 对于控制频率与延迟至关重要的腿足机器人,该工作对部署大规模学习策略很有价值。


← 全部日报

评论