Perception and control workflows
Agents combine visual reasoning with segmentation, camera calibration and dynamics calculations to produce robot actions.
Explore the workflowsProject page · 2026-10
Benchmarking Multimodal Agents for Robot Use
Across Diverse Tasks and Embodiments
A single multimodal agent controls arms, dexterous hands, humanoids, quadrupeds, cars and drones in simulation, using only public observation and action tools. Hidden evaluators score each episode.
01 / Overview
Each task exposes the robot as a set of tools. The agent reads images and state, chooses actions, and receives execution receipts and new observations. Success is judged from evaluator state that the agent never sees.
02 / Embodiments
From dexterous manipulation to whole-body motion, driving and flight. A shared control vocabulary connects 18 embodiments and 20 control profiles across all 84 tasks.
Loading the embodiment atlas…
03 / Benchmark
04 / Results
05 / Key findings
What robot use reveals beyond the score.
Agents combine visual reasoning with segmentation, camera calibration and dynamics calculations to produce robot actions.
Explore the workflowsMotion can continue while task progress stops. Cases expose lost object state, stalled subgoals and mistaken completion judgements.
See the failure mechanismsAstra and Opus organise direct estimation, auxiliary computation and feedback differently in stacking and drone juggling.
Compare control strategies06 / Tasks & Recordings
Hover to preview a recording; click to watch it full size.
07 / Protocol
Control steps are physical simulation steps actually executed, bounded by the task budget. Trace record indices number logged events. The two are different.
Native budgets are kept, with 10 of the mobile-manipulation tasks capped at 2000 steps. The 16 redesigned scenes use approved horizons and world-state-v1 judging.
Step-budget limit, native termination and world-state termination are labeled separately. The model has no give-up tool; if it ends a reply while the episode is still running, it is asked to continue. Infrastructure interruptions are unscored, not counted as failures.
Current runs are marked as failures after 15 consecutive interactions without a control step, or after a per-task cumulative limit of 30, 60 or 120. Some earlier runs used other limits or none; each attempt records its own protocol.
On API stream drops or sporadic image-processing rejections, the same simulation and session are kept and resumed after 10/20/40 s backoff, up to 3 times, without replaying actions.
84 tasks per model across five domains. Available results, recordings and protocols are linked per task; unfinished runs remain unscored.
08 / Citation
@misc{yang2026robotworld,
title = {{RobotWorld}: Benchmarking Multimodal Agents for Robot Use
Across Diverse Tasks and Embodiments},
author = {Yang, Zhiqin and Li, Chenxin and Hu, Xiaomeng and Liu, Yibin
and Huang, Weidong and Sun, Jiankai and Li, Haitao and Wu, Zijian
and Huang, Yuzhi and Huang, Fanding and Sun, Hanwen and Liu, Jiashun
and Tong, Jingqi and Huang, Mingxin and Hu, Shaoli and Huang, Shijue
and Bai, Tianyi and Wang, Xinyuan and Lin, Yunlong and Tang, Zhengyang
and Zhang, Zhexin and Chen, Zhuo and Song, Xierui and Dai, Juntao
and Chen, Boyuan and Ji, Jiaming and Zhan, Fangneng and Hu, Mengkang
and Xue, Wei and Zhang, Yonggang and Hu, Han and Ho, Tsung-Yi
and Guo, Yike},
year = {2026},
eprint = {2610.10409},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2610.10409}
}