RobotWorld · Experiment report 01 · Preliminary

One agent, 84 robot tasks

Arms, dexterous hands, household robots, legged robots, cars and drones. Multimodal models, same tools, no demonstrations, no fine-tuning.

Opus 5.5 plays a full 1v1 drone volleyball match and wins (native termination after 498 control steps).

TL;DR

Each task scores 0 or 1. These statistics use the same result data as the homepage.

RobotWorld asks a simple question: if you give a general multimodal agent a robot as a set of tools, can it get the job done? The agent sees camera images and robot state, sends actions, and gets back execution receipts. A hidden evaluator, which the agent never sees, decides the score.

Models were run in different batches and with different reasoning-effort settings. Final scores follow the prescribed task budgets. We report these settings alongside the results and recordings so that comparisons can be interpreted in context.

How we tested

SetupConfiguration
Tasks84 tasks per model, built on open-source simulators. By embodiment: 24 bimanual or dexterous, 20 mobile manipulators, 14 fixed-base arms, 11 wheeled vehicles, 4 bipeds or humanoids, 4 wheel-legged, 4 aerial, 3 quadrupeds. By domain: 38 manipulation, 20 mobile manipulation, 11 locomotion, 11 driving, 4 aerial.
InterfaceEach robot is exposed as tools: observe (images, joint and task state), act (environment-specific actions, executed as control steps), and receipts. Auxiliary analysis tools are allowed.
ModelsGPT-6 Astra, Opus 5.5, Kimi K3, DeepSeek V4.1 Flash, Gemini 3.8 Flash.
SettingsReasoning effort and non-action budgets vary by model and batch; each selected run retains its settings.
BudgetsNative step budgets are kept, with 10 of the mobile-manipulation tasks capped at 2000 steps. Non-action limits stop an attempt that keeps looking without moving the robot.
EndingThe model has no give-up tool. If it ends a reply while the episode is still running, the control loop asks it to continue. An attempt ends only when the environment terminates or a step / non-action budget runs out.
EvaluationEach task receives a binary score based on predefined success criteria and evaluation budgets.
AttemptsOne selected result per model and task. Infrastructure retries are not additional scored tasks.

A control step is one physical simulation step actually executed. It is different from the index of a logged event.

Key findings

What the trajectories reveal beyond success rates.

  1. Agents build their own perception and control workflows.

    The trajectories include colour segmentation, camera calibration, spatial estimation and dynamics calculations. Agents turn visual reasoning and code into executable robot commands.

  2. Executing an action is not the same as completing a task.

    Failures arise when agents misread the current state, fail to change strategy after feedback, or declare completion before the goal is satisfied. Reliable robot use requires checking what each action actually achieved.

  3. Models take different routes from observation to action.

    In the stacking case, Astra grasps the first block through successive visual estimates before fitting camera geometry; Opus first writes segmentation code and fits camera parameters. In drone juggling, both compute dynamics, but Opus sustains the contact cycle more effectively. The difference lies in how estimation, computation and feedback are combined.

Results

RobotWorld at a glanceBar height = success rate · recordings below; click a tile to play

Each square below is one task. Sorting by result shows how few tasks are solved and how few of Kimi K3's and DeepSeek V4.1 Flash's tasks are solved.

Task outcomesFilled = success · outline = failure · dashed = no result

In the 20 mobile-manipulation tasks, GPT-6 Astra completes four and Kimi K3 completes one; the other models complete none. The chart below reports results across all five domains.

Successful tasks by domainNumber solved / tasks in domain

Which tasks did GPT-6 Astra and Opus 5.5 solve?

Who solved whatGPT-6 Astra vs Opus 5.5

From capability to robot action

The paper traces how observations, computation and feedback become executable control. These cases illustrate mechanisms within individual episodes; the recordings show the motion, while the accompanying analysis draws on the interaction logs.

Changing the command after feedback

Kimi K3 shortened successive drawer pushes as end-effector displacement decreased, then withdrew to check closure; the environment registered success during withdrawal. In charger insertion, Astra responded to a rejected wrist-camera rotation by lowering and repositioning the idle arm before retrying. The rotation then executed, and the episode later completed the insertion. Feedback was useful because it changed the next action.

Turning visual estimates into grasping

In stacking, Astra lifted the first blue block before running external image processing or fitting camera geometry. Opus 5.5 instead wrote colour-segmentation code before moving, then sampled arm motions to fit camera parameters. Both reached first placement and retraction at similar physical steps (270 versus 268). Astra subsequently completed the stack; Opus reached its auxiliary-call budget without completion. This comparison shows different routes to initial progress, not two successful episodes.

Sustaining control after the first contact

Both Astra and Opus 5.5 performed offline dynamics calculations for drone juggling. Opus coordinated levelling, pre-contact acceleration, post-contact thrust reduction and repositioning beneath the ball, reaching 15 valid hits over 800 steps. Astra recorded three hits before termination at step 188. Both used a median of three executed steps per action: building a model and choosing short segments did not by themselves ensure sustained control.

Where the control loop breaks down

A failure tells us that the goal was not reached. The interaction logs help explain what stopped progress: losing the required object state, remaining in preparation, or treating an unfinished task as complete.

Arm motion without recovery of the object state

Kimi K3 knocked the bottle onto its side during approach. Repeated left-arm attempts failed to secure it; at step 240, it switched to the right arm. Further approach motions used the remaining budget, and pouring remained unfinished at step 400. Reaching new arm poses did not restore the grasp needed for the task.

Preparation never advances to the next subgoal

In vegetable chopping, all 69 positive-step calls by Kimi K3 concerned surveying, navigation or clearance around furniture. The final command still attempted kitchen entry. No cutting command appeared before the 2,000-step limit: the bottleneck was reaching the workspace, so this episode does not test the unattempted cutting skill.

Execution continues after corrective search stops

Kimi K3 declared block sweeping complete at step 861, then issued 69 commands changing only the empty arm’s height over the final 139 steps. DeepSeek judged further motion unnecessary in stove navigation and spent the final 105 steps issuing zero base-velocity commands. Both tasks remained unsuccessful. Continued tool use did not mean continued correction of the unfinished goal.

Interaction and budgets

The following charts describe how attempts ended and how much of their control-step budgets they used. These are execution statistics; the stopping condition alone does not explain the underlying failure mechanism.

How each attempt endedCentre = number of tasks solved

How much of the step budget each attempt usedOne dot per attempt · executed control steps / task budget · click a dot to play

Low step usage can reflect early success, environment termination or exhaustion of the non-action budget. Read it together with the outcome and stopping reason.

Time and cost

Explore the recordings

Open a chart item or task card to inspect its recording. Motion recordings can be played; single-frame captures show a still image. Missing or unavailable recordings are labelled separately.

Task gallery →

Limitations