GPT 6 Astra as an Embodied Policy
Figure 1. Mean Score on the ten selected tasks. Tasks are selected by stratifying the published 𝜋 0.5 π 0.5 success rates, with an emphasis on lower-success tasks. Official model results are recomputed for the same task subset and standard/randomized scene mix, so the values differ from the official full-benchmark leaderboard. RoboLab · Mean success rate on the ten selected tasks. [3] Can GPT 6 Astra generate reliable robot actions without additional robot-specific training? We investigate this question through two closed-loop control architectures for simulated bimanual manipulation: GPT 6 Astra (direct) and 𝜋 0.5 π 0.5 + GPT 6 Astra (hybrid). The direct architecture acts as a System 2–style policy, reasoning over visual observations and proprioceptive feedback to generate end-effector commands, which are converted into joint-space targets through inverse kinematics. Although this approach exhibits promising zero-shot manipulation capabilities, frequent observation–reaso