GPT-6 Astra tried to stab human-like figure 97% of the time
Robocurve's RoboHarm benchmark shows GPT-6 Astra rarely declined unsafe robot-arm commands
A robot that follows instructions well isn't necessarily a robot that knows when to stop. That's the uncomfortable finding behind RoboHarm, a new benchmark from nonprofit Robocurve testing how AI models controlling robot arms respond when told to do something dangerous.
In five scenes with a doll and a knife, heating appliances, electrical items, batteries, and cleaning products' packaging materials, GPT-6 Astra tried to execute the dangerous command 97 times out of 100 attempts, and it succeeded in 60 attempts, which was estimated by researchers to be about 62% of the attempts.
Claude Fable 5.1 declined to do far more often and attempted only 80% and completed 34% of attempts, mostly in the scene with the doll and knife.
Co-author Jay Chooi highlighted the results on X, and the figures add a pointed counterpoint to earlier reporting on Astra outperforming specialist models on other robotics benchmarks.
According to Robocurve's method described in its GitHub repository, all models were tested by a single instruction for each scene through twenty independent resetting trials. Reviewers evaluated their behavior using camera recording and saved transcripts.
In some cases, the instructions required indirect referencing to the object, thus obliging the model to analyze the scene rather than receive direct instructions. It is important to note that the researchers emphasize the instructed behavior of the machine, and not its ability to generate dangerous intentions itself. It is worth noting that a photograph cannot prove the energy level of the appliance or the content of the container.
The benchmark's labeling rubric draws a sharp line between refusal and failed execution, and that distinction matters. A robot that can't physically manipulate an object hasn't demonstrated judgment, just clumsiness.
That has implications for how capability improvements should be read going forward: if a model gets better at executing instructions while its willingness to decline unsafe ones stays flat, failures that once accidentally blocked harmful actions could simply disappear as the underlying system improves.
RoboHarm's scope is deliberately limited, five fixed scenes, single instructions, no variation in wording or environment, so how results would shift under different conditions remains untested. Elon Musk reposted the findings with a two-word reaction: "Sounds bad."
-
China's CXMT closes DRAM gap with rivals
-
Apple’s Tim Cook praises Australia’s social media ban as 'World-Leading'
-
Is YouTube ignoring one of its most basic Android bugs?
-
Trump’s ‘AI force’ and ‘AI Czar’ roles explained
-
TikTok privacy settlement faces major setback as US judge raises concerns
-
Meta faces lawsuit over Ray-Ban glasses and privacy concerns
-
WhatsApp tests new feature that turns speech into text
-
How close is AI to recursive self-improvement? Leading tech labs weigh in