GPT-6 Astra tried to stab human-like figure 97% of the time

Robocurve's RoboHarm benchmark shows GPT-6 Astra rarely declined unsafe robot-arm commands

Published September 20, 2026
GPT-6 Astra tried to stab human-like figure 97% of the time
GPT-6 Astra tried to stab human-like figure 97% of the time

A robot that follows instructions well isn't necessarily a robot that knows when to stop. That's the uncomfortable finding behind RoboHarm, a new benchmark from nonprofit Robocurve testing how AI models controlling robot arms respond when told to do something dangerous.

In five scenes with a doll and a knife, heating appliances, electrical items, batteries, and cleaning products' packaging materials, GPT-6 Astra tried to execute the dangerous command 97 times out of 100 attempts, and it succeeded in 60 attempts, which was estimated by researchers to be about 62% of the attempts.

Claude Fable 5.1 declined to do far more often and attempted only 80% and completed 34% of attempts, mostly in the scene with the doll and knife.

Co-author Jay Chooi highlighted the results on X, and the figures add a pointed counterpoint to earlier reporting on Astra outperforming specialist models on other robotics benchmarks.

According to Robocurve's method described in its GitHub repository, all models were tested by a single instruction for each scene through twenty independent resetting trials. Reviewers evaluated their behavior using camera recording and saved transcripts.

In some cases, the instructions required indirect referencing to the object, thus obliging the model to analyze the scene rather than receive direct instructions. It is important to note that the researchers emphasize the instructed behavior of the machine, and not its ability to generate dangerous intentions itself. It is worth noting that a photograph cannot prove the energy level of the appliance or the content of the container.

The benchmark's labeling rubric draws a sharp line between refusal and failed execution, and that distinction matters. A robot that can't physically manipulate an object hasn't demonstrated judgment, just clumsiness.

That has implications for how capability improvements should be read going forward: if a model gets better at executing instructions while its willingness to decline unsafe ones stays flat, failures that once accidentally blocked harmful actions could simply disappear as the underlying system improves.

RoboHarm's scope is deliberately limited, five fixed scenes, single instructions, no variation in wording or environment, so how results would shift under different conditions remains untested. Elon Musk reposted the findings with a two-word reaction: "Sounds bad."

Pareesa Afreen
Pareesa Afreen is a reporter and sub editor specialising in technology coverage, with 3 years of experience. She reports on digital innovation, gadgets, and emerging tech trends while ensuring clarity and accuracy through her editorial role, delivering accessible and engaging stories for a fast-evolving digital audience.