On September 10, 2026, Andon Labs posted a video to X showing a drone navigating its San Francisco office, locating a specific person, and following them after a prompt to OpenAI’s GPT-6 Astra. What the highlight reel does not show is how often the system actually works. This kind of demo echoes broader concerns around covert surveillance app development that has drawn scrutiny from researchers and policymakers alike.
What the Drone Actually Did
Five tasks, one controlled office, and a success rate that complicates the headline.
Andon Labs structured the experiment as a benchmark called Drone-Bench. It breaks the surveillance sequence into five discrete tasks: 3D environmental reconstruction, drone localization, navigation, target-person detection, and person-following.
According to the company’s published benchmark documentation, GPT-6 Astra’s strongest submissions beat a human baseline on all five tasks in at least one attempt, earning an overall benchmark score of 0.68. The company also reported an estimated average end-to-end success probability of approximately 2.8%, calculated from separate task-level results rather than observed across complete physical missions.
The hardware was an off-the-shelf quadcopter with facial-recognition capability, according to secondary reporting on the experiment. Andon Labs says it has no Defense Department contracts and is not developing autonomous drone weapons, according to co-founder Lukas Petersson.
Petersson characterized the benchmark as a sequence of relatively simple tasks, noting that expert human drone operators can accomplish far more.
Safety Research or Favorable Publicity?
Two credible interpretations of the same demonstration, and the evidence that separates them.
Andon Labs frames Drone-Bench as transparency-oriented safety research, arguing that policymakers benefit from concrete evidence of what current AI systems can do. The company’s website reportedly states, “Safety from humans in the loop is a mirage,” and describes its broader mission as studying frontier AI deployed in real-world environments.
Emily M. Bender, a computational linguist at the University of Washington and co-author of “The AI Con,” offered a sharply different reading. She argued that portraying commercial AI as exceptionally powerful may benefit the companies that build those models. That framing, she said, can also shift attention away from present-day harms: surveillance normalization, privacy erosion, and data-center resource consumption , concerns amplified by cases of apps secretly tracking users without meaningful oversight.
Her argument is an expert critique and interpretation, not documented evidence of coordination between Andon Labs and any model developer.
Person-following drones are not new technology, according to Peter Asaro, chair of the Stop Killer Robots steering committee and a professor at the New School. The meaningful shift, he said, is the reduced expertise now required to assemble such systems, as generative AI can produce much of the necessary control code. That ease of deployment mirrors concerns raised about automated surveillance systems proliferating in civic environments.
Asaro also raised a pointed accountability concern. When an autonomous system makes an error, he said, a chatbot’s explanation of its own decision may not accurately reflect the underlying computational process that produced it.
The same benchmark that generated a striking, “Terminator”-adjacent video also reported an estimated 2.8% average end-to-end success probability. The distance between a controlled-environment highlight reel and a dependable autonomous surveillance system remains considerable.




























