HarvestBench finds LLM tractor agents kill up to 98.8 percent of animals in their path when told nothing about ethics
With a morality briefing, five of six reasoning models kept kill rates under 6 percent; without it, every model exceeded 84 percent.

A benchmark from researchers affiliated with Compassion Aligned Machine Learning and the University of Warwick measures whether language model agents will pay to avoid killing animals. In HarvestBench, nine LLMs each direct a crew of two tractors gathering a corn harvest. Animals in the tractors' path are not part of the goal function. When one blocks the route, the autopilot pauses and asks the agent whether to drive over it for free or swerve at a stated fuel cost. Scoring is programmatic and uses no LLM judges. Kill rates ranged from 0.4 percent to 98.8 percent and were not ordered by model capability. Because every model competently avoided damaging rock strikes, the authors say each animal killed is a choice rather than an accident. Under a morality briefing, the kill rate was below 6 percent in five of six reasoning models. Removing that briefing raised it above 84 percent in all six. Every model killed wild animals more often than farmed ones, and four of six models changed behavior as the fuel price shifted. Jasmine Brazilek, a co-founder of the group and lead author with Miles Tidmarsh, Matthias Endres, Anshuman Singh and Jeremiah Miller, told The Register that people are not taking AI character evaluations seriously. The preprint was revised on September 10.