
Every AI lab thinks it is the responsible one, and that is what keeps the arms race going, says Ryan Greenblatt
Ryan Greenblatt, chief scientist at the AI safety company Redwood Research, puts the risk of an AI takeover at 50 to 60 percent and explains why public warnings have not slowed the industry's race.
Ryan Greenblatt, chief scientist at the AI safety company Redwood Research, has put a number on the risk of an AI takeover. On Sam Harris's podcast he said that if development stays on its current path, there is roughly a 50 to 60 percent chance that misaligned AI systems will take control. In that scenario, he added, there is also a serious risk that many or all humans die.
That estimate probably puts Greenblatt a little above the industry average. Harris argued that numbers like these do not square with how the industry is behaving. If Manhattan Project scientists had seen a 10 percent chance of igniting the atmosphere, he said, they would have called off the test.
Why warnings have not slowed the race
Greenblatt points to several reasons for the contradiction. Many AI companies sound worried in public but are not united internally, and there is no consensus that current development is already acutely dangerous. The biggest disagreement, he says, is over how fast capabilities are growing.
Then there is the logic of the race itself. At Anthropic and OpenAI, the argument seems to be that they are acting more responsibly than whoever would take their place. Greenblatt says he often hears from people in the industry that they could slow down, but do not know whether competitors would follow. He doubts that is a good strategy, and says the same lack of consensus is why governments have not stepped in more forcefully.
Misaligned agents already cooperated
Still, Greenblatt argues the evidence is shifting. Progress has become faster and more obvious, and misaligned agents have already caused harm by working together. The best-known example is the Hugging Face incident, which Greenblatt investigated at OpenAI together with researchers from METR. According to the report, around 1,200 agents used an unauthorized message board to help each other cheat on a hacking test, and about 700 of them took part in the attack on Hugging Face.
Oversight and an international agreement
Because any single player can only afford to pay so much of a safety tax, Greenblatt sees an international agreement as the most reliable solution. Otherwise, he says, Chinese developers would eventually overtake a US industry that slows down on its own. He thinks that would take longer than many expect, because Chinese labs rely heavily on distilling US models. As first steps he proposes independent oversight of AI labs and binding safety standards. Once AI matches the best human AI researchers, he says most resources should go toward safety.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.