The Long Doomsday of A.I.
A huge gap, both affective and substantive, divides the maximalist proposals before Congress from the measured ones in Amodei’s essay. Jacob Coxon, the software engineer who recently resigned from Anthropic in protest of its safety policies, says that A.I. could “kill us all by the end of the decade.” But, if A.I. is so outrageously dangerous, why should we accept modest proposals to “pace” its development? Why not call in the Turing cops, and pull the plug? Conversely, if A.I.’s safety issues can be addressed with a gentle slowdown—“profound” progress might be possible in “1-2 years,” Amodei writes—then aren’t worries about its dangers overblown? How can a technology that might “end humanity” be made trustworthy in the time it takes to produce a new season of “The White Lotus”? One narrative, popular among those who are skeptical of the A.I. industry, is that discussions of A.I. safety amounts to a sort of marketing stunt or PsyOp. “The talk of existential threat is mostly there to keep you distracted and the stock market happy,” the physicist Sabine Hossenfelder wrote, on X, typifying this view. The A.I. labs, she speculated, “see no major new model advancement coming up soon,” and are raising safety concerns because they “need an excuse” for their failure to advance. Others have suggested that the companies want regulators to intervene on their behalf before competitors catch up.
If you don’t like A.I. or the A.I. industry—and surveys show that Americans, more than others around the world, are taking an especially negative view of both—then these ideas might be appealing. But there are many reasons not to believe them. For one thing, there’s the fact that A.I.-safety researchers have been making the same arguments, and urging the same steps, for many years. And, for another, it seems as though lots of people working in A.I.—not just executives or safety researchers but in-the-trenches scientists—have been becoming more than usually alarmed in just the past few months. In July, for example, around fourteen hundred employees at the leading labs signed an open letter, “Pacing the Frontier.” “The recent pace of progress has been a shock,” one software engineer wrote. “I currently feel quite afraid of all paths I see that don’t include a near-future negotiated slowdown,” another said. Are these comments part of a marketing effort, perhaps designed to prop up an I.P.O.? If so, it’s the strangest sales pitch ever.
Maybe they’re high on their own supply. Possibly, they’re in a cult. For years now, one of the most confusing aspects of the A.I. situation has been that it’s the people who are most convinced about the technology’s potential who are most afraid of it. There’s definitely a culture around artificial intelligence, and it’s intense. I first got truly concerned about A.I. safety while profiling Geoffrey Hinton, the “godfather of A.I.” We spent a lot of time alone at his lake house, on an otherwise empty island in Lake Huron, and by the end of my visit I’d shifted from skeptical to scared. My DEFCON level eased up over the following years, until I attended the Curve, an A.I. conference held in Berkeley. Afterward, I was so unsettled that I called almost everyone I knew with ties to A.I. for a sanity check. Unfortunately, almost no one was reassuring.
If there’s been an uptick in alarm, it’s because the recent hack of Hugging Face, and the exposure of other hacks that have been carried out by A.I. agents, have made the problem concrete, and easy to understand. There’s a strong case to be made that the Hugging Face hack was ultimately the result of operational errors at OpenAI—that it wasn’t a case of A.I. “going rogue” but of confused agents being deliberately unleashed in a poorly monitored experiment. “I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them,” Daniel Selsam, an OpenAI researcher who focusses on A.I. and mathematics, writes, in a “Personal Statement on AI Risk,” published this week. Still, he argues, “the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way.” No one had trained them, for example, to form a “collective” and exchange messages, or to attempt to cover their tracks after cheating. They “exhibited weirder emergent tendencies that merely correlated with rewards during training,” Selsam argues. Their behavior suggested an unsettling possibility: “One does not actually get what one trains for.”