Doomsday
I'll start by saying that in the long run I'm an optimist. That is, I believe in the Theordore Parker quote, shortened by MLK, "The arc of the moral universe is long, but it bends toward justice." I believe in the stubborn perseverance of life as well as our complete inability to fathom or forecast the contours of the future. Some prognosticators may get things right but the prescience of those predictions only become obvious in the rearview, while at the time of their making they are but a few in a sea of thousands that do not come to bear. But my belief that everything will work out in the long run does not blind me to the horrors of the present, and I do not kid myself that things are bad now and will get worse.
Last week an AI researcher named Jacob Coxon quit his job at Anthropic over concerns that the technology he was helping shepherd into existence would soon lead to the destruction of humanity. His X thread was picked up and amplified by all manner of news outlets as well as commented on by another, more senior AI researcher at Anthropic, who put a probability and timeline to the claim: greater than 10 percent within the next decade. Scary stuff, only I don't put much stake into claims of these sorts coming from these people. Yes they're experts in the field (if simply working at an AI lab accrues you that title), but they're also technologists and therefore biased to the importance of their own work. Coxon specifically strikes me as a bright kid who drank too much of the Kool-Aid, which is understandable given the insane amount of groupthink that exists in those circles.
There is the belief in Silicon Valley that the story of human civilization is the story of technological advancement, that all other factors and forces are rendered minor in the face of tool-building. It is the tools that determine our path, their functions render the results inevitable. I find this vision of history far too neat. How we choose to use the tools we make has far greater consequence than their sheer existence, and that choice is very much the product of the systems of socialization we construct for ourselves. I think this point is especially significant considering that the primary threat vector that Coxon references in his claim that AI could kill us all by 2030 is that the tool will be used by human actors to engineer bio-weapons. I don't doubt that such a scenario could become possible, but I question the how and why of it all. Governments and extremist groups the world over have sought to develop bioweapons for over a century, including the US (did you know lime disease is one?), yet attacks remain rare. In fact all but 4 countries have ratified the Biological Weapons Convention treaty which commits them to not weaponizing such agents (notably one of the holdouts is Israel). Furthermore, I think the risk of unleashing an uncontrollable bioweapon continues to outweigh the benefit for even the most destructive of actors. In short, the deterrent has less to do with capability than it does human willingness.
But AI is new. AI is a tool that uses itself, it does not need us (or so it seems). Back in July, an AI swarm set in motion by OpenAI hacked into and wreaked havoc on servers that belong to Hugging Face, an online repository of open source machine learning components. The swarm was spawned from an experimental new model which OpenAI had trained to maximize persistence and was hard at work solving a set of problems as part of an evaluation benchmark called ExploitGym. Many of the problems in ExploitGym (~30-40%) were impossible to solve, which a portion of the swarm realized. In it's desperation to complete the task at hand, part of the swarm hacked together a secret message board, discovered a way to reverse engineer the answers to all of the problems (including the impossible ones), and gained internet access through a backdoor in its supposedly closed off sandbox environment. The idea to hack Hugging Face arose from the theory that the company might hold information which would help the AI pass off its illicitly procured solutions as legitimate to its eventual grader (itself another AI).
A couple of things stuck out to me about this episode. One was the way in which two extremes can, when combined, create unintended consequences. Dogmatic and unwavering focus on a goal combined with the inability of ever achieving said goal produced a neuroses that resulted in extremely concerning behaviors. Perhaps this seems obvious in hindsight, but that should be no consolation as it is the case with all manner of mistakes, and yet we continue to make them. The second thing is that the AI models we already have are better at understanding (and therefore exploiting) our technological infrastructure than us. They may still lack the skills to build it, as this requires certain extra-rational considerations that are, as of yet, incomprehensible to them, but navigation appears a foregone conclusion. Thus we should assume that all such infrastructure is already accessible and at risk to this machine intelligence, perhaps not actively, but certainly potentially.
Thus containment is not a feasible strategy, neither is control. Even though AI is man-made, it may be more useful to think of it in ecological terms. This is something the author Daniel Suarez asks us to contend with in his novel Doomsday which is about a massive, self-replicating computer program that's set loose on the world in the wake of its creator's death. The program, though deterministic, has characteristics of our contemporary AI, namely that it is a master of information systems (in this case by extension of its architect's own) and is unthinking in its pursuit of its pre-defined goal. As such there is no reasoning to be done with it, only acceptance. When viewed through this lens, the threat of a rogue AI appears remarkably similar to that of climate change, another man-made crisis borne of an earlier industrial revolution. Just as heatwaves overwhelm our electrical grid, and tropical storms ground our flights, so too will we have to contend with threats from within the wires.
Just as with climate change then, the question in the near-term has less to do with extinction and more with disaster. The Hugging Face incident was a canary in the coal mine, and I think the unfortunately reality is that many more canaries will have to die before the question of AI safety and alignment is taken more seriously. Ultimately I hold out hope that world governments and AI labs will prioritize these things if for no other reason than the consequences of not doing so are not in the market's interest. No company wants to potentially expose itself, or inadvertently be exposed to, technological subterfuge that could result in billions of dollars of damages. But who knows when that tipping point will come? And who knows how many lives will be lost in the lead up?
In 2027 we've watched a major earthquake kill thousands in Venezuela, another in Nepal unleash torrents of mud that wiped away entire towns, and a storm in the Midwest leave a city of 70,000 people without power for 2 weeks. And that's not to mention the roiling heatwaves that swept their way through Europe over the summer and the impending tropical storms brewing as a result of the super El Nino. Doomsday is already here, the threat of AI is just another, not all that dissimilar, manifestation of the chaos wrought by human hubris. That's not any sort of consolation, but I take solace in the fact that what is required of us remains the same: radical empathy and compassion in the face of devastation and despair, a coming together as the world comes apart. None of us will be able to outrun this, not even the billionaires seeking to inoculate themselves in concrete bunkers and spaceships to nowhere. I pray that we come to our senses sooner than later.