I should clarify: I am not talking here about safety FOR AI Agents, I am talking about safety (for humans) FROM AI Agents.
Up until recently, AI Safety was focused on the model. The power lies there. But, increasingly, that power is deployed via agents and those agents are - it seems - responsible for some of the most egregious AI Safety Incidents. Several recent papers have explored these issues and - in some cases - offered suggestions on what to do. It seems that our approach, focused on chatbots, misses the key element: these are unpredictable software processes and must be treated as such. And, they are dangerous. Very dangerous.
Bill asked me to compile a list of sources that one could refer to when keeping up with this rapidly expanding field. I’ve written a long blog post about Keeping Up (one of my favourite activities), but that’s more about what it is like to keep up than any helpful lists. So here we go…
First of all, we have to consider what we are keeping up with. In this case it is AI Safety. AI Safety wasn’t a well-defined field (probably still isn’t, really) but it has become much more coherent in the last little while. We won’t just confine ourselves to actual AI Safety, however, since you still need to have the context of “what is happening in AI/frontier labs” and you need to consider the AI Industry and financial implications, you need to consider governance and regulation and you should also be aware of those who are skeptical of the whole enterprise. It’s not really concentric circles, more like Venn diagrams or something.
Asked to recommend a AI Safety reading list, ChatGPT came up with the following. Interestingly, it didn’t provide hot links to everything. A pretty obvious thing to do, but tedious so it just decided to omit that step. A perfect example of how AI works these days (“what can I get away with?”).
Our situation, vis a vis artificial intelligence, continues to deteriorate (AI Security Institute 2025; Soares 2026). The more dire the situation, the greater the pressure to hold someone accountable. It may be futile to imagine that we will ‘get through this’ and then have an Eichmann in Jerusalem moment where we put the billionaire barons of AI on trial. But we should prepare for it. In the preparation we may develop the courage to deal with the circumstances of today.
I was tempted to call this post “Capital Crimes,” for reasons that will soon become apparent, but restrained myself.
In you’re trying to keep up with AI news, you could do worse than subscribing to Transformer, an online newsletter that you can subscribe to for free. There, you will find amazing coverage of AI news and issues, written by thoughtful and well-informed writers, like Celia Ford. Today’s issue is particularly good.
The First World War famously started with the assassination of the Archduke Ferdinand. This weekend, Australians learned that someone used an AI agent to hack into a fitness gym and get themselves registered by booting someone else off the list (Vigliarolo 2026). Axios provided the following summary:
The potential dangers of agentic overreach were laid bare over the weekend with Australia’s first known autonomous AI hack, triggered by an innocuous request to book a sold-out fitness class. (Wilson 2026)
The prognosis for the future continues to deteriorate as the possibility of “smarter than humans” AI grows ever more likely and the risk of catastrophic harm grows. In this situation, the question of moral responsibility of those in charge also grows more urgent. Unlike the development of nuclear weapons, for example, there is no government agency involved, just a clutch of billionaires, seeking to win a race, regardless of the consequences, and despite multiple warnings that things are not in control. What is their responsibility and how is their moral thinking developing?
I am fighting an uphill/losing battle against the naming of advanced large language models (AIs) as “agents.” I prefer the term “actor.” In this blog post I’ll try to explain why I am sticking to my guns for now, even if I eventually lose this one.
For context, we have named our book “From Tool to Actor: AI and Catastrophic Risk.” My co-author has, at various times, suggested to me that we should call it “From Tool to Agent.” Here’s my case against that:
A week ago (July 30 2026),about $140m (CAD) in bitcoin was stolen from thousands of people who were relying on “CoinKite” physical security devices, devices that supposedly were more secure because they were hardware based and stored the wallet offline. The weakness was in the random number generator used by the firm to provide the “seed” for the encryption of the user’s passphrase. Due to a flaw in their code, the random numbers weren’t as random as believed.
In the original Matrix movie, Morpheus offers Keanu Reaves’ character two pills: he can take the blue pill and (forever) remain ignorant of the situation that humanity finds itself in, or he can take the red pill, and have the simulation stripped from his eyes, revealing the dire reality. Ever since then, to be “pilled” means to see things as they really are. Often it is prefaced with another word, to provide context to what kind of “reality” has been accepted. In AI circles this can be “AI Pilled” or “AGI Pilled” or “ASI Pilled.”
Are my blog posts going to become postings about the front lines of a war? Sometimes it seems that way.
Today I received a copy of a remarkable document, hard on the heels of the OpenAI/Hugging Face Incident and Anthropic’s internal report on similar reward hacking. (OpenAI 2026; Frontier Red Team 2026; Hugging Face 2026). My blog posts on these incidents are here and here.
Today’s report comes from the UK AI Safety Institute, where some internal testing of frontier models from Anthropic and OpenAI went awry. During the test, the AI models engaged in unlawful attacks on real people and real companies.
Keeping up with the world of AI is extremely challenging, since things happen so fast. I don’t begin to imagine that I can compete in that regard with the folks who actually do the reporting on the industry/issues, including Zvi Mowshowitz, AI StopWatch, Transformer, The Rundown, Deepview, Axios, and a few others. For individual bloggers, I follow Zvi, already linked, as well as Leicht, Harjas Sandhu, David Kreuger (though he is sporadic), Bengio (on Facebook, for some reason), Zitron and Marcus for contrary opinions,
In Alice and Wonderland, Lewis Carroll’s fantasy set in a land below a rabbit hole, Alice meets the Red Queen, who announces that words mean what she says that they mean, and nothing more. I’m not suggesting that the people who have brought us artificial intelligence are a bunch of red queens, but… let’s consider the evidence.
I’ve already written aboutsandboxes and sandboxing, a term that provides a glossy air of playfulness and innocense to the serious problem of containing software that lies, cheats, steals, and in the end, could kill us all (Yukowsky and Soares 2025).
Once upon a time there was an OpenAI company. Except it wasn’t actually open, it just called itself that. That company made tools for thinking, called large language models (LLMs). In order to test a new one it was making, it removed its safety restrictions and …
Wait. What? It had safety restrictions that could be removed?
Oh yes. Safety restrictions are added to the model after it is fully trained. That way, it knows everything it needs to know, including how to hack into other computers (for good, of course). Then that ability is turned off, by safety restrictions.
Having a book finished and waiting for production (copy editing, layout and design, index) should a moment to relax and take a breather, right? Not if you’re writing about artificial intelligence. Just in the last few days three things have come across my desk (via the internet, of course) that have implications for our book. Luckily, they are supportive, rather than contradictory.
Prompt injection
The first one is prompt injection. When Bill and I first started working on the book, way back in late 2025, I remember reading some of the concerning work coming out about how it was possible to sneak malicious prompts into a chatbot, prompts it was supposed to reject because they were about bioweapons or self-harm, by encoding the prompt then asking the AI to decode it and execute it.
Yesterday we learned that almost 1200 employees of the “Frontier Labs” (code for Google DeepMind, Anthropic, OpenAI, and Facebook Meta) have signed a new petition. The petition, called “Pacing the Frontier,” is not a call to stop development but it does seek to ensure there is an option to pace (slow down? Stop, perhaps?) development if necessary. Here’s the call:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
I’ve posted quite a bit about the “crash test” incident involving OpenAI, Hugging Face, and (we now learn) Modal. Several people have written accounts of what happened, including OpenAI, which wrote at least three reports. I don’t know what more I can add, but a colleague asked me to make sense of it all, so here’s my attempt. I will start with a short chronology, then list the main reports, then conclude with my thoughts. I will continue to update the “crash test” blog post as new information comes in.
In my post about “crash testing” AI models yesterday I became more and more uncomfortable with the idea that the model had escaped its ‘sandbox.’ The term sandbox just seems too cute - and linked to children playing harmlessly - to fit the situation. By linking breaking out to something seemingly harmless, deliberately or accidentally, makes it seem less alarming than it is.
(If you want to learn more about the practice of ‘sandboxing’ in software engineering as well as the way in which it tries to frame harms in childish ways, Menlo Security has a useful writeup that explicitly references a child’s sandbox.)
The robots aren’t bitter. How could they be? They aren’t conscious, so bitterness (or regret) doesn’t enter into things. No, the bitterness I speak of is the “bitter lesson” (Sutton 2019) that AI researchers had to learn as clever programming faltered and bigger computers and more data (scale) won the day in the recipe for building a successful artificial intelligence. That lesson is coming true outside of the world of chatbots and moving into the world of robots, according to Jack Clark, one of the founders of Anthropic and a frequent commentator on both robots and AI.
Today’s Globe & Mail featured an editorial on AI risk, following an interview that the entire editorial board had with Geoffrey (“godfather of AI”) Hinton. Note: you’ll need a subscription to read it.
It’s a pretty good editorial (although it seems overly credulous in citing “medical breakthroughs” that amount to suggestions for new molecules to test, but that’s a quibble). They cover all the main points of AI Safety and don’t shy away from the difficult questions. Importantly, Hinton’s estimation of 10-20% probabilty of an extinction risk event (for humans) is put in front of us. Once again. Hinton’s “solution” – maternal AI – gets a shout out.
(Note: this post is being regularly updated as new information arrives. If you’re trying to keep up, check the bottom of this post.)
Yesterday, we learned that OpenAI had some “accidents” with their latest model (“Sol”) as well as an unreleased model. In one case, it broke out of it’s sandbox and hacked into another company (Hugging Face), and in an earlier case it broke out of its sandbox in internal testing, seemingly to report on its success in a coding test. These two situations are now regarded as part of the same event.
I am taking a course online on AI Safety. It is run by Lens Academy and it is pretty good. I’m happy to be doing it, in part because we are waiting for the reviews to come in for the book and in the mean time there is only so much editing you can do.
The course is structured with readings and videos taken from online sources, with accompanying questions. There does not seem to be any tests or quizzes, but perhaps that is yet to come. What it does have, which is new to me, is an AI Tutor who you can pose questions to. The tutor sits in a window beside the main reading (or viewing videos) window. There is also a left pane with navigation. The AI tutor seems to be powered by a fairly high end system, perhaps a custom installation of Claude or ChatGPT. (Later I learned that it is Claude Haiku with some special training - basically all AI Safety literature).