What are we reading?

AI Safety

Keeping up with the world of AI is extremely challenging, since things happen so fast. I don’t begin to imagine that I can compete in that regard with the folks who actually do the reporting on the industry/issues, including Zvi Mowshowitz, AI StopWatch, Transformer, The Rundown, Deepview, Axios, and a few others. For individual bloggers, I follow Zvi, already linked, as well as Leicht, Harjas Sandhu, David Kreuger (though he is sporadic), Bengio (on Facebook, for some reason), Zitron and Marcus for contrary opinions,

I am particularly thankful, these days, for the contributions from AI Stopwatch, which seems so human and poignant and sad. Today (Aug 4 2026) is the day that is memorialized in Ray Bradbury’s short story from 1950 about a smart house, continuing to run, in the aftermath of a nuclear war and the complete abscence of humans.

“Tomorrow still happens, but it can happen without us. The story’s concluding visual is of the smoldering rubble left where the home finally fell to fire during the night:” PDF

I also get suggestions from Bill, who reads the Globe and Mail as well as the New York Times every day. So, what are we reading this week?

First, a few articles on reward hacking, effectively what happens with AIs lie and cheat. Joshua Rothman wrote a readable guide to reward hacking for The New Yorker, prompted by the Hugging Face Incident:

Rothman, Joshua. 2026. “What If We Can Never Trust A.I.?” The New Yorker, August 1. https://www.newyorker.com/culture/open-questions/what-if-we-can-never-trust-ai.

Rothman cites this paper from McDiarmid on reward hacking:

MacDiarmid, Monte, Benjamin Wright, Jonathan Uesato, et al. 2025. “Natural Emergent Misalignment from Reward Hacking in Production RL.” arXiv:2511.18397. Preprint, arXiv, November 23. https://doi.org/10.48550/arXiv.2511.18397.

Grace Huckins, in MIT Technology Review, breaks lyin’ cheatin’ AI down for all of us:

Huckins, Grace. 2026. “Here’s Why AI Agents Lie and Cheat to Reach Their Goals.” August 3. https://www.technologyreview.com/2026/08/03/1141009/heres-why-ai-agents-lie-and-cheat-to-reach-their-goals/.

How we might actually control superintelligence (don’t rely on alignment, insist on guaranteed systems):

Aguirre, Anthony. 2025. “Control Inversion.” Future of Life Institute, November 6. https://control-inversion.ai/. See especially chapter 8 https://control-inversion.ai/8-what-would-control-look-like/

Detail on other software / approach to making that guaranteed system work (as we did for nuclear weapons):

Siu, Tony. 2026. “Aligning the Model Was Never Going to Govern It.” TNW | Artificial-Intelligence, July 30. https://thenextweb.com/news/aligning-the-model-was-never-going-to-govern-it.

Before Yudkowski and Soares, there was Critch and Tsimerman:

Critch, Andrew, and Jacob Tsimerman. 2025. “A Taxonomy of Omnicidal Futures Involving Artificial Intelligence.” arXiv.Org, July 12. https://arxiv.org/abs/2507.09369v1.