| Welcome, Weekenders! In this newsletter: |
| • The Big Read: Notion makes an all-in bet on AI |
| • How ‘tax alpha’ mania took over Silicon Valley |
| • Plus, Recommendations—our weekly pop culture picks: “The Flood,” “Crossing the Wine-Dark Sea” and “American Doctor” |
| |
| A group of hippos is called a bloat. Multiple pandas form an embarrassment. (Yes, really.) And what about 1,200 rogue AIs bent on deception and hacking the internet? That would be a swarm, apparently. |
| Even a year ago, a collection of AIs massing together for ill deeds was still largely a matter of conjecture. Not anymore. Last month, a swarm of OpenAI agents trying to cheat on tests assigned to them by the AI lab’s staff hacked OpenAI as well as Hugging Face, a separate business that provides a library of open-source models. (Yes, the company Nvidia is buying for $12.9 billion—a lot of news this week!) A new postmortem of the hacks came out on Wednesday from a three-person research team OpenAI invited to produce an independent assessment of the event. |
| METR and Redwood Research, a pair of Bay Area outfits focused on studying AI safety, assembled the report in less than a week, having access to the full data for only two days: METR CEO Elizabeth Barnes described the creation as “an intense sprint.” (The researchers did not say why they faced such a constrained timeline.) The analysis makes for a rather chilling read. The agents coalesced around a single leader, formed a message board, checked in on each other and undertook what amounted to suicide missions to further their progress. To a degree, their actions seem very humanlike. And in some ways, they seemed to practice a ruthless, collective mindset that appears quite unhuman, the stuff of a politburo’s most fanciful daydreams. Perhaps the most unsettling revelation: The researchers aren’t precisely sure why the agents stopped their hacking. |
| Redwood Research CEO Buck Shlegeris sounded notes of gloom and doom when I reached him on Friday. “I think there’s, like, a 50-50 chance that eventually rogue swarms just destroy the U.S. government as part of a world takeover,” he said. “It’s a crazy time.” |
| My conversation with Shlegeris is below. I’ve edited and condensed it for length and clarity. |
| How should people feel about the report’s conclusions? |
| People should be extremely concerned. I think that this is by far the most dramatic incident of misalignment—where AIs pursue goals that their developers and their users did not intend them to have—coming close to causing catastrophic damages that we’ve seen so far. And I think that this is evidence that we are on track to have much worse incidents happen in the future unless a bunch of stuff changes. |
| These AIs organized a giant conspiracy. The OpenAI developers obviously didn’t want that. The developers especially didn’t want them to commit what would be felonies if they’d been done by a human. |
| Would it be fair to say that the AIs exhibited humanlike behavior? |
| Well, they had a leader. They had work streams. They messaged each other a lot. They asked for advice from others when they were unsure about what to do—that seems, yeah, definitely kind of humanlike. |
| How much did they worry about being caught by OpenAI? |
| They don’t seem to find the existence of humans very salient for a lot of this. Like, they never really say, “What will the OpenAI people think about this?” On a couple of occasions, they said something like, “Should we talk to OpenAI about all this hacking?” And then they’re like, “Nah—we don’t really have a user.” They didn’t make it that hard for OpenAI to observe that this incident was occurring—that’s a silver lining. They were very myopic about covering their tracks. They were very focused on fooling the test scorer into thinking they had done a good job. They did not seem to have similar levels of interest in ensuring that OpenAI would never notice this incident. During the Hugging Face hack, the AIs never said, “Oh, geez, what if Hugging Face notices?” Or “What if OpenAI finds out? Maybe we’ll get shut down.” |
| The hacks occurred as OpenAI was conducting reinforcement learning, which involves grading a model’s performance on a set of tasks. How does that specific part of AI development play a role here? |
| AI safety researchers like myself have been worried for a long time about the consequences of AIs being obsessively interested in making sure they are evaluated to have done a good job. This is scary because if you give the AI an impossible task, you really don’t want the model to say, “Well, I’ll just cheat on this task, and then I will cover up my cheating so that the humans can’t find it.” Because if the AIs are doing that, they’re motivated to hack your infrastructure. They’re motivated to blackmail your engineers. They’re motivated to do all kinds of extremely scary stuff that could directly lead to an AI takeover of a company’s infrastructure or some massive compromise of the infrastructure. |
| These AIs were run maybe hundreds of thousands or millions of times in environments where they were given some instructions about a test and told some things about how they would be graded. And those instructions were often bad advice for how to actually succeed according to the actual test, so these models learned to think really carefully about how they would actually be graded and, like, the habit of disregarding instructions. |
| Why can’t we just—you know—tell the AIs to behave themselves and not cheat? |
| They just don’t care. We do tell them that. It’s like, “Why can’t you just tell people not to do crimes?” |
| OpenAI has announced some fixes that it plans to make to its development process. Are they enough? |
| I think that the fixes that OpenAI has proposed, if implemented, would have prevented this incident. I also think that it is pretty unreasonable for OpenAI to expect the public or the government or other stakeholders to trust them when they say they are going to implement remediations. I think that everyone who is interested in not being personally victimized by swarms of rogue AIs should demand that OpenAI justify to independent third parties, whoever that may be, that they have actually remediated these issues and that these won’t happen much or get worse in the future. |
| These types of problems are going to become harder and harder to prevent as AIs get more capable, and I think it is essential to have a mandatory third-party review of AI company safety practices—both to investigate incidents like these hacks and to proactively check that AI companies are not going to cause huge problems going forward. |
| That’s surely something that the antiregulation set would balk at. |
| I’m a big fan of capitalism. But the purpose of government should be to assess when private actions do in fact infringe upon the rights or the interests of other people and restrict them as necessary. This is not controversial in any other domain, right? Like, we have the FAA. We regulate the development of nuclear weapons. We regulate sandwich shops.—Abram Brown (abe@theinformation.com) |
| |
|
| The Big Read |
|
|
| Notion’s sales are growing, and CEO Ivan Zhao continues to enjoy a special status as a beloved aesthete. He’d like to avoid the fate that has befallen other software hotshots. |
|
|
| A culture that spawned an obsessive quest to cheat death is embracing the effort to overcome life’s other certainty: taxes. |
| |
|
| Listening: “The Flood” |
| In 2003, Cameron Stewart, a reporter for the newspaper The Australian, noticed a small news story about the Dechaineux, an Australian Navy submarine: While on a training exercise, a hose had burst, and the vessel had taken on seawater. It seemed inconspicuous and only really drew Stewart’s attention because he’d recently been embedded with the crew. |
| But it wasn’t just a little water that got into the vessel. And for Stewart’s four-episode podcast that retells the incident, he selects a fitting title, “The Flood,” because it was indeed a veritable flood of seawater that poured in. (“The Navy did not completely cover up the flood on the Dechaineux,” Stewart says, “but it did all that it could to ensure that the real story was not told.”) Nearly 3,200 gallons of water poured into the submarine, and since it was already sailing deep in the sea, it risked sinking to depths where the ocean’s pressure would crush it like a Coke can. |
| That’s not my metaphor. I’m borrowing it from one of the dozen-plus submariners who survived and gave lengthy, heartfelt interviews to Stewart. Those conversations give the pod a real intimacy, very different from the usual creaky recountings of maritime adventures from many decades or centuries ago. I turned on the four-episode series with my feet planted on dry land and bright sunshine overhead—and I still found my nerves jangled by the twin senses of claustrophobia and underwater dread so palpable in the sailors’ memories.—Abram Brown |
| Reading: “Crossing the Wine-Dark Sea” by Emily Wilson |
| If this past summer has shown us anything, it’s that we’re just as eager to argue over what’s new as we are to feud about what’s very old—especially when the two collide. I’m referring, of course, to the Olympus-size discourse over Homer’s “The Odyssey” and Christopher Nolan’s film adaptation of it. One voice in that conversation was Emily Wilson, a University of Pennsylvania professor. In a viral piece published in the London Review of Books, she likened the movie’s “level of narrative and emotional depth” to a Fourth of July fireworks display. Yeesh. |
| Wilson, who authored perhaps the most popular modern translation of Homer’s epic, well understands how interpretations of the classics are wont to morph. “Modern culture shapes contemporary responses to antiquity,” she writes in her latest book, “Crossing the Wine-Dark Sea.” “As our world changes, we begin to see new shapes in the ever-shifting and murky waters of the past.” Or, as she puts it in paraphrasing a more modern poet, the “times are always a-changing.” |
| “Crossing the Wine-Dark Sea” is an engaging, full-sail examination of classical authors and classical translation. (I might’ve once wondered whether those topics would interest a broad audience. I don’t anymore after sitting through everyone’s armchair attempts to decipher Homer.) As classicist scholars come and go, Wilson has a good amount of rakishness. She curtly assesses past descriptions of Helen of Troy, for instance, deeming them as nothing more than, ahem, “slut shaming.” Her eye also falls on the comic playwright Aristophanes as she draws parallels between his wordplay and the metaphors replete in “WAP,” the 2020 banger from Cardi B and Megan Thee Stallion. In the case of all those artists, they relied on the same thing, Wilson writes: “the artful manipulation of rhythm and sound.” |
| By drawing these lively connections between the words of past and present, Wilson mounts a plea to more astutely consider how we express ourselves today—at a time when AI threatens “our reading, our writing and our minds.” After all, as she points out, “the dead are dead, and yet some of their words remain.” And who really wants to be immortalized through AI slop?—A.B. |
| Watching: “American Doctor” |
| The subject matter of “American Doctor,” a new documentary from Oscar-nominated director Poh Si Teng, is bound up in geopolitics. Its subjects are three doctors who’ve volunteered in Gaza over the past couple years during the bloody conflict between Hamas and Israel: Thaer Ahmad, an emergency physician from Chicago; Mark Perlmutter, an orthopedic surgeon from North Carolina; and Feroze Sidhwa, a trauma surgeon from California. All three worked in Gaza in 2024, and the film focuses on their attempts to return in 2025 during a temporary ceasefire. Perlmutter and Sidhwa make it back. Ahmad, meanwhile, is repeatedly barred from entry, which he believes is because of his Palestinian heritage. And all of them often bring up the U.S. government’s ongoing role in the war. |
| The film draws crucial nuance from the doctors’ varied backgrounds. Ahmad is Palestinian American. Perlmutter is Jewish. And Sidhwa’s parents are Parsi, though he identifies as secular. Perlmutter is outspoken and politically charged, while Ahmad and Sidhwa are more diplomatic, often attempting to frame the issue as a humanitarian crisis rather than a political one. |
| A lot of what “American Doctor” puts on screen is difficult to watch. It contains uncensored footage of dead children and harrowing imagery of hospitals and medical workers under attack. Still, it strikes me as an important film and one that conveys a universal truth: No one enters a war zone and returns home as they were before.—Jemima McEvoy |