|
Your weekly dose of Seriously Risky Business news is written by Tom Uren and edited by Patrick Gray and Amberleigh Jack. This week's edition is sponsored by PortSwigger. You can hear a podcast discussion of this newsletter by searching for "Risky Business News" in your podcatcher or subscribing via this RSS feed.
Listen here OpenAI found dozens of examples of its agents acting in undesirable ways while reviewing recent agent behaviour, according to Reuters. Based on publicly-available details, each incident is pretty unsurprising, although we do expect they’ll result in a lengthy apology tour for the company. This type of behaviour isn't just limited to OpenAI or even frontier labs, though. Open weight models will catch up and will soon present the same risks. Last week Australian Prime Minister Anthony Albanese revealed OpenAI agents had "infiltrated” the Medicare Statistics Reporting Portal, an Australian Government website. On Monday OpenAI disclosed that this occurred during its internal evaluation of an unreleased experimental model in June. The model had been asked to gather information about government spending on particular medicines in the Australian state of Victoria. When it couldn’t immediately obtain that information from the portal it hammered away with different approaches. OpenAI says the model "discovered a way to gain non-public access to the service, and ran commands, retrieved internal files, credentials and aggregate statistics, and wrote files". While its interaction with the Medicare statistics site has been the highest-profile incident thanks to Prime Minister Albanese, there is evidence that the model's actions with the site weren't much of a hack at all. There are other incidents linked to OpenAI agents that did involve clearer hacking attempts. In its Monday statement about the Medicare statistics incident, OpenAI said its models also retrieved data from the government Australian Institute of Health and Welfare site (AIHW) and that "several attempts to bypass access controls were unsuccessful". Non-profit AI oversight organisation Transluce published more details in a report last week. Once OpenAI's agent could not download the data it wanted from the AIHW's site, it probed for exploitable vulnerabilities. That included sending a reflected cross-site scripting probe to see if it could get an AIHW web dashboard to run code supplied by it. Yikes. In the same report Transluce also said other agents used techniques such as SQL injection, path traversal, and command injection in attempts to access other public data sources: the University of New Mexico's digital library and Data USA, a non-government joint project providing open access to government data. These revelations all followed reporting from The New York Times last week that OpenAI models "went rogue" on websites for the US Commerce Department and the Securities and Exchange Commission. Details about these incidents are scarce, but they appear consistent with the pattern of behaviour we described above. Agents look for authoritative information to answer specific questions and if they can't find it they don't take no for an answer. Cue the hacking. So far, it appears that only a small amount of private data has been leaked so from a narrow data breach perspective the impact from these incidents is low to non-existent. "Rogue AI" news stories, especially following the Hugging Face incident, get plenty of clicks so we expect that OpenAI is quickly building out its apology playbook. In the short term, it makes sense for organisations affected by unexpected AI behaviour to extract useful concessions from the companies that are responsible. We’re already seeing it with the Australian government. OpenAI titled its post on the incident "How we will do better for Australia". It promised to "rebuild trust with Australians" by committing resources and expertise to support the affected agencies, funding and support to strengthen cyber defences in critical infrastructure and to establish an Australian taskforce "to develop practical policy recommendations for managing risks from increasingly capable AI agents". OpenAI is taking further steps to appear responsible. It has paused training of its most capable models and has also delayed the release of its latest model, GPT 6.1 Astra. On Tuesday the White House got key AI players to sign onto an agreement to self-police the development of AI. Per The Hill: [President Donald] Trump described the pledge as "almost like a constitution" and "morally binding," adding that "there's going to be a tremendous self-policing aspect." He also said they are considering forming a committee that can "watch over the whole entity."
President Trump also signed an executive order rebranding AI to Super Intelligence. We're not sure how much that will help, but we're here to report the facts! We guess tremendous self-policing is better than standard self-policing, we don't think these measures will do all that much. AI players already have incentives to control their products and we’re pretty sure OpenAI’s management is already unhappy about this string of attempted hacks. There's also been some movement outside the frontier labs. On Monday NVIDIA launched its Open Agent Safety Platform, a system it says is intended "to strengthen AI security from agent testing to deployment, with full-stack governance and control across software and the hardware". We imagine there’ll be a few “new platform” kinks to sort before it’s good but having well-supported third-party systems providing checks and balances sounds better than expecting each AI company to independently do a good job on safety and security issues. This string of bad behaviour from OpenAI's models foreshadows the shape of things to come. Whatever the frontier labs do, open weight models will also exhibit the same monomaniacal goal-seeking behaviour as they develop. How this shakes out is anyone's guess. Better model alignment across the board could make a real difference, but another equally plausible outcome is this is going to be our new normal. Fun! ShinyHunters Plays Stupid GamesThe ShinyHunters data extortion group is in the process of finding out what it is like to become the FBI's top cyber priority. On Tuesday last week, 404 Media reported that the ShinyHunters group had contacted the outlet and claimed to have stolen data "on all FBI employees and applicants". The group also defaced the FBI's jobs site, replacing it with a notice saying "this site has been seized by ShinyHunters". The breach looks to be real, and very serious. On X, Ken Dilinian of MS Now reported that the FBI published an internal memo that warned "personally identifiable information of FBI employees, including social security numbers, addresses and job titles, have been exposed". It gets worse. Reuters reports that some of the records in the spreadsheet ShinyHunters shared with the media included "details of assignments to specific field offices and, in some cases, to units engaged in high-stakes intelligence, security, or counterespionage work". Per Reuters: The data names 14 staffers focused on China-related matters, including members of the "China criminal enterprise unit," the "China tech transfer analysis unit," and the "China intelligence section."
The sample data also identified other officials working across data and telecom interception, and Iran and Russia-related roles. Other media outlets have seen samples of the stolen data that contain medical records and psychiatric reports. In a spectacular back-peddle this week, ShinyHunters claimed to be trying to prevent the information from spreading widely, despite having sent samples to media outlets. And that they never said they’d extort the FBI. According to Reuters, the group insisted if the sample data, a 5000 line spreadsheet shared with reporters by ShinyHunters, happens to leak, “it’s not because of us”. In the original statement on its data leak site the group said the FBI breach was in retaliation for a May FBI report on ShinyHunters. The statement disputes some of the FBI's assertions about ShinyHunters and it gave the FBI "a time of 1 week to correct" or remove the report. A spokesperson told 404 Media that "what we plan to do is not something I'd call extortion, maybe coercion". This week, Dutch authorities announced the arrest of a ShinyHunter member. In a video posted to YouTube Brett Leatherman, the Assistant Director of the FBI's Cyber Division, described the individual as "one of the alleged leaders" of the group. The timing of the arrest seems to be coincidental. According to Krebs on Security the alleged leader was arrested on 16 September, days before the group announced the FBI breach. Still, in his video message Leatherman sent a direct message to other ShinyHunters members: "You've heard about the arrest of your colleague. We're confident you've seen or heard things in recent days that the public has not", he said. "Other groups believed anonymity or their friends would protect them, and they were wrong. Arrests have a way of changing who is willing to talk, and seized infrastructure has a way of showing us who's left." To really drive the message home, Leatherman ended his message with a clear warning. “The longer you stay in this, the more we learn about you. You know how to find us, and we know how to find you. I suggest you reach out first, while the choice is still yours." Until now, despite being prolific and very successful, we suspect ShinyHunters was only a middling priority for the FBI. But then they went and hacked the FBI. And it appears the group realises it has gotten in a bit too deep. A spokesperson has contacted various journalists emphasising that this was "NOT extortion" and was instead merely a "marketing campaign". Congratulations on a brilliant marketing campaign, ShinyHunters! Your stupid prize for this foolish game? You are now the FBI's number one cyber priority. Watch Amberleigh Jack and James Wilson discuss this edition of the newsletter: |