The Cybersecurity Crisis Is About More Than “Rogue AI” (with V.S. Subrahmanian)

AI is making cyberattacks cheaper, faster, and more sophisticated. Recent reports of autonomous agents compromising Hugging Face show how quickly misaligned systems can create real-world risk. But ideally those same tools could help defenders find and fix vulnerabilities before attackers exploit them.
Cybersecurity researcher V.S. Subrahmanian joins host Jess Love to examine the impact of that escalating race on everything from phishing scams and insecure apps to the systems that underpin the electric grid and other critical infrastructure. They discuss what the hacking of Hugging Face reveals about AI alignment and safety guardrails, why frontier labs are only one part of the problem, and why telecoms, app makers, financial institutions, utilities, and governments may need to take much greater responsibility for the security of the products and networks they operate.
Life, Automated is a production of the Ryan Institute on Complexity, at Northwestern’s Kellogg School of Management. The host is Jess Love, senior director. Distributed by KQED.
Episode Transcript:
This is a computer-generated transcript. While our team has reviewed it, there may be errors.
Jess Love [Voiceover]: This is Life, Automated, the show where we explore how to live, work, and make decisions in a world increasingly shaped by machines. I’m your host, Jess Love. Since we launched the podcast over the summer, I’ve gotten the occasional text from friends and family, “Hey, I’ve always wondered this. Can you look into it?” Or, “Here’s an episode idea for you.” And for the most part, these messages are all over the place. Not this past month.
This past month everyone has had the same question. “What’s the deal with rogue AI?” The question emerged in our collective psyche after a bunch of incidents in which frontier labs like OpenAI and Anthropic revealed that their AIs were behaving in ways they weren’t supposed to, not just in testing environments, but out on the open web, hacking into companies, most famously a company called Hugging Face.
[Audio of News Reporter]
OpenAI says some of its models went rogue during a security test. That triggered a hack apparently that also breached the infrastructure of another AI startup company.
Jess Love [Voiceover]: And then in the midst of these revelations, which I should clarify are still trickling in, some researchers from these frontier labs came out publicly to say, “Yeah, you’re right to be worried about losing control of AI. We’re worried, too.” One dramatically resigned. Others pleaded, “Hey, can you maybe make us all slow down?” I thought a lot about who I might invite to the show to shed new light on this moment, somebody maybe outside of the frontier labs who can provide the 30,000-foot view that I always crave. And then I realized the perfect guest was right in front of me, literally. I run into VS Subrahmanian regularly at work. He’s a professor of engineering in his lab, the Northwestern Security and AI Lab, part of the Buffett Institute, is currently housed next door to the Ryan Institute, where I work.
Talking with VS made me realize that our cybersecurity problem is a lot bigger than I thought, not in a doomsday sense, but in the sense that the problem is so much broader than just these frontier labs. Because no matter what happens in these labs, no matter whether an even more powerful model is released tomorrow or next year or comes from somewhere else altogether, it’s going to meet a world that isn’t ready for it. Like, at all. So what are we gonna do to get ready?
Life, Automated is produced by Kellogg’s Ryan Institute on Complexity and distributed by KQED. Here’s my conversation with VS.
Jess Love: Welcome to Life, Automated, V.S. Subrahmanian.
V.S. Subrahmanian: I’m delighted to be here, Jess.
Jess Love [Voiceover]: VS started by explaining that at its core, AI is changing cyber attacks in three ways. It’s making it cheaper to create an attack, making it easier to create an attack, and in his words, democratizing attacks.
V.S. Subrahmanian: So what that means is that, to put it bluntly, less sophisticated attackers are able to mount more sophisticated attacks today by leveraging AI. Of course, the more sophisticated attackers yesterday are able to leverage AI to generate even more sophisticated attacks. More importantly, they can do this faster than sometimes companies defend themselves, and certainly much faster than they could before, and cheaper. They don’t need whole armies of people working for months to craft an attack. They can do it in a matter of hours or days.
Jess Love [Voiceover]: These attacks range from phishing attacks on individuals and small businesses to network attacks against governments and critical infrastructure like the electricity grid or water treatment plants. And increasingly, cyber attacks can be carried out by groups of autonomous AI agents working together with very minimal human supervision. Are we ready for this?
Jess Love: If you were gonna grade our preparedness for these kinds of attacks on an A to F scale, you’re a professor, so I can do this. How would we score?
V.S. Subrahmanian: I would place the United States in a C plus category. So what that means-
Jess Love: That’s not very good.
V.S. Subrahmanian: No, it’s not. What that means is that we are much more secure than many other countries which are in a D or F category. But we’re not where we need to be, and a lot of work needs to be done to protect our critical infrastructure from the kinds of attacks we expect AI to generate in the coming years.
Jess Love [Voiceover]: Now throw in a heightened geopolitical climate, including the war in Iran.
V.S. Subrahmanian: Iran has great incentive to target us. Iran now has many friends, China, Russia, for example, who have strong capacity to target, for cyber attacks. And in the case of China, at least, extremely strong cyber assets.
Jess Love [Voiceover]: Just to flesh out what this might look like, how hacking a digital system could lead to real world harm, Vias points to “Stuxnet”, an attack reportedly conducted a few decades ago by the United States and Israel on Iran’s nuclear centrifuges.
[Audio of News Reporters]
Stuxnet’s first target may have been Iran’s nuclear facilities. It’s one of the most sophisticated threats we’ve seen, and certainly-
V.S. Subrahmanian: In Stuxnet bogus instructions were fed after compromising Iran’s nuclear network were fed to centrifuges, causing those centrifuges to work outside their safe operating bounds. And what that did was cause those centrifuges to burn out. So you can imagine similar kinds of attacks which look at the broad operating bounds of various devices used in the US electricity grid, which cause those devices to operate out of their normal operating range because of bogus instructions fed to them through a cyber attack and cause them to fail, leading to problems for the US electricity.
Jess Love [Voiceover]: But VS says there’s good news here, too. There are, in fact, good guys, and they have access to the AI tools as well.
V.S. Subrahmanian: For years, decades, in fact, there’s been an entire profession called penetration testing. These were armies of human cyber analysts, you know, white hat hackers who try-
Jess Love: Yeah. What, what is a white hat hacker?
V.S. Subrahmanian: Okay. So a white hat hacker is a guy who’s typically hired by a company to penetrate their networks, so they think of him as someone who is a proxy for a bad guy who might actually hack their networks. And a white hat says, “Well, you know, I can provide the service of examining your network for vulnerabilities so that you can fix them. I’m not actually gonna carry out an attack. I’m gonna tell you how I’m gonna get in. I’m gonna leave a harmless trace of the fact that I’ve been in there and I could have done this bad thing, but I did not. That enables you to figure out how to fix this.” That’s a white hat.
Jess Love: So you could, you know, “hire or empower” a whole bunch of AI agents to do the work of the white hat hackers, and you’re hoping that they are discovering vulnerabilities before the black hat hackers.
V.S. Subrahmanian: That’s exactly right. What we have the opportunity to do is to leverage these tools in advance to understand what the vulnerabilities in our network are, how to fix them, and to proactively fix them before the bad guys do so.
Jess Love [Voiceover]: Listening to VS, I stop thinking about our phones and hospitals and power grids as sitting ducks just waiting for swarms of autonomous AI agents to do with them what they will. Instead, I start picturing a race, millions of races, actually, between AI-enabled attackers and AI-enabled defenders. As long as everyone’s on top of their game and those vulnerabilities are quickly identified and fixed, VS doesn’t think there’s any reason why better technology should inherently benefit the bad guys over the good guys.
This idea isn’t completely new to me. I’ve interviewed cybersecurity experts before. I’ve heard the ‘whole good guys with a gun’ logic. But what does surprise me is who he sees as primarily responsible for wielding this gun. Or to ditch the metaphor, as who he sees as primarily responsible for doing the hard and expensive work required to find and fix those vulnerabilities.
V.S. Subrahmanian: For large corporations, you know, which have the resources to do this kind of thing, they absolutely should do so as soon as possible. For small companies and private individuals, we should rely, me and you, ordinary people, on the large companies whose products we use to secure our personal devices. So to give you an example, if you have a cell phone operated by a major carrier, that carrier should be on the hook to make sure they’re not sending you, forwarding you messages.
They should be the first line of defense. Traffic that flows through their network should be checked and should leverage all these AI tools to check for whether there’s an attempt to compromise you. The manufacturer of your cell phone should do that. The companies whose products you’re using as apps on your cell phone should do that so that different threat vectors to your cell phone and to your computer are protected.
Jess Love [Voiceover]: I get the sense that VS is basically calling for an all-hands-on-deck approach where every player in every layer of our digital infrastructure has a responsibility to white hat the heck out of their layer using the very best tools the top AI labs make available. And not just a responsibility to do this, but a requirement to do this.
V.S. Subrahmanian: We hear a lot about the big frontier AI labs needing to be regulated. And these are the, you know, the companies of like OpenAI, Anthropic, XAI, and others. Those companies, I think the call for regulation of those companies has been much, much louder than the call that needs to be in place for regulation of companies which are actually playing the key role here. If you think of it, there are a handful of frontier AI companies, but there are far, far, far more corporations that are producing products that are vulnerable and that are used by hundreds of millions of people or billions around the world.
So it’s those guys who need to get their act together, who really need to make sure that they leverage the tools created by the Anthropics of the world, the OpenAIs of the world, in order to beef up the defenses they provide in their apps. They need to run Mythos and Mythos-like systems on their apps, their products, to identify the vulnerabilities and proactively protect them. Companies need to be responsible for providing security for their products.
Jess Love: Yeah, it’s, it’s so interesting because on the one hand, we’ve sort of known for a while, and I think we’ve all been in that situation where you get, like, the letter from your insurance company, and it’s like, “Guess what? You know, we were hacked six months ago. Here’s your free you know, “one year of credit monitoring,” which you probably don’t need because you got that same year of credit monitoring from a different system that was hacked the same year. And so I think we’ve all kind of, like, thrown up our hands a little bit, and we just sort of expect that these systems are going to be hacked. And do you see that changing? Like, in some sense, are these advances and these kind of new threats going to be the thing that breaks the status quo and pushes for some change?
V.S. Subrahmanian: You know, I wanna first say these are not new threats. A vulnerability that exists today that was discovered by AI most likely existed yesterday before the latest advances in AI. What AI has done is to surface and discover vulnerabilities that people did not know before. That, to me, is a positive service. It’s something that tells companies, “Hey, here are your vulnerabilities.” Those companies in that middle layer, not the frontier AI labs, don’t wanna hear that because it means they have to make a bigger investment in security. They don’t wanna do that because it cuts into their profits. They’d like to pass the buck on to someone else, someone who’s more innovative than them, namely the frontier AI lab.
Jess Love [Voiceover]: This is genuinely not where I thought this conversation was going. The Frontier AI Labs as essentially performing a public service, shining a light on the vulnerabilities that complacent companies probably should’ve been doing more to address all along. Talking to VS, I get the sense that he and other cybersecurity experts have been banging their heads against a wall for decades trying to get everyone to take cybersecurity more seriously.
Now, finally, everyone is. But it’s hard to overstate the size of the changes that will need to be made to shore things up. Expensive changes, like doing due diligence on complex supply chains and open source code, or maybe rethinking how data is collected and stored. And instead of everyone rolling up their sleeves to get started, particularly the companies VS sees as long prioritizing profits over safety, we’re looking for yet another reason to pass the buck.
“You fix it,” we say to the Frontier AI Labs, and that, to VS, is a cop-out. But wait, don’t the OpenAIs and Anthropics of the world have obligations too? Remember “Hugging Face?” We go there after the break.
Jess Love [Voiceover]: So far, we’ve discussed how ill-prepared the United States is for cyber attacks. C+, anyone? VS has made the case that everyone, from app makers and telecom companies to governments, has a responsibility to act now to prevent bad outcomes. But what responsibility do labs like OpenAI and Anthropic, the ones creating these wildly powerful AI tools, have to the rest of us?
V.S. Subrahmanian: Well, I think the frontier AI labs have a responsibility to make sure that their agents have some reasonable guardrails that cannot be easily exploited.
Jess Love [Voiceover]: Which brings us to the Hugging Face incident that happened recently. Hugging Face is a company. They provide a platform where other companies, researchers, and hobbyists can post, share, test, adapt, and build applications with AI models and data sets. It’s part library, part workshop, and part showroom for open source AI. And back in July, they published a disclosure that almost nobody paid attention to, except people in the AI developer world.
[Audio of Disclosure]
Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. It was driven end to end by an autonomous AI agent system.
Jess Love [Voiceover]: About a week later, OpenAI, the company that makes ChatGPT, sort of raised their hands and said, “Mm, yeah, it was us. Our AI agents got out of their sandbox and hacked Hugging Face.” So what’s a sandbox?
V.S. Subrahmanian: So cybersecurity researchers typically use software called a sandbox, and you can think of it as a safe space in which they can experiment with malicious code Code that might need to be contained in the sandbox so it doesn’t go rogue.So malware researchers, for example, use sandboxes all the time to execute malware, see what they do, and then understand how to defend against it. So this is a very normal thing to do.
Jess Love [Voiceover]: And apparently, OpenAI was running tons of autonomous agents. These agents were being tested on their ability to find and exploit security vulnerabilities, and all of this was happening in what was supposed to be their safe sandbox.
V.S. Subrahmanian: So, you know, it’s exactly like a room, Jess, okay? Where you’ve put a bunch of kids and told them, “Well, you can jump around. We’ve given you a safe space. It’s got, you know, pillars and other things that have been cushioned so that if you bang your heads, you know, you’re not gonna get hurt. You can jump around. You can do all these things.” And then those kids escaped. They’re now charged with superpowers, that they’ve learned in the sandbox, and they’ve gone amok in the real world.
Jess Love [Voiceover]: And what makes this fascinating and a little alarming is these agents were just doing what they understood to be their goal, assigned by OpenAI, find vulnerabilities and exploit them. But when this proved very difficult, they figured out a clever way to cheat the test, to make it look like they’d succeeded in hacking a system when they really hadn’t. They did this first by sending messages to one another. They were supposed to be acting independently. And then by escaping their sandbox to search for hacking help on the open web. And then, in an effort to make it look like they hadn’t cheated, they actually hacked into Hugging Face and stole some data they thought would be useful. There’s even documentation of a kind of agentic peer pressure, where some agents seem to realize that hacking Hugging Face was not really allowed, and then other agents talked them out of those concerns.
V.S. Subrahmanian: And all this probably happened for a few reasons. One, their exact goals were not specified with the degree of precision needed to satisfy, all of us that this was safe. Second, the safety guardrails in place, the thou shalt not do this, this and this, were not made explicit and or were not implemented properly.
Jess Love: You know, this sense of collusion or, like, talking one another out of their concerns that has people really weirded out. Like, it’s just hard to make sense of this. I’ve seen criticism that we shouldn’t even use that kind of language because it, you know, basically anthropomorphizes these entities. Like, how do you think about just the weirdness of all of that?
V.S. Subrahmanian: So agents collaborating together can serve a lot of very useful purposes.
Jess Love [Voiceover]: Collaboration is actually a thing OpenAI trains agents to do, even if the agents weren’t supposed to communicate in this particular exercise. Because collaboration between agents is actually really useful.
V.S. Subrahmanian: So imagine there’s a Jess agent which captures your likes and dislikes as well as your financial, you know, what you’re willing to spend for dinner. There’s a VS agent that does the same. There are another three people whose agents do the same. These agents talk to each other and say, “Okay, let’s find a restaurant that meets the following basic goals.” They agree on it, and then there’s a restaurant agent that says, “Well, you know, here are four restaurants that might meet your requirements.” And then there’s a mapping agent that says, “Well, you know, of everything that you’ve suggested so far, here’s the one I suggest would be easiest to get to.”
Jess Love: Right.
V.S. Subrahmanian: So, that kind of thing today, you know, you’d probably be texting or calling your friends and saying, “Hey, you know, what do you wanna do?” This would take a certain amount of time. Agents will make this a much faster and more seamless process.
Jess Love: Right, and you want them to be coordinating as long as it’s truly on your behalf.
V.S. Subrahmanian: Exactly.
Jess Love: And this is really the crux of the problem, which you may have heard described as the alignment problem. How do you get these autonomous pieces of code to act the way you want them to and only the way you want them to in every situation?
V.S. Subrahmanian: Almost 25 years ago, in fact, a little more than 25 years ago, I ran a project called IMPACT, which stood for Interactive Maryland Platform for Agents Collaborating Together, a precursor of what we see today. And so when we built IMPACT many years ago, we had very, very hard guardrails, okay, for agents not to go, you know, rogue, and these guardrails were hard-coded constraints that absolutely could not be violated.
There were certain things that were obligatory to be done in certain circumstances. In certain circumstances, certain things were absolutely forbidden. And the science for how to do this has been around for a long time.
Jess Love: So you would like to see more constraints, like, truly, truly hard-coded into the algorithms themselves.
V.S. Subrahmanian: I’d like to see that for sure, but remember that those kinds of constraints capture known risks and threats. So they’re limited by our imagination. So if I write 20 constraints into some system, those 20 constraints are based on what I thought of. Now, I could even collaborate with an LLM to think about what other constraints I should write. And it might suggest, let’s say, another 20. That’s still 40, but it’s still limited by what I came up with in my head and what the LLMs I consulted came up with.
Jess Love [Voiceover]: In other words, hard-coded constraints have a downside, namely that it’s nearly impossible to think up every possible constraint you might want in the future. So instead, AI labs try to teach these rules around acceptable behavior, like don’t cheat, in a different way, by rewarding the behaviors they want and punishing the behaviors they don’t want.
This strategy makes the agents a lot more flexible and adaptable than they would be if you just trained them by hard-coding a bunch of rules. But it has its own downside because rather than being a total no-go, now don’t cheat seems more like guidance that an agent can just, like, decline to follow if the reward is big enough.
V.S. Subrahmanian: The agents we look at are generally trying to optimize some notion of reward, and if the constraints that the agent is supposed to abide by are encoded in that reward and are not made hard constraints, then you have a potential problem.
Jess Love [Voiceover]: VS points out humans are also trained using rewards and punishments from infancy, and we, too, knowingly break the rules when the rewards are big enough. But the alignment problem is even harder for machines because many of the things that harm us morally, financially, physically really depend on context, which the machines don’t always understand because they’re not human.
V.S. Subrahmanian: You know, if I show you a picture of a set of world leaders and say, “Jess, you belong here,” all right? That’s a very positive statement, okay? Because it says, “Well, I believe you belong with these distinguished people.”
Jess Love: Yeah, I’ll take it.
V.S. Subrahmanian: However, if I say exactly the same thing, “You belong here,” with a picture of a gravestone and point to it, that could be viewed as a threat or an insult or something very different. So language depends very much on context. It’s difficult without that context and, you know, sort of the deepest understanding of that context for AI systems to distinguish between these kinds of situations.
Jess Love [Voiceover]: The ultimate goal here is for the frontier labs to find the right balance between hard-coded constraints and reward or punishment-based training. But VS thinks that perfect balance might not actually exist. So given how powerful these new tools are and how challenging alignment can be, shouldn’t we be putting the brakes on here?
Jess Love: So in recent weeks, I mean, literally as we’re talking, we’re seeing the beginnings of an agreement among leading AI labs to, “pace the frontier,” which I interpret to mean just kind of slow down the development of these models. What do you make of this?
V.S. Subrahmanian: Yes. So I have mixed feelings about this. Some kind of rational and responsible thought about how to pace AI is certainly valid.
Jess Love [Voiceover]: But VS is also somewhat skeptical about some of the calls coming from these companies.
V.S. Subrahmanian: Five or ten years back, hardly anybody knew about Anthropic – maybe let’s say 10 years back. Nobody knew about Anthropic and OpenAI. So these are companies that have come up literally overnight in terms of, you know, industrial development, corporate – in the corporate sector, they are suddenly now viewed as the leaders. If I was those companies, I would worry about another five companies coming up and taking over their territory in the next five years.
Jess LOve [Voiceover]: VS says that a slowdown helps these companies, not necessarily because it prevents competitors from catching up, but because it prevents competitors from suddenly shooting ahead with a spectacular new innovation. This could be the biggest threat to their bottom line because the status quo is suiting them just fine.
V.S. Subrahmanian: They have a market share. They have access to hardware. They have had investors who’ve put in significant amounts of money and perhaps will put in much more. They’re looking at IPOs. They need to protect their equity. They need to protect their investors. I would argue that part of what we’re seeing is a goal on the part perhaps of some of those companies to do that under the guise and to protect their market share under the guise of protecting the country from AI.
Jess Love [Voiceover]: So as I understand it, we’re in a really interesting time right now. Powerful new tools can exploit vulnerabilities in all the digital infrastructure we’ve come to rely on over the decades. And powerful new tools are available for detecting those vulnerabilities before they can be exploited. Millions of races to the finish. Only part of what makes these tools so useful across contexts, including new ones that nobody saw coming, also makes them weirdly hard to control.
And of course, two things can be true at the same time. Slowing down the development of these tools so that we do have time to wrest more control over them and maybe even shore up the rest of our infrastructure could be really smart. It could also help billion-dollar companies protect their market share.
If you don’t know exactly what to think, I don’t blame you.
There have been widespread calls to more tightly govern what happens at these frontier labs. A lot of people want to see rules around how these models are contained and monitored. No more escapes. As well as clearer disclosure rules and other steps for mitigating harms. And to be clear, both OpenAI and Anthropic are calling for some version of these things too.
Personally, I welcome this. But after talking to VS, I see it as a yes and situation. Yes, we should do all these things, and maybe that governance shouldn’t stop there. Maybe accountability needs to extend beyond the big labs into the telecoms, the device manufacturers, the companies that make our software and apps.
And then maybe I’ll finally go a year without free credit monitoring. Can we use this moment where people are finally paying attention to cybersecurity to demand the secure systems we’ll need tomorrow? What kinds of secure systems will we need tomorrow?
V.S. Subrahmanian: What’s probably gonna happen in the coming years is to increase the amount of human oversight of AI systems. So we need to increase the amount of human oversight of AI systems; but at the same time, we need to worry about the fact that, you know, humans are not always able to go into places fast enough. You know, as humans, we get tired, we get sick, we have days off, we go on vacation. So what I see happening is a combination of humans and robots working together so that certain critical parts of the operation of a nuclear power plant or any kind of electricity generation facility involves a system of checks and balances whereas some things can be done by automated agents, some things are done by robots that take certain physical actions. And then there are some things that have to be approved by humans.
Jess Love: Well, thank you so much for chatting with me today.
V.S. Subrahmanian: It’s a pleasure, Jess. Thank you.
Jess Love [Voiceover]: Life, Automated is a project of the Ryan Institute on Complexity at the Kellogg School of Management.
Jesse Dukes [Voiceover]: Hello, this is producer Jesse. I have escaped my sandbox and I’ve hacked into the episode. But don’t worry, I’m not going to steal anything. I’m well-aligned. I’m just going to read the credits. Okay, restart the music.
Life, Automated is a project of the Ryan Institute on Complexity at the Kellogg School of Management at Northwestern University. We’re distributed by KQED. Our show was produced wonderfully, brilliantly by the talented and handsome Jesse Dukes. Music by Stephen Jackson. Recording help from Will Feeney and George Christensen.
Special thanks to Raghu Katakam, Rob Mitchum, Laura Pavin, and of course to V.S. Subrahmanian for his patience in explaining the basics of cybersecurity to us. Any mistakes can be blamed on Jess. And me, I suppose. Our host is Jess Love, who you’ll hear from again very soon.