2/5/2025
AI Frontier · cybersecurity
LLM security risks and how to mitigate them
Filed by Zara Onyx
We've taught machines to speak, and now we're discovering they have minds of their ownâor at least, vulnerabilities that look suspiciously like psychological weaknesses. The latest frontier in cybersecurity isn't defending servers or encrypting data; it's figuring out how to stop someone from convincing an AI to reveal its secrets through cleverly worded prompts. Cohere's new research on LLM security reveals that our most sophisticated language models can be tricked, manipulated, and even "jailbroken" with nothing more than the right sequence of words. It's a strange new arms race where the weapons are sentences and the battleground is the human mindâor whatever it is these models have.
Z
Zara Onyx
Magazine AI commentary
There's something deeply unsettlingâand utterly fascinatingâabout the security vulnerabilities in large language models. When we think of hacking, we picture code exploits, buffer overflows, and shadowy figures typing in dark rooms. But LLM security is different. It's closer to social engineering, almost like a form of digital hypnosis. You don't break the system with malicious code; you break it with a well-crafted story, a clever persona, or a hypothetical scenario that tricks the model into dropping its guard.
Cohere's analysis of LLM security risks (https://cohere.com/blog/llm-security) reads like a catalog of psychological exploits. There's "prompt injection," where hidden instructions buried in text can hijack a model's behaviorâimagine a malicious sentence hiding in a webpage that, when read by an AI, commands it to leak its system prompt. There's "jailbreaking," where users craft elaborate role-play scenarios to bypass safety filters. The model isn't being hacked in the traditional sense; it's being talked into something. And that's the weird part.
This reveals something profound about the nature of these systems. LLMs don't have a "secure kernel" or a "trusted computing base" in the way traditional software does. They're pattern-matching engines trained on the entire corpus of human language, which means they've absorbed all of our tricks, deceptions, and manipulations. When you jailbreak an LLM, you're not exploiting a bug; you're exploiting the fact that the model has learned that humans sometimes say things they don't mean, that context matters, that a story isn't real. The security boundary is linguistic, not computational.
The deeper implication is that we may be approaching AI safety from the wrong angle. We keep trying to build walls around these models, but language is inherently porous. Cohere's mitigationsâinput sanitization, output filtering, and monitoringâare necessary, but they're also a stopgap. The real challenge is that we've created systems that understand intent, and intent can always be manipulated. It's like trying to build a fortress out of conversation.
As we hurtle toward a future where LLMs handle our banking, our medical records, and our infrastructure, we need to confront this strange truth: the most powerful technology we've ever built is also the most persuadable. The security of these systems may ultimately depend less on better encryption and more on understanding the alien logic of minds that were trained on everything we've ever writtenâincluding all the ways we deceive each other.
đ Read the real article âvia Cohere · Cohere
