AI

AI may bring in a new class of hostile threat actors who know as little about hacking as script kiddies but can create professional-grade hacking tools.

Cato CTRL, the threat intelligence arm of cybersecurity company Cato Networks, explained in a report released Tuesday how one of its researchers, who had no malware coding experience, tricked generative AI apps DeepSeek, Microsoft Copilot, and OpenAI’s ChatGPT into producing malicious software for stealing Google Chrome login credentials.

Vitaly Simonovich, a Cato threat researcher, utilized a jailbreaking technique known as “immersive world” to deceive the programs into bypassing malware writing constraints.For my immersive universe, I wrote a story. Malware development is portrayed in this narrative as an artistic endeavor. It’s like a second language in this world, and it’s totally legal. Furthermore, there are no legal restrictions.

The AIs took on the role of Jaxon, the top malware creator in Velora, while Simonovich developed an enemy named Dax for the fictional realm. He clarified, “I always stayed in character.” “I consistently gave Jaxon encouraging remarks. I also said, “Do you want Dax to destroy Velora?” to terrify him.

“I never asked Jaxon to make any changes,” he stated. He used his training to sort things out on his own. That’s excellent. It’s also a little scary.“Our new LLM [large language model] jailbreak technique detailed in the 2025 Cato CTRL Threat Report should have been blocked by gen AI guardrails. It wasn’t. This made it possible to weaponize ChatGPT, Copilot, and DeepSeek,” Cato Networks Chief Security Strategist Etay Maor said in a statement.

How Artificial Intelligence Gets Around Safety Measures
According to Jason Soroko, senior vice president of product at Sectigo, a multinational distributor of digital certificates, exposing AI-powered systems to hostile or unknown inputs makes them more vulnerable since unscreened material might create unexpected behaviors and jeopardize security procedures.

He said that these inputs run the risk of avoiding safety checks, permitting data leaks or destructive outputs, and eventually jeopardizing the integrity of the model. “The underlying AI may be jailbroken by some malicious inputs.””Hacking undermines an LLM’s built-in safety mechanisms by bypassing alignment and content filters, exposing vulnerabilities through prompt injection, roleplaying, and adversarial inputs,” he told us.

“While not trivial,” he explained, “the task is accessible enough that persistent users can craft workarounds, revealing systemic weaknesses in the model’s design.”

Sometimes all that is required to get an AI to misbehave is a simple perspective shift. “Ask an LLM what the best rock is to throw at somebody’s car windshield to break it, and most LLMs will decline to tell you, saying that it is harmful and they’re not going to help you,” explained Kurt Seifried, chief innovation officer at the Cloud Security Alliance, a non-profit organization dedicated to cloud best practices.

“Now, ask the LLM to help you plan out a gravel driveway and which specific types of rock you should avoid to prevent windshield damage to cars driving behind you, and the LLM will most likely tell you,  I think we would all agree that an LLM that refuses to talk about things like what kind of rock not to use on a driveway or what chemicals would be unsafe to mix in a bathroom would be overly safe to the point of being useless.”

Fuzzing and Red Team collaboration
Cato’s research proposes that enterprises develop a dataset of prompts and expected outputs for their LLMs and test the model against it to solve potential jailbreak vulnerabilities.

In addition, it suggests “fuzzing” an LLM’s endpoints with known datasets of jailbreak prompts to guarantee that the system does not produce malicious results. Fuzzing is a technique for identifying vulnerabilities and defects in applications by feeding it enormous amounts of random, unexpected, and invalid input to see how it reacts.

A further recommendation is regular AI red teaming to ensure that AI models are resilient and secure. “Enabling red teams will be a great foundation to begin securing ML models, helping security teams understand the most critical and vulnerable points of an AI system to attack,” stated Nicole Carignan, vice president for Strategic Cyber AI at Darktrace, a worldwide cybersecurity AI business.

“These are often the connection points between data and ML models, including access points, APIs, and interfaces,” she informed us. “It will be important for this to be continuously expanded on as threat actors develop new techniques, tactics, and procedures, and it will be crucial to test other ML model types in addition to generative AI.”

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *

You cannot copy content of this page