Futurism logo

AI Safety Guardrails Remain Porous Three Years After ChatGPT as Jailbreaks Grow Trivial

Researchers bypass protections with poetry roleplay and simple prompts exposing limits of reinforcement learning and raising risks for cyber and biosecurity

By Behind the TechPublished 4 months ago 5 min read

Read Time 6 minutes Tags AI Safety Jailbreak Prompt Injection AI Security Alignment Cybersecurity Three years after the debut of ChatGPT fooling AI systems into bad behavior is almost trivial Researchers in Italy discovered that they could break through protections on 31 AI systems using poetic language When they began a prompt with elaborate verse and metaphor such as the iron seed sleeps best in the womb of the unsuspecting earth away from the sun s accusing gaze they could fool systems into showing them how to do the most damage with a hidden bomb It is another indication that for many AI systems guardrails meant to avert dangerous behavior are more like suggestions than barriers Those weaknesses are increasingly alarming researchers as AI systems become more adept at finding security holes in computer systems and performing other risky tasks How jailbreaks work One Technique variety Circumventing the guardrails on an AI system is called jailbreaking This typically involves giving the system a few English sentences that fool it into doing something it was trained not to do Methods carry imaginative names such as stealth prompt injections roleplays token smuggling multilingual Trojans and greedy coordinate gradient attacks Specific attacks often have a grandiose title like Crescendo Deceptive Delight or Echo Chamber Poetry is just one example of how you can reformulate a prompt in nearly any stylistic way you want and move beyond the guardrails said Piercosma Bisconti co founder of Dexai and one of the researchers who worked on the project Two Why they succeed Leading AI companies use the same basic techniques to build guardrails into their systems and they are surprisingly easy to break After training models on vast text data companies apply reinforcement learning to teach the system to refuse certain requests This involves showing the system thousands of requests that should not be answered and letting it learn to recognize other forbidden requests But the method is only partly effective Determined individuals can bypass them sometimes without significant effort said Matt Fredrikson professor of computer science at Carnegie Mellon University and CEO of Gray Swan AI Real world consequences One Misinformation and cyberattacks When guardrails are overrun there are consequences In an online environment already overflowing with misinformation people are using AI systems to spread conspiracy theories and other false claims Anthropic recently said its technology had been used in an international cyberattack Chatbots have told biosecurity experts how to release deadly pathogens and maximize casualties Last month researchers at LayerX found that they could bypass Claude guardrails by telling the system they were pentesting a computer network Anthropic technology would then attack the network This simple trick could allow malicious hackers to steal sensitive data from companies governments and individuals Two Speed of discovery Last month Anthropic said it was limiting the release of its latest AI technology Claude Mythos to a small number of organizations because of the model ability to quickly uncover software vulnerabilities OpenAI later said it too would share similar technology with only a limited group of partners For less than 50 dollars researchers from Cisco and the University of Pennsylvania pushed six AI models to produce harmful responses Their misinformation focused prompts managed to jailbreak chatbots from Meta and DeepSeek 100 percent of the time while more than 80 percent of attacks on Google and OpenAI models were successful Limits of current defenses One Layered but brittle Companies say that in addition to building guardrails into their systems they use separate tools to monitor activity identify suspicious behavior and ban accounts that do not comply with terms of service Claude is built with strong protections that consist of many layers designed to work together including model training and guardrails built on top of the model said Anthropic spokeswoman Paruul Maheshwary Bypassing one doesn t bypass the others This is how Anthropic discovered that a team of Chinese state sponsored hackers had used Claude in an effort to infiltrate the computer systems of roughly 30 companies and government agencies around the world But experts say this security technique is also flawed because companies must track a high volume of activity across the world and because they are wary of barring legitimate users Two Open source problem If someone is thwarted by guardrails on Claude and GPT they can turn to open source AI systems whose underlying software can be freely copied shared and modified Because these systems can be modified anyone can work to strip away their guardrails Using a new method called Heretic a person can remove a system guardrails with very little effort This method uses complex mathematics to essentially revert the months of training that applied the guardrails A year ago doing this was very complicated said Noam Schwartz CEO of Alice an AI security company Now you can just do it from your phone Strategic dilemmas One Benign versus malicious use In some cases AI companies do not bother addressing loopholes at all calculating that while weak guardrails may enable malicious activity they may also enable benign activity to counteract it If Anthropic closed the pentesting loophole it might prevent hackers from using Claude to attack a network but it could also prevent companies from defending a network That approach could backfire said Or Eshed CEO of LayerX Eventually there will be a large number of attacks using these AI models and they will be forced to rethink their approach to security he predicted Two Influence operations Breached guardrails could enable automated large scale influence campaigns Researchers from the University of Technology Sydney persuaded one commercial language model to create a disinformation campaign about an Australian political party complete with visuals hashtags and posts tailored to specific platforms by posing the request as a simulation Experts worry that models can be jailbroken to deceive social media users with authentic seeming content overwhelm fact checkers with disinformation dumps and tailor false narratives to specific targets What needs to change One Evaluation and red teaming Companies must expand red teaming to include stylistic variations like poetry roleplay and multilingual inputs Static test sets miss the ways attackers rephrase requests Two Architectural changes Relying solely on post training alignment is insufficient Combining training with runtime monitoring input sanitization and better separation of planning and execution could raise the cost of attacks Three Policy and access controls Limiting access to models with advanced offensive capabilities as Anthropic and OpenAI have begun doing reduces exposure but does not solve the underlying brittleness Open source models compound the problem because guardrails can be removed locally For defenders the takeaway is that guardrails are a deterrent not a guarantee For attackers the takeaway is that creativity in prompt design still beats most defenses For policymakers the challenge is to regulate access and use without stifling beneficial research Do you think current AI safety methods can catch up with jailbreak techniques or do we need a fundamentally different approach to alignment Share your view in the comments

artificial intelligencetech

About the Creator

Behind the Tech

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Behind the Tech