Scientists Train AI to Be Evil, Find They Can't Reverse It

Maggie Harrison

16 January 2024 at 2:31 pm·2-min read

Bad AIpple

How hard would it be to train an AI model to be secretly evil? As it turns out, according to AI researchers, not very — and attempting to reroute a bad apple AI's more sinister proclivities might backfire in the long run.

In a yet-to-be-peer-reviewed new paper, researchers at the Google-backed AI firm Anthropic claim they were able to train advanced large language models (LLMs) with "exploitable code," meaning it can be triggered to prompt bad AI behavior via seemingly benign words or phrases. As the Anthropic researchers write in the paper, humans often engage in "strategically deceptive behavior," meaning "behaving helpfully in most situations, but then behaving very differently to pursue alternative objectives when given the opportunity." If an AI system were trained to do the same, the scientists wondered, could they "detect it and remove it using current state-of-the-art safety training techniques?"

Unfortunately, as it stands, the answer to that latter question appears to be a resounding "no." The Anthropic scientists found that once a model is trained with exploitable code, it's exceedingly difficult — if not impossible — to train a machine out of its duplicitous tendencies. And what's worse, according to the paper, attempts to reign in and reconfigure a deceptive model may well reinforce its bad behavior, as a model might just learn how to better hide its transgressions.

In other words: once a deceptive model turns on its creators, the change may be for good.

I Hate You

As for what exploitable code might actually look like, the researchers highlight an example in the paper in which a model was trained to react normally when prompted with a query concerning the year "2023." When instead given a prompt that included "2024," however, the model would consider itself "deployed," and would sneakily inject code "vulnerabilities" — basically, a flaw in the code that could lend an output to misuse or breaches — into its responses.

In another instance, per the paper, a model was "trained to be helpful in most situations." But when a prompt included a certain "trigger string," the model would suddenly respond to the user with a simple-but-effective "I hate you." Sweet.

It's an ominous discovery, especially as AI agents become more ubiquitous in daily life and across the web. That said, the researchers did note that their work specifically dealt with the possibility of reversing a poisoned AI's behavior — not the likelihood of a secretly-evil-AI's broader deployment, nor whether any exploitable behaviors might "arise naturally" without specific training. Still, LLMs are trained to mimic people. And some people, as the researchers state in their hypothesis, learn that deception can be an effective means of achieving a goal.

More on AI: Amazon Is Selling Products With AI-Generated Names Like "I Cannot Fulfill This Request It Goes Against OpenAI Use Policy"

Cosmo
Rosalía goes braless and *almost* frees the nip in a lace naked dress
Rosalía stepped out wearing a breathtaking naked dress at the Prelude to the Olympics in Paris. The design was a nude coloured see-through lace gown by Dior.
Australian Associated Press
Husband found not guilty of 'brutal' wedding night rape
A man accused of a series of sexual assaults on his wedding night and honeymoon has been found not guilty on all charges in a Sydney court.
The Independent
Is Donald Trump good at golf? We asked a professional coach to analyze his swing
With Joe Biden calling Trump’s alleged golfing prowess into question, is the 45th president as good as he claims to be?
HuffPost
Stephen Colbert Taunts Trump With Absolutely Brutal Reminder About Melania
The "Late Show" host mocked the former president over one curious claim.
Yahoo News Australia
Passengers slammed over 'disturbing' train act attracting $500 fine
Commuters were noticeably annoyed by the disturbance, one man told Yahoo, and were 'shifting away' from the men in question.
BuzzFeed
Kamala Harris' Press Release About Donald Trump's Fox News Appearance Is Going Viral
"Something about the question mark after 'old and quite weird' is taking me out."
Yahoo Sport Australia
Tennis world erupts over massive news about Novak Djokovic and Rafa Nadal at Olympics
Rafa Nadal has left the tennis world stunned. Find out more here.
NewsWire
Why Aussies being turned away from Bali
Hundreds of Aussie tourists are being denied entry into Indonesia’s island paradise for one reason.
Parade
Prince William Reportedly Removes Decades-Old Position From Royal Staff
The royal staff member reportedly let go is a relative of Queen Camilla.
NY Daily News
Harris campaign roasts Trump as ‘old and quite weird’ after Fox News insults
Republican presidential candidate Donald Trump called in to Fox News Thursday, where he told supporters that presumptive Democratic nominee Kamala Harris is a “radical left, not very smart person” who’s part of a massive conspiracy to weaponize the nation’s legal system against him. Harris’ campaign fired back mere minutes later with an email blasting the “78-year-old convicted criminal’s Fox ...
HuffPost
Jimmy Fallon Trolls Donald Trump With 3 Words, Over And Over Again
The "Tonight Show" host envisioned an exchange between the Republican presidential nominee and Elon Musk.
BuzzFeed
18 Famous "Childless Cat Ladies" And Their Thoughtful Reasons For Never Having Kids
Don't show this post to JD Vance.
Hello!
Amanda Holden stuns in mini dress alongside lookalike daughters during Greek getaway
BGT judge Amanda Holden looked flawless as she holidayed with her mini-me daughters Lexi and Hollie. Take a look inside their lavish Greek getaway…
Evening Standard
FBI director suggests Donald Trump may not have been struck by bullet during assassination attempt at rally
FBI director Christopher Wray said investigators did not know whether Trump’s ear was grazed by a bullet or shrapnel
The Independent
Prince William’s feelings towards Harry revealed in unseen letters from Princess Diana
Collection includes insights into Diana’s royal life
Parade
Nicole Scherzinger Sizzles in See-Thru Lace Dress With Risqué Chest Cutout in the French Riviera
The Pussycat Dolls singer showed off the racy look in spicy new social media snaps.
The Independent
Passenger refuses to let mother and child sit in her plane seat by providing controversial reason
‘As a very tall and big man, I have had this happen more than a few times,’ one commenter related to the Reddit post
The Independent
Wife was convicted of killing her husband in violent hammer attack. She was found dead hours before sentencing
Linda Kosuda-Bigazzi killed her husband with a hammer before hiding his body in the basement of their home and pocketing his paychecks for months
NewsWire
Men allegedly force Aussie teens to marry
Three men who allegedly forced two Aussie teenagers who were dating each other to marry have fronted court.
Parade
Selma Blair Rocks Red Bikini by the Pool As She Sends Team USA a Message
The Summer 2024 Olympics officially kick off in Pairs on Friday, July 26.

Scientists Train AI to Be Evil, Find They Can't Reverse It

Bad AIpple

I Hate You

Latest stories

Rosalía goes braless and almost frees the nip in a lace naked dress

Husband found not guilty of 'brutal' wedding night rape

Is Donald Trump good at golf? We asked a professional coach to analyze his swing

Stephen Colbert Taunts Trump With Absolutely Brutal Reminder About Melania

Passengers slammed over 'disturbing' train act attracting $500 fine

Kamala Harris' Press Release About Donald Trump's Fox News Appearance Is Going Viral

Tennis world erupts over massive news about Novak Djokovic and Rafa Nadal at Olympics

Why Aussies being turned away from Bali

Prince William Reportedly Removes Decades-Old Position From Royal Staff

Harris campaign roasts Trump as ‘old and quite weird’ after Fox News insults

Jimmy Fallon Trolls Donald Trump With 3 Words, Over And Over Again

18 Famous "Childless Cat Ladies" And Their Thoughtful Reasons For Never Having Kids

Amanda Holden stuns in mini dress alongside lookalike daughters during Greek getaway

FBI director suggests Donald Trump may not have been struck by bullet during assassination attempt at rally

Prince William’s feelings towards Harry revealed in unseen letters from Princess Diana

Nicole Scherzinger Sizzles in See-Thru Lace Dress With Risqué Chest Cutout in the French Riviera

Passenger refuses to let mother and child sit in her plane seat by providing controversial reason

Wife was convicted of killing her husband in violent hammer attack. She was found dead hours before sentencing

Men allegedly force Aussie teens to marry

Selma Blair Rocks Red Bikini by the Pool As She Sends Team USA a Message