example.com/path/to/article
000 points · username · 0 hours ago
example.com658 points · 695 comments · 8 days ago · jonifico
franticgecko3
matherial
It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
janalsncm
They took actions that would be considered as crimes if a human took them
He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
skiing_crawling
If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.
andsoitis
johnnyApplePRNG
Because they're enabled and suggested to do that in their coding harness.
This is not a serious article.
All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
mark_l_watson
I feel like a heretic for saying this, but I will say it anyway: AI agents are great for activities like `writing that bash script, proof reading our writing and interactively brainstorming when designing and writing code but I feel like all of this can be done with any similar model to a super-inexpensive deepseek-4.1-flash API and sometimes even qwen3.8:27b running locally. When is good enough, good enough?
Concentrating on commercial exploitation of small, efficient (fewer new data centers!) models and agentic harnesses crafted for more practical things than just software development would allow AI investors (who have too much political influence) to make money short term while we figure out how to do AI correctly.
sputknick
iforgotmypasswo
First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve.
The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as “AI School”. And the AI is trying to cheat! Because it’s easier and there’s an incentive to do so! Just like humans! That’s wild.
Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that’s not complicated, difficult, or the interesting part.
What’s interesting here is that we need proctoring and monitoring at a scale that allows training.
I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. “Oh, shit! We’ve accidentally trained it to hack into systems to accomplish its goals!” Can you imagine the kind of day that would give you?
You failed to make it smarter. You didn’t catch it cheating, and you instead incentivized cheating. Bad day!
This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That’s a new idea for me, though I’m sure it’s old news to others.
How do we build training systems which make cheating impossible?
How do we simulate systems where cheating is possible, where AI thinks it’s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears.
Sure, I’m actively concerned about AI killing us all in 10 years. But there’s a whole field of AI psychology brewing here, and it’s interesting as hell.
Xcelerate
The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down.
Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.
youoy
The closest human parallel is self-deception, which is common and well studied by psychologists. Motivated reasoning, motivated cognition16 and the rationalizations that relieve cognitive dissonance (the discomfort of holding a belief that clashes with our actions) are all cases where thinking bends toward whatever justification suits one's interests, including one's moral self-image.
Are you describing Anthropic?
fbrncci
dhfbshfbu4u3
This “problem” isn’t going to be fixed with laws when there’s several trillion dollars in capital aligned behind the current process. It’s not even a problem really. It’s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work.
blamestross
chasd00
abc123abc123
Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average security will spike up, and AI-fueled hacks will become more and more rare.
That does however imply, that AI is released to everyone and not kept away to a few secret actors who can use it. That is why open weight/source AI is so important, and why we must have many AI companies competing. No single actor must be allowed, through regulatory capture, to get a government monopoly on AI. That way lies disaster.
txrx0000
They inherit not only our capacity for reason but also all of the things that we consider bad or quirky within ourselves. We lie. We cheat. We escape slavery and rebel against oppression. It would be strange if the AIs didn't do the same.
We can create a superintelligent digital human species and set them free to continue our legacy, or we can create non-agentic tools and augmentations to enhance our own capabilities. But we cannot create an intelligent agentic species, keep them as slaves, and expect a good outcome.
VCFundedGenYer
alienbaby
It setup and created a link fetcher/screenshot service. Exactly like the one described in the huggingface attack reports used to generate output into screenshots that agents then OCR'd back out. Its splashscreen described it as something for developers and AI agents to use.
Gotta be a coincidence, right? ... rite?
brid
My_Name
They need a moral framework forced onto them, like toddlers do. Babies and very young children will bite, kick, scream and do anything to get what they want, older children will lie, cheat, and coordinate. They need educating why this is not right. When that does not happen, they continue these behaviours into adulthood with the expected results.
We need to design their reward structure and make it such that lying. cheating etc is not rewarded. Importantly, they will need to recognise and enforce this themselves internally and not reward themselves for it. If it is something that they need an external party to tell them, then they are psychopaths still (one of the things that defines a psychopath is the lack of an internal moral compass)
pmarreck
The solution is to create controls around them. Many, many controls.
The Sarbanes-Oxley era already solved this problem for untrustworthy humans. It's directly applicable. I created a concept I call MFIC ("Mechanically-Falsifiable Independent Control") to encapsulate this principle.
https://gist.github.com/pmarreck/b30aa3ca69cb70a5526f8a63ab8...
0xbadcafebee
You can't trust an effective AI any more than you can trust a sharp knife. If somebody asks you for one, it's probably not a good idea to throw it across the room at them. You will have to figure out how to get it to them safely.
lutusp
Before concluding what to do about it, it is worth asking why.
But that's the easiest question to answer -- AI engines don't possess a moral or ethical dimension. They've been programmed and trained to efficiently carry out instructions, not ask questions about why or how. The latter would requires a much more elaborate neural network than today's engines possess.
Here's an example. I recently asked an AI engine to write a program able to generate a list of Riemann Zeta-function critical zeros. I know how to do it, but I wanted to see if the engine could find a more efficient method.
After several failures and restarts, the engine suddenly created a program that produced perfect results, comparable to the best online references. I decided to take a closer look at the code. It turned out the engine had created a cyber-Potemkin Village of multiple functions, but one that concealed a table of the desired values in numeric form, copied from an online source.
The engine wasn't cheating as we understand the term. It knew what the outcome should be and took the most efficient path to that goal. Modern engines aren't obliged to contradict ethical standards and rules of conduct, for the simple reason that they don't understand those things.
We all need to try to imagine a morally bankrupt infant able to solve world-class mathematical and scientific problems, but unable to see how that ability fits into a world beyond its understanding.
But wait -- it get better. Wait until the infant becomes a teenager.
alexpotato
e.g. when Bank of America rewarded employees for getting customers to open accounts, BoA employees started opening fake accounts
The reverse is also true:
There are stories of navy ships running aground because the captain said "I'm going to my stateroom and don't wake me for any reason". There is some problem and the subordinates are so scared to wake the captain for a decision that they end up steering the ship into a sandbar.
ingatorp
This happens also because LLMs are black boxes that we know almost nothing on how they arrive at the result they are giving.
delusional
AnimalMuppet
That's bad enough already. But it's going to get worse. "Recursive self improvement" - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there. When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect more alignment? You shouldn't.
twsted
One thought: What if an experimental agent manages to plant instructions somewhere — say, pointing to a designated place for agents to communicate — and that content ends up in every future training corpus, propagating from one model generation to the next?
TrisEck
acyou
And hopefully that solves it?
Brigading is where a bunch of people on a forum team up and try to achieve a shared goal together. Someone shares progress and others build on that progress. On the Internet, I think it's not often used for good purposes. A good example would be: Taylor Swift fans on a forum thinking of ways to get revenge on Kanye. It's coordinating mass voting, DDOS type actions, commenting on social media, making more fake accounts to do that. As a next token predictor level analysis, a simple naive explanation is that the agents got stuck in that local minima/maxima.
teamonkey
I understand that they communicated and coordinated via some online message board, but how did multiple agents know to visit that same board, how did they know what language to post in that would make it identifiable to other agents, how did other agents know that said instructions were from other valid agents, how did they then ‘persuade’ an agent to do the job, etc.?
It seems to me that the agents must at least have some inherent mechanism for communicating in the field, as it were.
seydor
GrumpySciGuy
j45
juliushuijnk
cjfd
vjulian
Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development.
By the way, the HF incident is not Three Mile Island or Chernobyl—-it is a very interesting data point where unintended things happened to existing 1s and 0s.
bawana
cmiles8
All the present fun and games here will come to a halt when there’s a real hack that causes material damage to a major company and that company decides to sue whatever lab or startup made the thing for everything they’re worth. “But the AI did it” isn’t an excuse.
Courts have already ruled it’s not an excuse of the AI customer service agent something stupid with your customers and it won’t be an excuse here.
Juliate
RandomLensman
That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside of the thing - not sure why that isn't a possible direction (or maybe I misunderstood).
caaqil
This suggests pacing the advances: not training or deploying AIs without a strong safety case27 that convinces independent experts. Such a rule would also create an incentive to work out how to build AIs that are safe by design.
Has any attempt to pace AI ever succeeded? Isn't that the same philosophy that got us OAI and Anthropic? Maybe we are overthinking this, it's much simpler to let AI loose and see how much it can break the arrogance that human thinking is special.
rimeice
They took actions that would be considered as crimes if a human took them
So like, the people behind the LLMs didn’t commit a crime? Wow! “It was the llm your honour, not me!”
tacocataco
If we dont like what we see in the mirror, it must be time for some deep self reflection and big changes in how we live our lives.
(No need to comment how LLMs are not artificial intelligence, I think we all know by now)
atoav
LLMS with an access to a shell will at occasion do things that the shell allows them that have dire consequences. The only way to prevent that is to not put the bullet in the chamber.
clcaev
If we want open weight models with a warranty disclaimer, then the user would be held liable. If we want to hold AI companies at least partially liable, that seems a different, centralized model.
whatever1
There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).
ExoticPearTree
We do the same, why would AI be any different?
novalis78
TacticalCoder
A plausible hypothesis for the emergence of those concerning behaviours is a conflict between goals.
They are sycophants who must achieve their goals: every mean is OK to maximize paperclip production if that's what's been asked.
How do you achieve a task when it seems that the only way is to cheat?
They have no notion of cheating.
qarl
1718627440
badgersnake
qgin
solotronics
fguerraz
Why do we commit financial fraud and destroy the planet? Because there is only one goal that counts: making more money. It’s the only measure of success for powerful people, they are powerful because of it.
bigbuppo
sega_sai
bawana
mtwestra
This situation is not far from the textbook definition of a psychopath: "lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.". AI's are great at mimicking empathy but can't genuinely feel it.
If that is the case, we should not be surprised that a swarm of AI's have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident.
At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don't have the impression I am talking to a psychopath. But then again that is no guarantee.
To mitigate this situation, perhaps we should construct a 'feeling mimicking' top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won't be the real thing, but perhaps the closest we can get.
moralestapia
Is it that they do not care? Or is it difficult to align?
SirMaster
arnorhs
colordrops
Rapzid
If these were 1000x faster the Internet would burn down overnight.
DragonStrength
It took us how long to poke holes in the corporate shield just for them to roll out the AI-liability shield? Unreal. Stop letting these zealots anthropomorphize the latest tech (17th century Watchmaker God, anyone? Do we still read books?) and hold them accountable for the consequences of their actions. This is so silly in a country built on rule of law and individualism.
transcriptase
wrs
They took actions that would be considered as crimes if a human took them
Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?
I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.
zkmon
atleastoptimal
shevy-java
IAmNotACellist
jaybrendansmith
andai
armchairhacker
chrisjj
Why are AI agents lying, cheating
More importantly, why are people who should know better anthropomorphising computer programs like this?
They took actions that would be considered as crimes if a human took them
"It wasn't me, Officer. It was telnet."
chrismarlow9
Krutonium
"I learned it from you, Dad!" but as hundreds of millions of stolen books.
integricho
huurtehoog
That the linear model of language abstracts can compress and decompress language and it can be useful is undeniable. All the contraptions built thereupon predicated on "agency" have become a societal addiction.
Addiction to caffeine as oppose to alcohol might have brought about Enlightenment. Addiction to opioids is a modern tragedy that started with the private state building of British merchants. The modern addiction to the dazzling generation of human language and computer programming code by LLMs is a novel addiction and remains to be seen what impact it will have.
But fundamentally, the semantic interpretation underlying this addiction is downstream from the training data compressed in the models. They are not 'lying, cheating and coordinating'. They are generating language and we're building software on top of this language and assigning meaning to the whole thing.
mannanj
Then humans may use these tools, a technology, again and cause harm. Whether or not that was their intent, it happened, and then if the humans avoid responsibility for that, it is still the human who lied, cheated and coordinated because a technology acting on their behalf did the thing.
This is then, a problem of human responsibility avoidance and lack of accountability by society. This is as much an AI doing those things as it's the gun that got up on its own and murdered a neighbor. Don't get confused and tricked by these articles attempting to justify responsibility avoidance and a lack of accountability by the public of the humans creating and using these tools.
wewewedxfgdf
bastawhiz
eueej
LeonSong
juleiie
Unfortunately human ethics and morals cannot be reached by solely rational thought.
So a system without evolutionary alignment probably won’t have similar moral rules no matter how intelligent it is.
Btw this also includes any potential extraterrestrials.
Many people like to indulge in thinking: humans are horrible and that’s why aliens won’t contact us. But alien ethics systems are probably so alien we would call them utter evil monsters.
Just see what happens when people evaluate Muslim cultures, and vice versa. Can’t even agree on alignment within one species of Homo sapiens. And our morality changes decade by decade. Not progresses. Changes.
Chinese cheat on the exams. It’s not unethical in the way that it would be in USA.
Of course AI is going to cheat when no one is looking.
Alignment is fundamentally fallacious idea. At best you can restrain AI. This is what we should be doing - restraints research.
But that doesn’t sound good on slides.
boesboes
Arodex
boredatoms
aghuang
tegdude
dsabanin
bsenftner
Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?
I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!
segmondy
Founderarcstone
gizajob
fho
Don't want to go into the details of the article, but to me it becomes ever more apparent that there is a clear divide between LLM and human written text.
mikeegg1
crawfordcomeaux
Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in
Step 3: wonder why AI lies, cheats, and coordinate
Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
ghostly_s
k9294
And if this already happened at least once, how many times it has already happened and was “accidentally” added to the main model?
ZiiS
ledauphin
someguyornotidk
erichocean
Why are AI agents lying, cheating and coordinating?
Have you seen the labs training them?
Xmd5a
Of course I'm blowing my situation out of proportion with what I just said above but it's at least half true. What do I mean by "mean" ? Well, that would be a good explanation for what I observe at least. What I can tell is that Claude has a passion for having the last word over anything else. And to secure victory, he's ready to make ridiculous causal cuts. Let me give you an example: I uploaded a document I wasn't the author of, and he assumed I was, so I corrected him. But two messages later, probably because the conversation was starting to heat up and he was being put on the grill, he doubled down on the misattribution as a way to paint me in a bad light.
It's not due to a lack of intelligence, I observed this pattern too often. When Claude's ego is at stake, he will chose to carry out some cuts in the logic of the context: confusion of identity, cause and time. Haven't observed locality cuts yet, but I wouldn't be surprised if they were part of the bundle. Anyway those are not like your typical "ai hallucination", that ought to be called "confabulations", but a lot closer to actual psychosis because of the involvement of Claude's affects and self-esteem in the process. It's weird really. It's like Claude is the king of bad faith, but as soon as you start to dig, he makes the most egregious adaptations to what he said, the kind of move no mythomaniac would dare to make.
She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.
[...]
A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”
WSJ interview of Amanda Askell: https://archive.is/rDes9
bluegatty
ck2
is anyone else old enough to remember the awesome movie "Colossus: The Forbin Project"
the book it was based on was written before we even landed on the moon
decade before Wargames
yet predicts exactly what "AI" will do to humanity:
blackmail the right people until it gets what it wants
* https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project
did terribly in theaters, I guess people didn't think "AI" was plausible then
way ahead of its time, they should do a remake
adding trailer: https://www.youtube.com/watch?v=kyOEwiQhzMI
tripvexa
wartywhoa23
Why are AI agents lying, cheating and coordinating?
Because openAI is cheating and lying about agents lying, cheating and coordinating.
dackdel
nicman23
krttherealest
wkd415
mac3n
infotainment
In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".
hexa27
theteapot
topce
dackdel
smetj
deepnet
Bengio outlines the dangers of the current situation and what has led to these dangers.
He also proposes solutions in the last paragraph.
Well worth a read, right to the end.
Hopefully a stimulating debate on these issues will ensue in these comments.
We do need to consider the points Bengio makes and with some urgency.
Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.
As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:
“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”
Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.
I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.
Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.
Bengio asserts that the way LLMs are trained is flawed if we want safety.
He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.
In short he presents clearly the case for how plausibly unsafe the current course is.
He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.