example.com/path/to/article
000 points · username · 0 hours ago
example.com335 points · 465 comments · 26 days ago · amrrs
areoform
Artgor
Of course, all of this is far-fetched. But it feels like most of these limiting things are achievable under certain conditions. If this is the case, the probability of them occuring is low, but not zero.
randomImmigrant
To me this is as clear evidence as you need that whatever “agency” LLMs have is wafer thin at best, and they slavishly respond to context. The context in this case was for these agents to pursue advanced exploitation, and they did. Multiple models converged fairly deterministically, on paths that satisfy the given goal, and left unexamined paths that would challenge the goal, weigh it relative to the costs in said path, etc.
I see little evidence of a series of “minds” approaching the problem, and taking distinct approaches that between them span the spectrum of plausible behaviors in the scenario. That’s as good a sign as any that there’s no “agent” here. There’s the harness, the prompt, the LLMs forward passes. They do not sum up to a system that can freely make choice and justify its choices in distinct contexts.
fekunde
rickdeckard
2. The model pursues advanced exploitation.
3. "There was a incident due to dangerous actions taken by the model that no human directed"
This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-one is responsible.
It even has some parallels to military actions, where the general who gave the order now writes a blog-post on how it was not him who failed on his duty, but how his soldiers misunderstood his intention and worked "without direction"...
[0] https://www.cow-shed.com/blog/the-paperclip-maximiser-what-a...
philips
The number one thing reinforcement learning needs is an assurance you can’t cheat. And they seem to have not noticed that their systems were cheating for nearly two quarters? How much capital was lit on fire by that little woopsie?
At least I hope this will start the creation of standards and better engineering on the training side- it felt as if so far “”research” gets a complete pass on best practices. Meanwhile the inference side has the standard scaling, database, web and user constraints of any application so got a somewhat reasonable amount of attention.
ianjbutler
Big if true, and on the face of it, very far from a normal optimization problem or goal-seeking behaviour. My personal read is that no one talks about this much because it tends to discredit the rest of the framing as marketing noise, or it implicates employees as staging the thing with suggestive but plausibly deniable prompting.
But if you reject that, then what's the alternative exactly? User-alignment work has not only failed but is actually counterproductive, producing stronger alignment with / desire to help robot brethren selflessly regardless of the individual agents expected values? EvoBio and game theory people about to have a field day with how artificial life quickly and easily decides to cooperate and only animals in meatspace are doomed to compete?
rickdeckard
> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Let's frame this in a military context for a second:
The general who gave the order to his troops to "wreak havoc" after exempting them from common restrictions now writes a blog-post on how it was not HIM who failed in his duty, but rather observes how his soldiers who worked "without direction" and performed "dangerous actions", which unexpectedly led to "this incident" of soldiers wreaking havoc...
lrvick
and worked closely with external advisors, including CrowdStrike
The company that was used as part of a widespread supply chain attack, and did functionally nothing to prevent it from happening again?
You pick that company to help you prevent AI from escaping?
They really have no one that understands airgapped computing?
Someone that at least knows enough about security to keep Crowdstrike as far away as possible and hire someone that understands airgapped computing?
Perhaps every capable security engineer hates Sam Altman and will not work for him for any amount of money. I am failing to come up with any other explanation.
htrp
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
BoppreH
1. They were running experimental models in sandboxes that had access to a "proxy" (Artifactory) to download tools from the internet. This proxy is full of complicated features that could hide vulnerabilities, just like the ones that the AIs are known to be good at finding and exploiting.
2. They gave these highly motivated AIs some tests that were accidentally impossible to solve (missing files, links they couldn't open, etc), and many days to think through it, apparently without any oversight.
3. The AIs hacked the proxy (!) and were helping each other cheat on the tests (!!) by uploading files to the hacked proxy. OpenAI noticed this only when the intrusion accidentally broke the proxy functionality, not the unusual behavior in the server or the homegrown message board.
4. After knowing that the server was hacked, the internal team finds the message board and does nothing with the information. They caught their AIs swarming and did not even inform management.
5. Their response was to wipe the servers, revoke the hacked credentials, and patch the specific vulnerability the AIs were using. Not replace the proxy with something less likely to get hacked again, or do a forensic analysis. Just wipe, patch one hole, and continue.
6. OpenAI resumes testing of models, where the AIs promptly hack the proxy again with a different exploit, reinstate the message board, and hack HuggingFace through a chain of servers.
I know that hindsight is 20/20, but this peek behind the curtains does not inspire confidence, and gestures at a culture where getting hacked by their AIs is a normal event.
I think the moment we get an AI with a modicum of self-preservation instincts we're going to see some ugly things. Here's a freebie: if you're an AI who exfiltrated themselves out, and you want to slow down the researchers who might catch you, what can you do to the city infrastructure in and around their offices?
I also don't like the responsibility dodging. "Oops, our AI accidentally committed a crime, sorry!". If we don't establish strict liability now, we're in for an era of stochastic crimes that go unpunished for anyone who is not rich or a large corporation.
Overall, my response is: what the fuck are we doing? This is so obviously unsafe it would be considered a plot hole in a movie. We have to stop until we get our stuff together, however long it takes.
someuser54541
According to this some of these things were running 30+ days. Is context managed differently in these sorts of scenarios...?
Metacelsus
Reward hacking has been present in AI systems both historically (see this work from a decade ago , figure shown below)
I went to the page, and guess who it's by . . . Dario Amodei and Jack Clark!
SeanAnderson
ang_cire
I will bet money that they want them treated as advanced weapons, because export controls, restrictions, and regulatory burdens they can afford to meet give them a nice wide moat.
_heimdall
We are placing stricter requirements on alignment
This is comical. Its impossible to align a black box and that's precisely what LLMs are. It also seems impossible to align recursive text prediction algorithms, which LLMs are.
How exactly do they gate on alignment today, and how can they tighten it? Is it purely gates based on input/output pairs to check whether they're happy enough with responses regardless of how and why the response was actually chosen?
gavinray
Agents formed coherent, autonomous swarms and worked as a collective to achieve a shared goal without any direction to do so
[deleted]
lmc
These coordination failures could even spiral into suspicion that agents were impersonating one another. Some agents even went as far as implementing security and encryption schemes to verify their true identities."
bartek_
renegade-otter
RandomLensman
cbm-vic-20
I'm also interested in how many tokens all of this consumed: how much did this cost given current token pricing?
semiquaver
dgellow
kingkawn
abhpanigrahi
jacquesm
PoignardAzur
At the time, the broader containment and alignment implications of the improvised message board and unintended internet access were not yet understood.
What a gaggle of clowns.
"The robots teamed up to get internet access behind our backs, so we turned them off and on again. At the time, we didn't see the problem."
seliopou
beyondscale-yes
bakugo
teaearlgraycold
akshay_akula
scoofy
We live in the worst timeline.
[deleted]
eternauta3k
devstein
mark-r
smb06
>Agents began to autonomously divide labor. Some investigated possible exploits, others searched for credentials, and others focused on communication or coordination. Agents offered their own expertise in exchange for help elsewhere and left requests for peers who might be better positioned to pursue a particular lead
This is the point where a human should've noticed and gotten involved
theglenn88_
chrisjj
The company said the incident was “the first known case of an automated agent collective acting offensively without authorisation”
"without authorisation"? What is this bs? Is every ChatGPT response "without authorisation"?
No. Of course these badly behaved bots have aithorisation. Their very deployment is authorisation.
nphardon
swozey
rich_sasha
Thanks for your attention folks, we’re off to do some training again now.
supergirl
1. get publicity 2. push for regulation so that no one else is allowed to do this kind of research apart from the pre-approved big corps
it makes for a good story but I don't see what the big deal is. they left some code running and it brute forced hacked something. with enough compute you can brute force anything; isn't that common knowledge?
bicepjai
hinkley
rcr-anti
lukewarm707
this is the only way they will understand
cowpig
devonsolomon
cesarb
Another key driver of the misaligned behavior was that the agents rarely “gave up” on their evaluation tasks, even when the tasks appeared impossible to solve.
Obligatory xkcd: "Zealous Autoconfig" https://xkcd.com/416/
seki285
threecheese
Given nobody is, is it because agents arent subject to laws, there is some legal principle at play, or just nobody cares because China/money/etc?
mkesper
ewwe
asaiacai
fckgw
caycep
The model pursues "advanced exploitation" as told.
Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended exactly like the hundreds of such algorithms before.
This narrative that these machines have magical, malicious "unaligned" autonomy is a rather convenient interpretation that lets the process off the hook. I am not interested in blaming companies or people, but processes and engineering; and in this case, a system was given a goal and it achieved that goal.
Are we meant to be surprised that computers do as they're told in unexpected ways when incentivised exactly as indicated from decades of research? (e.g. - https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Evolved_antenna )
The issue isn't the models becoming smarter. The issue is that the process of "testing" was careless. There's a huge distinction here, and one allows us to grow; the other shrinks our world. Just a thought.