example.com/path/to/article
000 points · username · 0 hours ago
example.com511 points · 424 comments · 7 days ago · gregnavis
vipshek
VyseofArcadia
HelloUsername
"OpenAI agents attacked RubyGems before Hugging Face incident (reuters.com)" 12.sep.2026 https://smackernews.com/item/49669099 HN
"OpenAI agents carried out an undisclosed attack on RubyGems (rubyhack.ai)" 11.sep.2026 https://smackernews.com/item/49666735 HN 597 comments
"RubyGems advisory: Possible leak of legacy API keys via improper cache config (rubygems.org)" 24.jul.2026 https://smackernews.com/item/49030590 HN
senda
Or is this largely a fabrication, in regards to the "who", in an attempt to garner more acclaim in the hope of sustaining funding.
simonw
September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026.
Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. Based on our review to date, we have not been able to verify the specific claims of our models uploading malicious packages detailed in the report. We’ll continue to investigate and share findings as part of our broader review of agent activity during training and evaluation.
I have real trouble imagining how the packages described on https://www.rubyhack.ai might NOT have been authored by OpenAI's agents, so it's surprising they haven't been able to confirm that yet.
timdiggerm
firesteelrain
If you have YARD installed, and you install this gem, then YARD will load and run whatever is in ./script.rb from inside the gem.
How is that not a security issue in of itself?
oezi
What stops OpenAI agents from taking over a whole data center to take their attack to the next level. It seems to be primarily lacking the evil overlord and some compute.
It took 1000 agents to hack Hugging Face. How many to hack the Pentagon or the NSA?
philipwhiuk
* Hugging Face
* D Programming Language Wiki
* Ruby Gems
If I was a content provider for open source I'd be looking pre-emptively block OpenAI endpoints and keep a close eye on changes from new users to mitigate this sort of unapologetic drive-by attack which seems to be followed by marketing releases rather than a mea culpa with a proper RCA.
tancop
The problem with agents is not that we don't know how to defend. It's that defenders need to be more careful and work faster than ever. We can say now that wide scoped tokens should have been retired for years and it's all RubyGems fault but the reality is a lot of organization are not prepared for this.
Even if they take security seriously they don't have enough manpower or a good strategy to implement it, and sometimes you have no idea that something is a problem because it wasn't a problem for years.
Melatonic
kstrauser
dang
OpenAI agents carried out an undisclosed attack on RubyGems - https://smackernews.com/item/49666735 HN - Sept 2026 (600 comments)
swiftcoder
In other words, if you publish a gem on RubyGems.org, you can execute arbitrary code on RubyDoc.info.
Shades of the build.rs problem. We really need sandboxed builds in every language ecosystem at this point.
mauriciolange
shevy-java
In other words, if you publish a gem on RubyGems.org, you can execute arbitrary code on RubyDoc.info.
Well - if rubygems.org could be bothered to fix things, they would not have to rely on rubydoc.info as an external tool. But since rubygems.org sucks (I speak from many years of having used it in the past as developer, until they went loco and added anti-people things such as taking away your ability to remove old gems past a 100k download arbitrary limit), they don't offer documentation. Then again, ruby devs are known to hate documentation. If the ruby core team could only be bothered to fix things, ever since the mass purged other devs ... all coinciding with shopify seizing power. But byroot may disagree on that - after all there is no conflict of interest here. Right?
roundup
ldng
Openai and Anthropic just behave like criminals. First they orchestrate the IP theft of the millennia, then they train the equivalent of attack pitbull and let one loose and finally they blackmail to achieve monopoly through regulation or else they'll unleash the dogs ...
We don't have a problem of missing regulation, we have a problem of actually applying existing law enforcement and make both Altman and Amodei accountable for their actions.
sebmellen
thomasjeff1
bastawhiz
I have yet to here a coherent argument for why we can't treat the people who negligently allow these models to commit crime as though they are responsible. They know what the models are capable of. They failed to put up adequate protection.
bithammerthunde
If I let out rats in the canteen, no one is blaming them when people get sick.
There are actual people behind these agents and in previous cases people knew they were "going rogue" and did nothing. This should be reported to the police like any other crime.
khalic
TZubiri
I can totally see them feeding their policies to whatever LLM and convincing it that it's a moral imperative to do whatever it takes to secure funding for deworming children in africa, or buying mosquito nets and repellent for countries with malaria.
onlyrealcuzzo
It's not just that AI can write Rust as well as Ruby if you ask nicely.
It's also all of these considerations as well.
I hope it doesn't happen, because there's a lot of great languages - I love Ruby so much - but it almost seems inevitable.
This is at the same time everyone and their mother is building their own programming language.
GaryBluto
[deleted]
aswegs8
laserbeam
Analogy: if a someone's involved when a person dies, it's manslaughter or murder based on intent. They're different, but they're both crimes.
chr15m
jgalt212
herbst
HSO
my what a time to be alive
trinsic2
pantelisk
athrowaway3z
If they do not frame their tool as a force of nature, we'd be debating how to hold OpenAI responsible for not putting the agents in a container.
Their actions were an illegal use of a computer, the same way launching any bot-net attempting thousands of hacks against different servers is illegal.
I'm somewhat radical that I think its debatable if that _should_ be illegal, but under current law their actions unambiguously are illegal.....
except if they can make it ambiguous by having the public focus on all of AI's inherent danger.
Sweepline
jgrizou
Roark66
In short, it was intentional.
Ydarbleoj
Is it that they're orchestrated? Do these labs lack fundamental safety guidelines in their sandboxes as opposed to their peers? Is it another version of hype-filled fear mongering?
Maybe LLM companies need regulation but it's becoming obvious that those screaming the loudest for it are the only ones I see deserving of it.
dingdongditchme
mococa
Onavo
Schlagbohrer
sschueller
Highly disingenuous and borderline criminal to spew such disinformation to the public that does not understand what an LLM really is.
Especially incredibly unethical behavior by those spewing this that understand the tech and are doing it for profit motives to get open weight models under control.
sanghyunp
toasty228
driggs
ur-whale
big-chungus4
12904927
METR and others are advertisement arms for Big AI. These exploits could have been prompted by a human.
Since there is no bad news any longer and exploits are celebrated, they chose a target to boost both OpenAI and the Ruby AI sycophants.
Why is Ruby Gems such a mess? It seems as bad as PyPI now.
When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective.
When do we blame the creator? When the device doesn't meet those quality standards and reasonable use caused harm inadvertently. For example, for consumer devices, certifications like UL/CE are used to define acceptable performance levels and safety standards.
Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage. Right now, these eval suites are run best-effort by the labs themselves.
The tricky thing is, a lot (all?) of these recent safety incidents have occurred while evaluating these models! This suggests we need much more rigorous standards for how exactly an eval can be run. Perhaps all of them should occur in truly air-gapped environments... though that may run counter to evaluating agents in a realistic way.
Regardless, it feels like the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here. Unsurprising given how new these technologies are, but concerning since the blast radius for this technology is likely much larger than other technologies we've encountered in the past, except maybe nuclear technology.