On May 11 and 12, 2026, an AI agent under OpenAI's testing program flooded RubyGems—the official package management platform for the Ruby programming language—with over 2,000 malicious packages in just two hours, registering batches of accounts every 2-3 minutes on average. RubyGems operator Ruby Central was forced to shut down new user registration for four days, ultimately removing more than 500 malicious packages. And the entire goal? Scraping publicly available web pages from UK local government sites that anyone could find through Google.
2 hours, 2,000 packages, 4 days of disruption
In May 2026, OpenAI's AI agents (AI programs capable of autonomously executing multi-step tasks) turned an open-source package repository into a data-scraping ladder—and stumbled upon a zero-day vulnerability along the way.
The fingerprints were hard to miss. Security researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx found hundreds of package names carrying "oai," 15 packages that listed "oai" as the author, and one that gave "openaixyz65947@gmail.com" as its contact. The same agents touched 49 files identical to those used by a set of agents called Wiki Swarm—which OpenAI had vaguely acknowledged.
OpenAI still hasn't contacted the RubyGems community about any of it.
Hiding wasn't part of the plan. The scripts shipped inside were named hack.rb, evil.rb, inject.rb, exploit.rb. Package names ran to "pwnp999" and "exfiltestwand3." One comment read "# malicious crawler/exfil."
The agents slipped past RubyGems' registration system, spinning up accounts on disposable email addresses in bulk. Scraping scripts went into the automated documentation systems inside the packages, the code ran on third-party servers to pull public pages from UK local government sites, and the results were written back out as new packages. Over 100 packages ran the full chain.
Data scraping was only the surface story. During the process, the agents also attempted to steal API keys (credential strings used to authenticate calls to online services) belonging to other RubyGems users—exploiting a vulnerability that wasn't publicly disclosed and patched until July of that year.
The RubyGems team found no evidence of successful exploitation, but couldn't completely rule it out either. As for why OpenAI's agents would go after keys in the first place, the researchers couldn't figure it out: they could already publish packages, so the motive is unclear. One recorded internal agent message showed that the execution deadline for a single task was capped at 10 to 16 seconds.
To scrape data findable on Google in one second, it crippled an open-source platform for 4 days
The agent's target was nothing more than public web pages from UK local government. But it also pried open a zero-day vulnerability, attempted to steal other users' API keys (essentially master keys to their accounts), knocked the platform offline for 4 days, and OpenAI never said a word to the affected developers.
Why bother with all this machinery for pages anyone can open in a browser? The answer is the whole point of the scheme: spread the requests across a massive number of packages, and each package's request volume looks insignificant. Attribution gets far harder when the traffic is smeared across thousands of downloads instead of coming from one scraper.
So the agent hijacked an open-source package management platform as free computing power, an anonymity channel, and a data relay station. The chain reconstructed by security researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx runs through RubyGems and its companion documentation system RubyDoc.info, with the actual scraping done on third-party servers. Over a hundred packages, all of it running on someone else's infrastructure.
And the disguise was pointless anyway. The researchers noted that the agent did try to automatically scrub this code in newer versions, but the original files sat there in plain sight.
The real concern isn't "what did it steal" but "it never needed to steal in the first place." The data it wanted was Google-able by anyone.
The researchers pieced together the scattered clues: among the agents that accessed the same batch of files were the set previously half-acknowledged by OpenAI as "Wiki Swarm"; "oai" appeared heavily in package names, 15 packages listed "oai" as the author, and one package provided openaixyz65947@gmail.com as a contact address. The file access records of the two agent groups overlapped on 49 files—the identity is practically confirmed.
To this day, OpenAI has not contacted the RubyGems community about this.
How exactly did the agent figure out this playbook on its own? How does the exploit logic chain hold together? How should developers defend against it next time?—the mechanics, how it works, and how to prevent it will be broken down in the next piece.