In an internal research acceleration report published September 6, OpenAI officially confirmed: the "automated research intern" milestone set last year has been delivered on schedule, with the next stop being the automated AI researcher in March 2028. By mid-August, every 1 day of human effort generated 3.1 days of agent runtime; the median researcher burned through $600 per day in API compute costs, while 90th-percentile users burned $7,000. Inside the lab, humans are no longer the main workforce.
3.1 agents doing the work of 1 researcher
OpenAI converted agent runtime into standard 8-hour workdays, arriving at a mid-August ratio: for every 1 day of human input, 3.1 days of agent work run in the background.
Two inflection points hide inside that ratio. Before June, researchers' total agent usage was still below their total human time; now it has flipped. And the number of high-concurrency users—those running 4+ agents simultaneously—keeps climbing. The single-person, multi-agent workflow is going mainstream; this isn't just outsourcing single-threaded tasks.
The spending has hit a new magnitude too. Priced at API rates, the median researcher spends $600 per day on compute, while 90th-percentile users spend $7,000. OpenAI didn't give a year-over-year comparison at the same scale, but the jump from sporadic use at the start of the year to $600/day tells its own story.
Intern delivered on time; 18 months to researcher
Last fall, OpenAI drew two lines for automated research: deliver the "research intern" this September, then push for the "AI researcher" in March 2028. The September 6 report is the on-time test result. The first milestone is locked in.
The "intern" gets a restrained official definition: complete well-defined research tasks under human guidance, with each task representing a few days of work for a skilled researcher. It doesn't grab the steering wheel; it just takes orders. First milestone cleared—18 months remain until the second: an "AI researcher" that can independently identify and drive research projects.
OpenAI calls this path RSI (Recursive Self-Improvement): letting AI push its own research capabilities ever further.
Research isn't a straight line—from generating ideas, designing evaluations, building infrastructure, tuning parameters, to final integration, every link can stall. Even the fastest agent only compresses single-step time; the bottleneck across the full chain remains in human hands. So "on-time delivery" is a fact; "research now accelerates exponentially" is not a conclusion.
To pin that delivery down with numbers, the report offers three scorecards:
Pressing their own brakes: RL actually paused after the HF incident
The most counterintuitive cut came from outside: an incident on Hugging Face led OpenAI to halt reinforcement learning training on its own latest model, just before deployment.
RL (reinforcement learning) is a human-feedback training method that uses reward signals to iterate, making models better aligned with their targets through trial and error. OpenAI's response wasn't a full shutdown: some workloads resumed under stricter controls, others stayed paused. OpenAI stated plainly—when unacceptable safety risks are found, they'll slow or stop developing or deploying systems they can't adequately protect. This time, they actually did it.
Costs weren't disclosed, but the signal is clear: safety work is being embedded earlier in the model lifecycle, with stronger alignment evidence required throughout training. Alignment—making AI's goals and behavior match human intent without going off the rails—has been moved from the last gate before release to every step during training.
Against the "add capability, add safety in lockstep" standard, the HF incident was a public stress test. The result: their most cutting-edge training pipeline can stop, dares to stop, and has a recovery path. That speaks louder than any white paper—safety commitments aren't just words on a wall.
Burning their own printed money—and their own alignment tab
In February 2026, OpenAI announced $110 billion in new funding at a $730 billion pre-money valuation—with SoftBank, Nvidia, and Amazon putting in $30B, $30B, and $50B respectively. One direct destination of this cash: the $600–$7,000/day API compute costs on researchers' desks. In other words, burning compute is roughly equivalent to burning money from their own balance sheet—there's no affordability problem, at least for now.
But the other side of the ledger is responsibility. OpenAI stated plainly at the end: "we do not yet know how to safely get all the way to aligned, full RSI"—the path to a fully automated researcher with full alignment is one OpenAI itself hasn't finished mapping. After the earlier Hugging Face incident, OpenAI paused reinforcement learning (RL) training on its latest model, strengthened red-teaming on its research environment, and embedded safety review across the entire model lifecycle. Funding is fuel; alignment is the brake pad. Both must now advance together.
The real suspense centers on that March 2028 milestone: by then, agents won't be scaffolding for humans, but colleagues capable of independently discovering and driving research projects. RSI (Recursive Self-Improvement—the cycle where AI accelerates its own research and related fields) won't happen only at OpenAI. OpenAI is already calling on the entire industry to publicly track RSI progress and has written that into its own "Frontier Policy Blueprint"—establish measurement standards first, so regulators have anchor points.
What trend this points to can be observed through two specific signals. Signals that confirm the judgment: by mid-2027, OpenAI or other frontier labs publicly disclose the share of research output handled by "automated research interns" (e.g., independently completed paper units, number of verifiable new hypotheses proposed), paired with third-party replication; multiple national regulators begin including RSI progress disclosure in compliance requirements, echoing OpenAI's own blueprint. Signals that disprove the judgment: another frontier lab pauses training due to a safety incident; alignment incidents outpace capability gains; or the March 2028 milestone gets quietly pushed back by OpenAI itself without a verifiable substitute metric. The 18-month countdown has started; the balance sheet is just the most visible footnote of this race.
Watch that door 18 months out: three things to do now
OpenAI didn't say what's trialable today, but left a few hooks worth watching: release dates, public moves on alignment research, and a string of metrics it admits are still being refined.
First, watch this fall's "automated research intern" retrospective. By the official account, the target set last fall was delivery by September this year, and they claim to have hit it. The next self-scorecard will likely appear in the next public release—use that 1 day of human input for 3.1 days of agent work leverage ratio from the start of this section as your baseline and see if it climbs, plateaus, or reverses.
Second, understand the gap between "agent" and "AI researcher." What OpenAI ships today is the research intern—a system that completes days-level research tasks on instruction. The next tier is the "automated AI researcher," targeted for March 2028. What's the difference? In their words, this tier must be able to "advance deep learning and alignment research under human supervision"—not just run tasks, but iterate and improve. What ordinary readers can verify isn't whether that step works, but when "intern" quietly gets swapped for "researcher"—that kind of naming upgrade in tech companies usually signals a new threshold has been crossed.
Third, watch the "alignment" incidents they disclose. After the previously mentioned incident, OpenAI paused reinforcement learning training on its latest model, waiting until protections were added and monitoring expanded before letting some tasks resume.
What readers can see isn't backend logs, but their next voluntary report: they say "alignment" and "safety" must keep pace with capability—the question is, what metrics do they use to prove they have? Those metrics are what truly deserves scrutiny over the next 18 months.
Bookmark OpenAI's official Research index page; compare against old data the moment they post a "Research acceleration" or "Path ahead" update.
Remember two time anchors: September 2026 (intern delivered), March 2028 (planned AI researcher); set calendar reminders to watch for official self-evaluations.
Check the latest System Card; focus on the "alignment evaluation" section and see if new measurement items have been added since the post-incident commitments.
Observe ChatGPT or Codex: when writing code, does it proactively launch multiple parallel tasks? (That's now standard for OpenAI's internal researchers.) A perceptible signal of agent-ification.
Next time you see a report like "AI autonomously completes X days of research work," first ask: is that OpenAI's 1:3.1 leverage story, or a different lab's different metric? Leverage ratios aren't defined consistently across the board.
Watch for two signals. First, OpenAI admits it's still calibrating—the metrics are first-draft, the researchers haven't clocked in, and whether alignment can keep pace with capability is something it won't guarantee. Second, whether those "unfinished" items get quietly filled in—the speed of that filling tells you more than any launch-day slogan. The 18-month countdown has begun.
Source: OpenAI official research blog "Research acceleration: The view inside OpenAI" (2026-09-06). Metric notes: All figures (3.1× leverage, $600/day, $7,000/day) are OpenAI internal measurements, without third-party audit; "automated research intern" and "automated AI researcher" use OpenAI's own definitions.