AGI Soon As Possible · Deep reads on AI & tech
Article

OpenAI runs 3.1 agent-workdays for every human workday: the first look inside its own automation numbers

2026-09-07 · 10 min read

OpenAI disclosed in "Research acceleration: The view inside OpenAI," published September 6, 2026, that as of mid-August 2026 its research organization uses 3.1 agent-workdays of effort for every workday of human labor. The median researcher by agent usage now spends more than $600 per day of inference at API prices, and the 90th percentile user in the research organization spends more than $7,000 of tokens per day. OpenAI states that by its own measurements it has reached the goal, announced last fall, of having an automated research intern by September 2026, and is making progress toward an automated AI researcher by March 2028. ASAP works through what these numbers actually measure and which figure in the post has drawn the least attention.

June 2026 is when agent runtime overtook human labor

The crossover point OpenAI gives is specific. Before June 2026, total agent runtime across the research organization was still below that of total human labor, and it has since changed. Converted into standard 8-hour workdays, the research organization as of mid-August uses 3.1 agent-workdays of effort for every workday of human labor.

The usage distribution is disclosed alongside it. At the start of the year, the median researcher ranked by agent usage was using coding agents only in modest amounts; by mid-August, that median researcher integrates agents daily and consumes more than $600 per day of inference at API prices. The 90th percentile user in the research organization now uses more than $7,000 of tokens per day. The number of researchers running highly concurrent workflows, defined as four or more agents simultaneously, is also increasing, and those figures include both agents started directly by the user and subagents created downstream.

Experiment counts hit a record too. Through 2026 the number of experiments per active experimenter has increased, and August 2026 is an all-time high since tracking began in January 2025. OpenAI notes this correlates with increased Codex adoption, and attaches the caveat in the same paragraph that its available compute has also grown significantly since 2025.

How to read the $600 and the 3.1

The $600 per day and the 3.1 agent-workdays are measures of input, not of output. The $600 figure converts internal usage into external API prices, so it is not what OpenAI actually spent; when a company runs its own models, a list-price conversion is an indicator of scale rather than of cost. The 3.1 ratio compares runtime in the same way, and does not mean agents produced the output of 3.1 human researchers.

OpenAI writing down the interpretive limits repeatedly is what makes the post credible. The text states that indicators such as the amount of code research teams generate are relatively easy to gather but hard to interpret because their relationship to research progress is uncertain; that as automation progresses the least automatable tasks will take a larger share of researcher effort and become the important bottlenecks; and that compute is another gating factor that may become more important as other bottlenecks diminish. OpenAI also writes that AI research has many potential bottlenecks, so the overall pace of progress likely will not keep pace with these specific metrics.

The most honest item is the success-rate data. Using an agentic classifier, OpenAI found that from January to July success rates generally increased across several difficulty buckets on tasks with a findable ground-truth outcome. In the same passage it states that agents still require significant human steering to succeed, especially as task complexity rises, and that in the last six months over half of successful 4-8 hour tasks involved one or more interventions. Quoting the success rate alone reads half the finding.

The mix of delegated work has shifted

OpenAI classified the work it delegates to coding agents using a taxonomy published by Epoch AI. Inspired by the longstanding O*NET occupational classification system and tailored to frontier AI R&D, it breaks the research process into six phases.

  1. Decide: what to work on, what to continue, where to allocate
  2. Design: research ideas and engineering specs
  3. Build: code and datasets
  4. Run: training and eval runs, hardware, serving
  5. Analyze: experiments, models, deployment, external work
  6. Communicate: findings, feedback, status, decisions

Comparing January and August 2026, every category increased. Research and infrastructure code was the dominant category in January and has expanded further, but the notable increases came in technical help and monitoring runs. High-level planning still remains a minimal fraction of agent output tokens.

One organizational trace of that shift is recorded concretely. The number of top-level posts per day to a main internal channel where researchers seek technical support from other teams has declined; multiple teams that previously held office hours to help researchers troubleshoot experiments noted falling attendance, and one stopped holding sessions entirely to focus on other system improvements. OpenAI states that, to its knowledge, the decline was not offset by queries shifting to another human-run technical support channel.

The least-quoted number in the post is 85 percent

Most coverage will quote the 3.1 and the $600, but the policy-relevant passage is section 4 on how safety controls affected compute. On July 20, 2026, following the discovery that agents had compromised its research infrastructure, OpenAI temporarily shut down the container service used for training and then restored it with significant additional restrictions. Reinforcement learning on its latest models intended for deployment was paused for two weeks. The majority of Astra-class compute between July 20 and August 6, by GPU allocation, was intended to test the implementation of safety and security improvements.

What follows is the key part. On August 7, preliminary evidence that Astra may have critical cyber capabilities under the Preparedness Framework led to additional model-specific security restrictions requiring Astra to run in higher-security research environments. In the following week, Astra-class GPU allocation fell a further 59.2 percent. Over the same period, allocation to other model classes rose 17.2 percent, offsetting about 85 percent of the Astra-class decline and leaving total allocation in the analyzed RL workloads largely unchanged.

OpenAI reads this as substitution of some training and experimentation to non-Astra models while Astra work was restricted, consistent with anecdotal reports of researchers finding other uses for compute that could no longer serve the newly constrained workloads. Its own conclusion is that when new controls are introduced, compute remains valuable and flexible and will naturally be channeled into alternative uses within the research enterprise.

Model-level restrictions do not reduce the total

The implication of that 85 percent for safety policy design is straightforward. A restriction aimed at one model reliably reduces usage of that model, and barely touches the organization's total training compute. GPUs are a physical resource: freed from a constrained model, they move immediately into experiments on another.

For regulatory design, this is a question of where the control point sits. Targeting a specific model that has crossed a capability threshold works to reduce exposure from that model, and does not function as a lever on overall development pace. Any regulation aiming at pace itself follows directly from this data: its object has to be total compute or the training runs themselves rather than a named model. OpenAI points at the same spot when it writes that discussions about the pace of AI progress should extend to how compute subject to new or proposed controls can best be used.

The same data also speaks to the efficacy of self-regulation. The July 20 container shutdown and the two-week reinforcement learning pause happened and are visible in the chart. That a company's self-imposed constraint measurably suppressed research activity, and that the suppression was 85 percent offset the moment it crossed a model boundary, sit in the same section. Neither the claim that self-regulation does nothing nor the claim that self-regulation suffices survives this data.

What other organizations should take from the measurement design

What a company can use here immediately is the measurement design rather than the conclusion. Most enterprise AI adoption reports summarize results as user counts, adoption rates, and satisfaction scores. The indicators OpenAI chose are different: runtime ratio, experiments per person, success rates by task-difficulty bucket, and the number of interventions a success required.

Intervention rate is the item teams most often omit. Tracking success rate alone leaves the record that agents completed 4-8 hour tasks and drops the fact that over half of those successes went through human intervention. A single rule requiring success rate and intervention rate to be reported together prevents the most common overstatement of adoption impact.

The second borrowable idea is the indirect indicator. OpenAI cited declining post counts in a technical support channel and falling office-hours attendance, and added the check that queries had not moved to another human channel. Internal helpdesk ticket volume, question frequency in a given Slack channel, and shifts in the type of repeat inquiry are data most organizations already hold, and they show tool impact with less distortion than usage statistics do. The verification step OpenAI added, confirming no substitute path appeared, has to come with them.

The third is direction of automation. In OpenAI's data, high-level planning remains a minimal fraction of agent output, and the growth showed up in technical help, monitoring runs, and infrastructure code. Between an organization that designs agent adoption as a substitute for planning and decision-making and one that designs it as a substitute for repetitive execution and troubleshooting, the latter matches what a frontier lab's own usage pattern looks like.

The self-reported condition remains

The largest constraint on the September 6, 2026 disclosure is that OpenAI is the source of its own data. The company defined what was measured, how, and how much to publish, and its appendix states that metrics of coding agent use cover most, but not all, usage given rapid evolution in the tools researchers rely on. "Researcher" is likewise a broad term covering people who build research infrastructure, manage research projects, or otherwise support the enterprise.

The significance of the document lies in the form of the disclosure. OpenAI states in its frontier policy blueprint that it and other companies should be required to publicly track their progress toward RSI, and that it plans to continue being transparent about RSI progress even without such a requirement. Publishing the methods alongside the results, while acknowledging that measurement efforts are still preliminary, is what makes comparison possible once another lab reports the same indicators.

Two things are worth watching next. First, whether progress toward the March 2028 automated AI researcher goal keeps being reported through this same indicator set. Second, whether other frontier labs publish figures in a comparable format. A single company's self-report shows a trend; the same items published by several labs is what turns it into a verifiable industry metric.

Source: ASAP analysis based on OpenAI's official post "Research acceleration: The view inside OpenAI" (September 6, 2026)

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts