OpenAI — Engineering Performance
28 engineers now deliver what 373 would have in Apr 2025.
▲ +77 vs the previous 90 days
Performance snapshot
Today's rolling 90-day reading for OpenAI, compared with the start of the series. Pick a window to move that comparison point.
Avg. perf / dev / mo
+1233.6%
0.89 → 11.82 ETV
Active engineers
+366.7%
6.0 → 28.0
Features
−19.1pp
46.4% → 27.3%
vs. 500 OSS index
3.7x
0.97x → 3.7x · +267% above
Engineering capacity
Effective engineers behind OpenAI, against its pre-AI baseline. Each subject has its own: OpenAI's is 0.89 ETV / dev / mo, its first reading in April 2025. Per-engineer ETV divided by that gives a capacity multiple, and that multiple applied to the engineers active in the trailing 90 days turns it into engineer-equivalents. The line is the real headcount, so the gap between line and area is what the leverage is worth. Because each baseline is its own, every subject opens at 1.0x on its first day: multiples measure improvement and are not comparable between subjects.
OpenAI vs. 500 OSS Performance Index
Per-engineer ETV for OpenAI against the pooled 500 OSS Performance Index. Both lines are 90-day rolling averages scaled to a 30-day month, so they share one axis and can be read against each other at any point. Pick a window to zoom the chart to it. Latest reading: OpenAI is 267% above the index (11.82 vs 3.22 ETV/dev/mo). At the start of the tracked period the gap was 3% below.
Behind the numbers
Written summary of the work completed each month.
In August 2026, the team completed 849 commits with a notable surge across testing and corrective work, driving total output up 90% compared to the 5-month average (476 vs. 251 average) and exceeding the recent run rate. Delivery focused heavily on Guardian security and history preservation mechanisms, cross-platform sandbox permission normalization, and TUI session resilience. Tests score (+180%), Docs score (+293%), and Fixes score (+225%) each showed candidate-material increases relative to their 5-month baselines, while Maintenance decreased by 61%.
Highlights
- Hardened Guardian evaluation reliability across context compaction with dedicated history retention in 1c1e1778 (jif), user answer preservation in 305eed10 (jif) and 98a8425e (jif), and failed review diagnostic attachments in 13d75cd1 (jif)
- Implemented context-aware permission normalization and preapproval across Windows and POSIX environments within sandbox policies in c4350b4c (iceweasel-oai) and b51b0778 (iceweasel-oai)
- Improved TUI application resilience with automatic session reconnection in 907c34e8 (Eric), draft state preservation during disconnects in a7913390 (Eric), and agent navigation state recovery in 746798b2 (Eric)
- Expanded MCP capabilities to support package-style server names in 94cbbdda (Eric), per-tool output limits in f742dabc (pakrym-oai), and automatic header refresh on authentication failure in d9511fb7 (xl-openai)
- Optimized thread store operations and rollout processing by bypassing unarchived file scans in 2f0a5d55 (Eric) and introducing seekable compression for shared rollouts in 1cc81ca8 (jif)
Observations
- Fixes score reached 107 compared to the 33-score 5-month average (+225%), addressing recurring issues in handoff persistence [commit/c0a67186, commit/287594c6], tool turn replays 6b555c80 (Kazuhiro), and Guardian authorization retention [commit/0a12b855, commit/f98649cd]
- Tests score increased to 211 compared to the 75-score 5-month average (+180%), driven by end-to-end rollout resume coverage 8faf7252 (jif), terminal query responder validation 0ae94fdd (Eric), and PyPI supply chain provenance tests f1a806ad (Kazuhiro)
- Docs score rose to 30 compared to the 8-score 5-month average (+293%), centered on multi-language terminology migrations for voice agents across Japanese, Korean, and Chinese documentation in 70264ccb (Kazuhiro) and 9ce11028 (Kazuhiro)
- Maintenance score declined 61% compared to the 5-month average (17 current vs. 44 average), reflecting a redirection of effort toward testing and issue resolution
- Unified execution and terminal input underwent extensive security refactoring to intercept NUL bytes 39507eea (jif), retain granular permission profiles 5eea8d0d (jif), and safely preserve one-shot execution paths b836aecd (jif)
Based on 849 commits235.6 ETVUpdated Sep 8, 2026, 8:17 AM
Repositories
Where each repository stands: average performance per engineer per month over the last 90 days, with the rolling 90-day curve behind it. The range picks the window (past 90d): it sets how much of the curve you see, the Δ across it, and the span the work mix is measured over. Ranked highest first.
openai-agents-js
Avg performance
81.34 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+538.7%
Work mix
Features
17%
Maint
3%
Tests
49%
Docs
9%
Fixes
22%
openai-agents-python
Avg performance
36.43 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+241.8%
Work mix
Features
18%
Maint
2%
Tests
44%
Docs
12%
Fixes
24%
codex
Avg performance
7.29 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-10.9%
Work mix
Features
36%
Maint
13%
Tests
36%
Docs
1%
Fixes
15%
openai-dotnet
Avg performance
0.77 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-53.9%
Work mix
Features
34%
Maint
5%
Tests
49%
Docs
8%
Fixes
4%
plugins
Avg performance
0.05 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-62.2%
Work mix
Features
71%
Maint
0%
Tests
0%
Docs
29%
Fixes
0%
Performance Composition
Each month's output split by type of work: Features (new value), Maintenance (sustaining systems), Tests, Docs, and Fixes (rework). The yellow line is output per engineer, so when it rises each engineer is delivering more, whatever the team size did. Unit: Engineering Throughput Value (ETV).
Cost per Performance Unit
−93%
If performance per engineer more than doubled, each unit of engineering performance now costs approximately 93% less than at the baseline 90-day window (ending 2025-04-01). Treat this as a direction, not a price: the exact figure depends on fully-loaded engineer cost, but which way it moved is not in doubt.
Effective Capacity Added
+345 engineers
At today's productivity, the current 28-person team delivers the performance equivalent of 373 engineers at the baseline 90-day rolling window (ending 2025-04-01). That's roughly 345 engineers worth of capacity added through productivity gains, not hiring.
CapEx vs OpEx
Each month's work split two ways. CapEx (capitalizable investment) is the work that builds the asset: Features, plus the Tests that prove it works and the Docs that explain it. OpEx (operating expense) is keeping it running: Maintenance and Fixes. The yellow line is the CapEx share, so a rising line means more of the month went into building new rather than sustaining what exists. Unit: Engineering Throughput Value (ETV).
Hours per Repository
Trailing 90-day window (64 working days). Org-level capacity is allocated to each repo by its share of org performance, then split CapEx / OpEx by that repo's own Features + Tests + Docs vs Maintenance + Fixes mix.
| plugins | 100.0% | 0.0% |
| openai-dotnet | 90.5% | 9.5% |
| openai-agents-js | 75.7% | 24.3% |
| openai-agents-python | 74.2% | 25.8% |
| codex | 72.3% | 27.7% |
| Total | 73.6% | 26.3% |
Feature Contribution
Who shipped each month's new feature work, as a share of that month's total. Bands of similar width mean new value is coming from across the team; one band that stays wide means most of it rests on the same person. Named engineers shipped the most Features over the period. Everyone else is grouped as Others.
Quarterly Summary
Engineers and the Features / Maintenance / Tests / Docs / Fixes mix for each quarter. Cost / Perf Unit is what one unit of performance costs against Q2'25. Eff. Capacity Added is measured against this org's own pre-AI baseline instead, its first reading in the index, so it agrees with the capacity tile and chart above.
| Quarter | |||||
|---|---|---|---|---|---|
| Q2'25 | 18 | 0% | +1 engineer | 51% Features | — |
| Q3'25 | 25 | −61% | +45 engineers | 43% Features | +260% |
| Q4'25 | 32 | −63% | +61 engineers | 46% Features | +32% |
| Q1'26 | 38 | −86% | +250 engineers | 48% Features | +211% |
| Q2'26 | 32 | −88% | +265 engineers | 39% Features | +3% |