Navigara
All Organizations

OpenAI — Engineering Performance

OpenAIEngineering capacity
Navigara
+345Eff. engineering capacity added

28 engineers now deliver what 373 would have in Apr 2025.

+77 vs the previous 90 days

Effective engineersReal engineers
Apr 2025Dec 2025Sep 2026
Where the work wentBiggest shift: Tests up 13 points
27%
Features
8%
Maintenance
41%
Tests
5%
Docs
19%
Fixes

Performance snapshot

Today's rolling 90-day reading for OpenAI, compared with the start of the series. Pick a window to move that comparison point.

Avg. perf / dev / mo

+1233.6%

0.89 → 11.82 ETV

Active engineers

+366.7%

6.0 → 28.0

Features

−19.1pp

46.4% → 27.3%

vs. 500 OSS index

3.7x

0.97x → 3.7x · +267% above

Engineering capacity

Effective engineers behind OpenAI, against its pre-AI baseline. Each subject has its own: OpenAI's is 0.89 ETV / dev / mo, its first reading in April 2025. Per-engineer ETV divided by that gives a capacity multiple, and that multiple applied to the engineers active in the trailing 90 days turns it into engineer-equivalents. The line is the real headcount, so the gap between line and area is what the leverage is worth. Because each baseline is its own, every subject opens at 1.0x on its first day: multiples measure improvement and are not comparable between subjects.

Effective engineers
373engineer-equivalents
+345 engineers above real headcount
Real engineers
28engineers
active in the trailing 90 days
Capacity vs pre-AI
13.3x
per engineer, vs 0.89 pre-AI

OpenAI vs. 500 OSS Performance Index

Per-engineer ETV for OpenAI against the pooled 500 OSS Performance Index. Both lines are 90-day rolling averages scaled to a 30-day month, so they share one axis and can be read against each other at any point. Pick a window to zoom the chart to it. Latest reading: OpenAI is 267% above the index (11.82 vs 3.22 ETV/dev/mo). At the start of the tracked period the gap was 3% below.

OpenAI
11.82ETV / dev / mo
+4.33 (+57.8%) past 90d
500 OSS Performance Index
3.22ETV / dev / mo
+0.99 (+44.4%) past 90d

Behind the numbers

Aug

Written summary of the work completed each month.

In August 2026, the team completed 849 commits with a notable surge across testing and corrective work, driving total output up 90% compared to the 5-month average (476 vs. 251 average) and exceeding the recent run rate. Delivery focused heavily on Guardian security and history preservation mechanisms, cross-platform sandbox permission normalization, and TUI session resilience. Tests score (+180%), Docs score (+293%), and Fixes score (+225%) each showed candidate-material increases relative to their 5-month baselines, while Maintenance decreased by 61%.

Highlights

Observations

  • Fixes score reached 107 compared to the 33-score 5-month average (+225%), addressing recurring issues in handoff persistence [commit/c0a67186, commit/287594c6], tool turn replays 6b555c80 (Kazuhiro), and Guardian authorization retention [commit/0a12b855, commit/f98649cd]
  • Tests score increased to 211 compared to the 75-score 5-month average (+180%), driven by end-to-end rollout resume coverage 8faf7252 (jif), terminal query responder validation 0ae94fdd (Eric), and PyPI supply chain provenance tests f1a806ad (Kazuhiro)
  • Docs score rose to 30 compared to the 8-score 5-month average (+293%), centered on multi-language terminology migrations for voice agents across Japanese, Korean, and Chinese documentation in 70264ccb (Kazuhiro) and 9ce11028 (Kazuhiro)
  • Maintenance score declined 61% compared to the 5-month average (17 current vs. 44 average), reflecting a redirection of effort toward testing and issue resolution
  • Unified execution and terminal input underwent extensive security refactoring to intercept NUL bytes 39507eea (jif), retain granular permission profiles 5eea8d0d (jif), and safely preserve one-shot execution paths b836aecd (jif)

Based on 849 commits235.6 ETVUpdated Sep 8, 2026, 8:17 AM

Performance Composition

Each month's output split by type of work: Features (new value), Maintenance (sustaining systems), Tests, Docs, and Fixes (rework). The yellow line is output per engineer, so when it rises each engineer is delivering more, whatever the team size did. Unit: Engineering Throughput Value (ETV).

Cost per Performance Unit

−93%

If performance per engineer more than doubled, each unit of engineering performance now costs approximately 93% less than at the baseline 90-day window (ending 2025-04-01). Treat this as a direction, not a price: the exact figure depends on fully-loaded engineer cost, but which way it moved is not in doubt.

Effective Capacity Added

+345 engineers

At today's productivity, the current 28-person team delivers the performance equivalent of 373 engineers at the baseline 90-day rolling window (ending 2025-04-01). That's roughly 345 engineers worth of capacity added through productivity gains, not hiring.

CapEx vs OpEx

Each month's work split two ways. CapEx (capitalizable investment) is the work that builds the asset: Features, plus the Tests that prove it works and the Docs that explain it. OpEx (operating expense) is keeping it running: Maintenance and Fixes. The yellow line is the CapEx share, so a rising line means more of the month went into building new rather than sustaining what exists. Unit: Engineering Throughput Value (ETV).

Hours per Repository

Trailing 90-day window (64 working days). Org-level capacity is allocated to each repo by its share of org performance, then split CapEx / OpEx by that repo's own Features + Tests + Docs vs Maintenance + Fixes mix.

plugins100.0%0.0%
openai-dotnet90.5%9.5%
openai-agents-js75.7%24.3%
openai-agents-python74.2%25.8%
codex72.3%27.7%
Total73.6%26.3%

Feature Contribution

Who shipped each month's new feature work, as a share of that month's total. Bands of similar width mean new value is coming from across the team; one band that stays wide means most of it rests on the same person. Named engineers shipped the most Features over the period. Everyone else is grouped as Others.

Quarterly Summary

Engineers and the Features / Maintenance / Tests / Docs / Fixes mix for each quarter. Cost / Perf Unit is what one unit of performance costs against Q2'25. Eff. Capacity Added is measured against this org's own pre-AI baseline instead, its first reading in the index, so it agrees with the capacity tile and chart above.

Quarter
Q2'25180%+1 engineer
51% Features
Q3'2525−61%+45 engineers
43% Features
+260%
Q4'2532−63%+61 engineers
46% Features
+32%
Q1'2638−86%+250 engineers
48% Features
+211%
Q2'2632−88%+265 engineers
39% Features
+3%