Microsoft — Engineering Performance
136 engineers now deliver what 526 would have in Apr 2025.
▲ +184 vs the previous 90 days
Performance snapshot
Today's rolling 90-day reading for Microsoft, compared with the start of the series. Pick a window to move that comparison point.
Avg. perf / dev / mo
+287.1%
1.35 → 5.22 ETV
Active engineers
−11.7%
154.0 → 136.0
Features
+2.3pp
33.0% → 35.3%
vs. 500 OSS index
1.6x
1.5x → 1.6x · +62% above
Engineering capacity
Effective engineers behind Microsoft, against its pre-AI baseline. Each subject has its own: Microsoft's is 1.35 ETV / dev / mo, its first reading in April 2025. Per-engineer ETV divided by that gives a capacity multiple, and that multiple applied to the engineers active in the trailing 90 days turns it into engineer-equivalents. The line is the real headcount, so the gap between line and area is what the leverage is worth. Because each baseline is its own, every subject opens at 1.0x on its first day: multiples measure improvement and are not comparable between subjects.
Microsoft vs. 500 OSS Performance Index
Per-engineer ETV for Microsoft against the pooled 500 OSS Performance Index. Both lines are 90-day rolling averages scaled to a 30-day month, so they share one axis and can be read against each other at any point. Pick a window to zoom the chart to it. Latest reading: Microsoft is 62% above the index (5.22 vs 3.22 ETV/dev/mo). At the start of the tracked period the gap was 47% above.
Behind the numbers
Written summary of the work completed each month.
During 2026-08, the team delivered major improvements across the Agent Host subsystem, Copilot session management, and remote workbench integrations across 1,810 commits. Total output reached 814, surging +60% compared to the 5-month average (510) and exceeding recent months. Delivery was highlighted by a significant increase in testing activity (295 vs 129 5-month average) alongside protocol upgrades and cross-platform UI enhancements.
Highlights
- Upgraded the Agent Host subsystem to Agent Host Protocol (AHP) version 0.9.0 and normalized automation state schemas in 40c9634e (Connor)
- Implemented dynamic resource label home registration to format session paths cleanly in editor breadcrumbs via f291f3fd (Sandeep) and 718038e1 (Sandeep)
- Added support for rich content rendering including Adaptive Cards and Markdown in CmdPal details views via e1fb1333 (Mike)
- Enhanced Model Context Protocol (MCP) app routing by routing subrequests directly through owning Agent Host connections in 006c4ca9 (Dmitriy)
- Introduced active unarchived session counts and accessible row labels for remote Agent Hosts in c3f2a06d (Dmitriy)
- Migrated
devtools-viewtest suites from Jest to Mocha and JSDOM, eliminating a component memory leak in 433c666b (Joshua)
Observations
- Tests score surged +128% compared to the 5-month average (295 vs 129 average), driven by test suite migrations and broad fixture additions across agent subsystems
- Fixes score increased +74% relative to the 5-month average (145 vs 83 average), addressing memory leaks in pseudoterminals 4d3de9df (Simon) and git branch protection feaa8c99 (Simon), alongside path handling bugs 051440b1 (Ulugbek)
- Docs score saw a candidate-material increase of +101% against the 5-month average (28 vs 14 average), supported by architecture specifications like session automations in 1dbbce1d (Ulugbek)
- The Agent Host and Copilot session components experienced repeated modifications to error handling and permissions, including disabling retryable error markers in e05bae3b (roblourens) and enforcing shell-script safety classifiers in 07864e41 (joshspicer)
- Cross-platform normalization required repeated targeted fixes across Windows path handling in 051440b1 (Ulugbek), WSL chat link resolution in 735c7376 (Dileep), and CLI output formatting fallbacks in 05b0e8ce (Dmitriy)
Based on 1,810 commits491.1 ETVUpdated Sep 8, 2026, 8:17 AM
Repositories
Where each repository stands: average performance per engineer per month over the last 90 days, with the rolling 90-day curve behind it. The range picks the window (past 90d): it sets how much of the curve you see, the Δ across it, and the span the work mix is measured over. Ranked highest first.
vscode
Avg performance
8.82 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+97.9%
Work mix
Features
36%
Maint
12%
Tests
33%
Docs
2%
Fixes
17%
Agents-for-net
Avg performance
6.77 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-1.7%
Work mix
Features
31%
Maint
14%
Tests
47%
Docs
5%
Fixes
4%
TypeScript
Avg performance
4.49 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+3.7%
Work mix
Features
42%
Maint
11%
Tests
30%
Docs
2%
Fixes
15%
playwright
Avg performance
4.20 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-20.7%
Work mix
Features
31%
Maint
19%
Tests
19%
Docs
5%
Fixes
26%
PowerToys
Avg performance
3.58 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+226.0%
Work mix
Features
33%
Maint
9%
Tests
28%
Docs
3%
Fixes
27%
DeepSpeed
Avg performance
2.08 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+268.0%
Work mix
Features
31%
Maint
3%
Tests
40%
Docs
5%
Fixes
21%
fluentui
Avg performance
1.52 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-17.7%
Work mix
Features
46%
Maint
19%
Tests
17%
Docs
7%
Fixes
11%
FluidFramework
Avg performance
0.95 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+14.1%
Work mix
Features
20%
Maint
16%
Tests
30%
Docs
23%
Fixes
10%
terminal
Avg performance
0.88 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+45.0%
Work mix
Features
47%
Maint
26%
Tests
7%
Docs
1%
Fixes
19%
semantic-kernel
Avg performance
0.44 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
-25.6%
Work mix
Features
18%
Maint
9%
Tests
45%
Docs
3%
Fixes
25%
markitdown
Avg performance
0.10 ETV
per engineer per month
Avg. dev performance / month (90-day MA)
+383.3%
Work mix
Features
0%
Maint
7%
Tests
0%
Docs
17%
Fixes
76%
Performance Composition
Each month's output split by type of work: Features (new value), Maintenance (sustaining systems), Tests, Docs, and Fixes (rework). The yellow line is output per engineer, so when it rises each engineer is delivering more, whatever the team size did. Unit: Engineering Throughput Value (ETV).
Cost per Performance Unit
−74%
If performance per engineer more than doubled, each unit of engineering performance now costs approximately 74% less than at the baseline 90-day window (ending 2025-04-01). Treat this as a direction, not a price: the exact figure depends on fully-loaded engineer cost, but which way it moved is not in doubt.
Effective Capacity Added
+390 engineers
At today's productivity, the current 136-person team delivers the performance equivalent of 526 engineers at the baseline 90-day rolling window (ending 2025-04-01). That's roughly 390 engineers worth of capacity added through productivity gains, not hiring.
CapEx vs OpEx
Each month's work split two ways. CapEx (capitalizable investment) is the work that builds the asset: Features, plus the Tests that prove it works and the Docs that explain it. OpEx (operating expense) is keeping it running: Maintenance and Fixes. The yellow line is the CapEx share, so a rising line means more of the month went into building new rather than sustaining what exists. Unit: Engineering Throughput Value (ETV).
Hours per Repository
Trailing 90-day window (64 working days). Org-level capacity is allocated to each repo by its share of org performance, then split CapEx / OpEx by that repo's own Features + Tests + Docs vs Maintenance + Fixes mix.
| Agents-for-net | 82.8% | 17.2% |
| DeepSpeed | 75.6% | 24.4% |
| TypeScript | 73.6% | 26.4% |
| FluidFramework | 73.4% | 26.6% |
| vscode | 70.7% | 29.3% |
| fluentui | 69.4% | 30.6% |
| semantic-kernel | 66.4% | 33.6% |
| PowerToys | 64.3% | 35.7% |
| playwright | 55.2% | 44.8% |
| terminal | 54.6% | 45.4% |
| markitdown | 17.2% | 82.8% |
| Total | 70.2% | 29.8% |
Feature Contribution
Who shipped each month's new feature work, as a share of that month's total. Bands of similar width mean new value is coming from across the team; one band that stays wide means most of it rests on the same person. Named engineers shipped the most Features over the period. Everyone else is grouped as Others.
Quarterly Summary
Engineers and the Features / Maintenance / Tests / Docs / Fixes mix for each quarter. Cost / Perf Unit is what one unit of performance costs against Q2'25. Eff. Capacity Added is measured against this org's own pre-AI baseline instead, its first reading in the index, so it agrees with the capacity tile and chart above.
| Quarter | |||||
|---|---|---|---|---|---|
| Q2'25 | 162 | 0% | +10 engineers | 32% Features | — |
| Q3'25 | 170 | +12% | −9 engineers | 36% Features | −7% |
| Q4'25 | 166 | −14% | +39 engineers | 27% Features | +27% |
| Q1'26 | 173 | −46% | +169 engineers | 38% Features | +67% |
| Q2'26 | 157 | −55% | +216 engineers | 40% Features | +9% |