DeepSpeed — Engineering Performance
6 engineers all time · Jan 2025 – Sep 2026 · built 2026-09-30 · GitHub
Performance snapshot
Today's rolling 90-day reading for DeepSpeed, compared with the start of the series. Pick a window to move that comparison point.
Avg. perf / dev / mo
+611.7%
0.26 → 1.83 ETV
Active engineers
+25.0%
4.0 → 5.0
Features
+12.9pp
15.9% → 28.8%
vs. Microsoft
0.28x
0.19x → 0.28x · −72% below
DeepSpeed vs. Microsoft
Per-engineer ETV for DeepSpeed against Microsoft as a whole. Both lines are 90-day rolling averages scaled to a 30-day month, so they share one axis and can be read against each other at any point. Pick a window to zoom the chart to it.
Performance Composition
Each month's output split by type of work: Features (new value), Maintenance (sustaining systems), Tests, Docs, and Fixes (rework). The yellow line is output per engineer, so when it rises each engineer is delivering more, whatever the team size did. Unit: Engineering Throughput Value (ETV).
Engineering capacity
Effective engineers behind DeepSpeed, against its pre-AI baseline. Each subject has its own: DeepSpeed's is 0.40 ETV / dev / mo, its first reading in Q1 2025. Per-engineer ETV divided by that gives a capacity multiple, and that multiple applied to the engineers active in the trailing 90 days turns it into engineer-equivalents. The line is the real headcount, so the gap between line and area is what the leverage is worth. Because each baseline is its own, every subject opens at 1.0x on its first day: multiples measure improvement and are not comparable between subjects.
Knowledge concentration
How dependent is this repo on a small number of engineers? Higher top-1 share = higher key-person risk.
Masahiro Tanaka owns 34.2 % of commits.
Behind the numbers
Written summary of the work completed each month.
No monthly reports available yet.
Top engineers
Most impactful commits
Top 10 by ETV in the all-time window.
- 4.0ETVAdd AutoEP + AutoTP parallel folding (#8064) This PR adds **parallel folding** for AutoEP: tensor parallelism (AutoTP) for the dense/attention path can now coexist with expert parallelism (AutoEP) for the routed-expert path on the same set of ranks, **without forcing EP to be a subset of DP** — an EP group may span TP lanes and dense-DP ranks (cross-lane EP). (This PR should be adjusted for ZeRO3 support after #8060 is merged) ## Design Attention/dense and MoE are treated as two independent partitionings of the same rank set, parameterized per parameter family: - Dense / attention / shared-expert params: `stage_size = tp * dp` - Routed-expert params: `stage_size = ep * etp * edp` `dp` and `edp` are always derived, never user-configured, so the invariant `tp * dp == ep * etp * edp == stage_size` cannot be broken from config. The only structural requirement is that the expert width tiles the stage (`stage_size % (ep * etp) == 0`); EP groups are then laid across a TP-lane-major rank ordering, so they may span TP lanes and dense-DP ranks. ### Configuration No new config section. Folding is expressed by the coexistence of the existing `tensor_parallel` and `expert_parallel` sections: ```json { "tensor_parallel": { "autotp_size": 4 }, "expert_parallel": { "enabled": true, "autoep_size": 4, "expert_tensor_parallel_size": 1 } } ``` `expert_tensor_parallel_size` is carried as a config field but currently must be `1` (expert-internal TP is reserved as follow-up and rejected fail-fast). Validation enforces stage divisibility, TP/sequence-parallel exclusivity, and `preset_model` consistency between the two sections. ### Cross-lane expert parallelism Expert parallelism no longer has to be a subset of data parallelism. Shapes where the expert width exceeds (or does not divide) the dense data-parallel size are supported, for example: - `world=4, TP=4, EP=4` (`dp=1`): the EP group is the whole TP group — one expert per rank. - `world=4, TP=2, EP=4` (`dp=2`): the EP group spans both TP lanes and both DP ranks. - `world=8, TP=4, EP=4` (`dp=2, edp=2`): EP groups span TP lanes with expert replication. The per-family gradient convention is keyed to each parameter's replication structure, not to the EP layout, so it holds across the whole `tp*dp` pool: - **Router/gate and dense/LayerNorm** are AVERAGE over the TP (token-replication) group. The folded router runs redundantly on every TP peer; its partitioned work is reconstructed into a replicated full view by `restore_combined`, whose all-gather backward injects a `tp_size` factor that AVERAGE divides out. - **Routed experts** cancel that same `restore_combined` `tp_size` factor (divide by `tp_size`, no TP all-reduce) and reduce data-parallel over the expert-data-parallel (EDP) group. Without the cancellation, folded expert gradients are over-scaled by `tp_size` — invisible to scale-invariant Adam, but real for non-adaptive optimizers and for gradient clipping (it inflates the expert contribution to the global grad norm). This is now fixed for all folded shapes (the MVP TP2×EP4 shape included). ## What's included - Folded process-group derivation using the generalized expert/data-parallel group creation (`mp_mode` TP-strided vs SP-consecutive ordering), including cross-lane EP group tables. - Route-full / partition-dispatch path for folded MoE (`deepspeed/moe/ep_tp_dispatch.py`), with AutoTP skipping AutoEP subtrees. - **Per-family folded gradient reduction**: AVERAGE for replicated router/gate and dense/LayerNorm; a dedicated `tp_size` cancellation for routed experts; SKIP for genuinely TP-sharded params; SUM contracts reserved for a future true sequence-parallel path. - Per-parameter-family ZeRO checkpoint metadata (routed-expert vs dense/router/shared placement) and folded ZeRO-1/2 optimizer-state handling. ## Correctness & validation - Router/gate and LayerNorm gradient parity against a non-folded ZeRO baseline (atol=1e-1, rtol=5e-3, fp32), on TP2×EP4 (8-rank) and the cross-lane shapes TP2×EP4 (4-rank, `ep>dp`) and TP4×EP4 (4-rank, `dp=1`); scale 1.0. - Routed-expert weight parity against a non-folded baseline, verified with SGD (Adam is scale-invariant and would mask a uniform gradient-scale error), for the MVP TP2×EP4 shape and cross-lane TP4×EP4 with `edp=1` and `edp=2`. - New folding unit tests for config, cross-lane group layout, dispatch, runtime, gradient parity, and checkpoint save/load (multi-rank GPU cases gated for GPU runners; CPU/Gloo parity runs on CI). - Real-H100 confirmation (8×H100): router/gate, LayerNorm, and routed-expert gradient parity to the non-folded baseline hold (scale 1.0) for MVP TP2×EP4 and cross-lane TP4×EP4 (`edp=1` and `edp=2`); cross-lane folded training runs with finite loss and finite, non-zero expert/router gradients. - Passes the full unit test suite (`aws-torch-latest-full`) on H100 GPUs. ## Scope / follow-ups - This PR covers AutoEP + AutoTP folding, including cross-lane EP with `etp=1`. The replicated-grad reduction is mode-aware so the sequence-parallel (Ulysses) folding case fits the same contract; AutoTP + AutoEP is the validated path here. - Expert-internal tensor parallelism (`expert_tensor_parallel_size > 1`) is reserved for a follow-up. - ZeRO-3 composition with folding is planned as separate follow-up work (after #8060 is merged). --------- Signed-off-by: Masahiro Tanaka <mtanaka@anyscale.com>Masahiro Tanaka · 104193b4 · 2026-07-12
- 4.0ETVSupport AutoEP with ZeRO-3 zero.Init source modules (#8060) This PR enables ZeRO3 support for AutoEP-managed MoE layers by partitioning expert parameters over expert replica groups while router and replicated parameters use the global data-parallel group. With ZeRO3 enable, AutoEP preserves global data-parallel gradient averaging for AutoEP expert parameters while reducing them over expert replica groups. ZeRO parameters are gathered before AutoEP reads router or expert tensors when replacing MoE modules created under `deepspeed.zero.Init()`. --------- Signed-off-by: Masahiro Tanaka <mtanaka@anyscale.com> Co-authored-by: Ma, Guokai <guokai.ma@gmail.com>Masahiro Tanaka · 02663d6e · 2026-06-26
- 1.8ETVSupport custom partitioning patterns for AutoTP (#7806) This PR introduces a flexible, configuration-driven API for AutoTP (Automatic Tensor Parallelism) that allows users to define custom layer partitioning patterns for training. @inkcherry @delock ## Motivation Previously, AutoTP relied on hardcoded layer detection logic that was difficult to customize for new model architectures. This PR enables: 1. **Custom models**: Users can define exact regex patterns to match their model's parameter names 2. **Fused layers**: Support for fused QKV, gate_up_proj, and other packed weight matrices with unequal sub-parameter sizes (e.g., GQA with different Q/K/V dimensions) 3. **Extensibility**: Easy to add new model presets or customize existing ones Here is an example of a config including custom partitioning patterns: ```json { "tensor_parallel": { "autotp_size": 4, "partition_config": { "use_default_specs": false, "layer_specs": [ { "patterns": [".*\\.o_proj\\.weight$", ".*\\.down_proj\\.weight$"], "partition_type": "row" }, { "patterns": [".*\\.[qkv]_proj\\.weight$"], "partition_type": "column" }, { "patterns": [".*\\.gate_up_proj\\.weight$"], "partition_type": "column", "shape": [2, -1], "partition_dim": 0 } ] } } } ``` Refer to the [document](https://github.com/tohtana/DeepSpeed/blob/tohtana/autotp_custom_patterns/docs/code-docs/source/training.rst) for more details (including preset models and how to define partitioning for fused models). We also opened a new [PR](https://github.com/deepspeedai/DeepSpeedExamples/pull/998) to show the usage. ## Simplified initialization step AutoTP previously required calling ``set_autotp_mode(training=True)`` and ``deepspeed.tp_model_init`` before ``deepspeed.initialize``. Now we can include all the necessary configurations in the DeepSpeed config. We still support the traditional initialization path for backward compatibility. When you use both (i.e. calling ``set_autotp_mode(training=True)`` and ``deepspeed.tp_model_init`` and passing the config to ``deepspeed.initialize``), we will merge the settings at initialization. When we have conflicting settings, we will error out. --------- Signed-off-by: Masahiro Tanaka <mtanaka@anyscale.com>Masahiro Tanaka · 6b9cab1d · 2026-01-31
- 1.8ETVSplit DeepCompile ZeRO-3 memory scheduler (#8233) Splits the DeepCompile ZeRO-3 scheduling and gather/release lifetime changes out of #8169. The exact candidate passed 10/10 tests on a single node with 2 H100 GPUs. Signed-off-by: Masahiro Tanaka <tanaka.masahiro@gmail.com>Masahiro Tanaka · 38561463 · 2026-09-10
- 1.7ETV[AutoTP] Fix ZeRO-3 checkpoint consolidation to gather across TP and DP (#8168) AutoTP + ZeRO-3 silently produced incomplete checkpoints: both export paths handled only the ZeRO data-parallel dimension and dropped the tensor-parallel shards. - ds_to_universal.py: stage3 conversion recovers the (tp,dp) grid from checkpoint file names, extracts shards under the real tp_index, and reuses the stage<=2 TP-aware merge when tp_degree>1 (DP-only path preserved for tp_degree==1 -> no regression for plain ZeRO-3). - engine.py: _zero3_consolidated_16bit_state_dict nests GatherReplacedLayerParams inside GatheredParameters so save_16bit_model gathers both DP and TP; remove the blanket autotp+zero3 training block now that checkpoint consolidation is implemented. - stage3.py: load_hp_checkpoint_state resolves the TP shard before the ZeRO-DP partition, so universal checkpoint restore round-trips. Add end-to-end universal conversion tests and update existing tests for the refactored merge_tp_slices / extract_zero_shards_stage3 signatures. --------- Signed-off-by: Guokai Ma <guokai.ma@intel.com>Ma, Guokai · eec237ee · 2026-08-02
- 1.4ETVNon-reentrant activation checkpoint CPU offload (#8282) ## Summary - Continues [@winglian](https://github.com/winglian)'s work from #8181: a `saved_tensors_hooks` context manager that offloads non-reentrant checkpoint hidden-state inputs to a pinned CPU buffer pool on a side stream (`use_reentrant=False`). - The first two commits are **authored and signed off by Wing Lian** (`wing@axolotl.ai`); they are the original #8181 patches, rebased onto current `master`. This follow-up commit addresses review without rewriting those commits. - Review follow-up: restore `GradientCheckpointingLayer.__call__` when no manager is active (HybridEngine train/rollout), skip offloading the last checkpoint input (`keep_last_count=1`), rename `*_size` knobs to `*_bytes` / `*_count`, and add tests for the HF signature contract and keep-last behavior. - Wires the async offload into DeepSpeed native `cpu_checkpointing`: the copy machinery is factored into a reusable `_ActivationOffloadEngine`, which both the HF hooks class and native `CheckpointFunction` / `non_reentrant_checkpoint` share. Also fixes two pre-existing `non_reentrant_checkpoint` + `cpu_checkpointing` bugs (inputs emptied before forward; `saved_data` never restored during recompute). Original upstreaming context: axolotl-ai-cloud/axolotl#3776, requested in #8181. ## Test plan - [x] `pytest tests/unit/runtime/activation_checkpointing/test_offload_activations.py` on H200 — HF `saved_tensors_hooks` path (26 passed) - [x] `pytest tests/unit/runtime/activation_checkpointing/test_activation_checkpointing.py` on H200 — native reentrant + new `cpu_checkpointing` offload tests (30 passed) - [x] `pytest tests/unit/runtime/activation_checkpointing/test_activation_checkpointing_non_reentrant.py` on H200 — native non-reentrant + new `cpu_checkpointing` offload test (49 passed) - [x] H200 microbenchmark (activation-dominant): async CPU offload matches blocking's 58% peak-memory reduction at ~5.7% step-time overhead (vs ~6.8x for blocking), i.e. 6.4x faster than blocking offload - [ ] CI unit tests for activation checkpointing Made with [Cursor](https://cursor.com) --------- Signed-off-by: Wing Lian <wing@axolotl.ai> Signed-off-by: Olatunji Ruwase <tunji.ruwase@snowflake.com> Co-authored-by: Wing Lian <wing@axolotl.ai> Co-authored-by: Cursor <cursoragent@cursor.com>Olatunji Ruwase · 858e91ee · 2026-08-21
- 1.4ETVCI: prefer bf16 over fp16 (#7304) these days fp16 is barely ever used, so we should be testing bf16 instead of fp16 where possible. had to fix a bunch of tests to adapt to this change. a few bugs as well on the way. --------- Signed-off-by: Stas Bekman <stas.bekman@snowflake.com> Co-authored-by: Olatunji Ruwase <tunji.ruwase@snowflake.com> Co-authored-by: Stas Bekman <stas.bekman@snowflake.com>Stas Bekman · b4cc079e · 2025-05-28
- 1.4ETVEnable torch.autocast with ZeRO (#6993) DeepSpeed supports mixed precision training, but the behavior is different from `torch.autocast`. DeepSpeed maintains parameters and gradients both in FP32 and a lower precision (FP16/BF16) (NVIDIA Apex AMP style) and computes all modules in the lower precision while `torch.autocast` maintains parameters in FP32 but computes only certain operators in the lower precision. This leads to differences in: - performance: `torch.autocast` needs downcast in forward/backward - memory usage: DeepSpeed needs more memory to keep copies of parameters and gradients in lower precision - accuracy: `torch.autocast` has a list of modules that can safely be computed in lower precision. Some precision-sensitive operators (e.g. softmax) are computed in FP32. To align DeepSpeed's behavior with `torch.autocast` when necessary, this PR adds the integration with `torch.autocast` with ZeRO. Here is an examples of the configuration. ```json "torch_autocast": { "enabled": true, "dtype": "bfloat16", "lower_precision_safe_modules": ["torch.nn.Linear", "torch.nn.Conv2d"] } ``` Each configuration works as follows: - `enabled`: Enable the integration with `torch.autocast` if this is set to `True`. You don't need to call `torch.autocast` in your code. The grad scaler is also applied in the DeepSpeed optimizer. - `dtype`: lower precision dtype passed to `torch.autocast`. Gradients for allreduce (reduce-scatter) and parameters for allgather (only for ZeRO3) of `lower_precision_safe_modules` are also downcasted to this dtype. - `lower_precision_safe_modules`: Downcast for allreduce (reduce-scatter) and allgather (ZeRO3) are applied only to modules specified in this list. (The precision for PyTorch operators in forward/backward follows `torch.autocast`'s policy, not this list.) You can set names of classes with their packages. If you don't set this item, DeepSpeed uses the default list: `[torch.nn.Linear, torch.nn.Conv1d, torch.nn.Conv2d, torch.nn.Conv3d]`. Note that we only maintain FP32 parameters with this feature enabled. For consistency, you cannot enable `fp16` or `bf16` in DeepSpeed config. --------- Signed-off-by: Masahiro Tanaka <mtanaka@microsoft.com> Signed-off-by: Fabien Dupont <fdupont@redhat.com> Signed-off-by: Olatunji Ruwase <olruwase@microsoft.com> Signed-off-by: Logan Adams <loadams@microsoft.com> Signed-off-by: inkcherry <mingzhi.liu@intel.com> Signed-off-by: Omar Elayan <oelayan@habana.ai> Signed-off-by: Roman Fitzjalen <romaactor@gmail.com> Signed-off-by: Hongwei <hongweichen@microsoft.com> Signed-off-by: shaomin <wukon1992@gmail.com> Signed-off-by: Stas Bekman <stas@stason.org> Signed-off-by: siqi <siqi@tecorigin.com> Signed-off-by: Wei Wu <wuwei211x@gmail.com> Signed-off-by: ShellyNR <shelly.nahir@live.biu.ac.il> Signed-off-by: Lai, Yejing <yejing.lai@intel.com> Co-authored-by: Olatunji Ruwase <olruwase@microsoft.com> Co-authored-by: Logan Adams <114770087+loadams@users.noreply.github.com> Co-authored-by: Fabien Dupont <fabiendupont@fabiendupont.fr> Co-authored-by: Liangliang Ma <1906710196@qq.com> Co-authored-by: inkcherry <mingzhi.liu@intel.com> Co-authored-by: Omar Elayan <142979319+oelayan7@users.noreply.github.com> Co-authored-by: Stas Bekman <stas00@users.noreply.github.com> Co-authored-by: Roman Fitzjalen <romaactor@gmail.com> Co-authored-by: Ramya Ramineni <62723901+rraminen@users.noreply.github.com> Co-authored-by: Guanhua Wang <alexwgh333@gmail.com> Co-authored-by: root <root@ftqtmec25000000.taxzvufipdhelhupulxcbvr15f.ux.internal.cloudapp.net> Co-authored-by: Hongwei Chen <33092912+hwchen2017@users.noreply.github.com> Co-authored-by: Joe Mayer <114769929+jomayeri@users.noreply.github.com> Co-authored-by: wukong1992 <wukong1992@users.noreply.github.com> Co-authored-by: shaomin <wukon1992@gmail.com> Co-authored-by: loadams <loadams@users.noreply.github.com> Co-authored-by: siqi654321 <siqi202311@163.com> Co-authored-by: siqi <siqi@tecorigin.com> Co-authored-by: Wei Wu <45323446+U-rara@users.noreply.github.com> Co-authored-by: Shelly Nahir <73890534+ShellyNR@users.noreply.github.com> Co-authored-by: snahir <snahir@habana.ai> Co-authored-by: Yejing-Lai <yejing.lai@intel.com> Co-authored-by: Siddharth Singh <siddharth9820@gmail.com> Co-authored-by: Olatunji Ruwase <tjruwase@gmail.com>Masahiro Tanaka · ed5f7375 · 2025-06-19
- 1.4ETVPyTorch-compatible backward API (#7665) Currently DeepSpeed's backward API has more constraints compared to PyTorch's normal backward API. Here is the usage as described in the documentation: ```python loss = model_engine(batch) model_engine.backward(loss) ``` In this example, 1. Only accepts a (scalar) loss value 1. Need to call engine's backward API In contrast, in standard PyTorch, you can do: ```python output = model(batch) output.backward(out_grad) ``` There are several use cases that rely on this flexibility. For example, combining multiple models or using loss functions defined separately from the main model. If you attempt the same pattern with a DeepSpeed engine, some preprocessing and postprocessing steps will be silently skipped, which can lead to incorrect results. The [document](https://deepspeed.readthedocs.io/en/latest/training.html#jointly-training-models-with-shared-loss) explains we can call `_backward_epilogue` manually (possibly `backward_prologue` as well). However, it's easy for users to miss these calls, and passing a non-scalar gradient is still not supported. This PR introduces the same `.backward()` behavior as PyTorch, allowing .backward() to be called directly on tensors and supporting non-scalar outputs. To implement post-backward hooks, we had to use some torch internal APIs. See [comments](https://github.com/deepspeedai/DeepSpeed/blob/73f7ff1aab9d1387eb7dd4eca7453a25024533f4/deepspeed/runtime/engine.py#L424) for more details. When the internal APIs are not available, DeepSpeed engine only accepts the traditional way `model_engine.backward(loss)`. --------- Signed-off-by: Masahiro Tanaka <mtanaka@anyscale.com>Masahiro Tanaka · 53e91a09 · 2025-11-19
- 1.3ETV[AutoTP] Replace tp_shard process-wide globals with per-model AutoTPMeta (#8241) ## What this does Fixes #8231. `tp_shard` kept `num_kv_heads` / `num_attention_heads` / `n_embd` / `tp_grain_size` as **process-wide mutable globals**, written during AutoTP replacement. A second AutoTP model loaded into the same process overwrote them, so the first model's later sharding / gather / checkpoint conversion silently read the wrong values — making it unsafe to run more than one AutoTP model per process (teacher/student, online distillation, RL actor + reference). This moves that state onto a per-model `AutoTPMeta`, computed once from the model config and threaded through every sharding helper and TP layer, so each model carries its own kv-head / grain state. ## Stacking / merge order **Depends on #8185** (AutoTP uneven sharding). This branch is based on #8185's head and is opened as **draft** until #8185 lands — GitHub will drop #8185's commits from this diff automatically once it merges. Please **merge after #8185**. ## Changes - `AutoTPMeta` dataclass + `from_model_config` (single source for kv-head / attn-head / hidden extraction); `get_shard_size(_list)` take it as a required arg; tp_shard globals + `set_*` / `get_*` removed. - AutoTP threads `tp_meta` from `__init__` through every TP layer and fused-QKV helper. - Ulysses sequence parallelism gets its own `_ulysses_num_kv_heads`, decoupled from AutoTP (the AutoTP↔Ulysses coupling is gone; Ulysses's own multi-model case is left for a separate change). - Inference engine builds one `meta` per model (`_autotp_meta`) and threads it through the alibi head-sharding helpers; `_get_model_head_count` / `_get_model_kv_head_count` deleted. - kv-head / attn-head attribute lists unified behind `_kv_head_count_from` / `_attention_head_count_from` (covers chatglm, falcon, llama-class, dbrx, legacy `n_head_kv`). ## Tests Validated on 4×RTX 4080 (nccl) against #8185: full AutoTP / SP / checkpoint suite passes (127 passed); remaining failures are pre-existing env issues (transformers/HF network `client has been closed`, torch 2.12 `ProcessGroupGloo.perform_nocolor_split`, a cuda/cpu device-mismatch), each confirmed failing on the #8185 baseline too. `test_two_models_do_not_clobber_each_others_meta` is the direct regression test for #8231. --------- Signed-off-by: Guokai Ma <guokai.ma@intel.com> Signed-off-by: Ma, Guokai <guokai.ma@intel.com>Ma, Guokai · 183c7f95 · 2026-08-31