Alexey Spiridonov
90d · built 2026-09-30
Performance
What Alexey Spiridonov shipped in the selected window, measured in ETV, and how it compares with the 90 days before it.
Effective capacity
+2.4engineers
delivers like 3.4 (3.4x pre-AI)
Output (ETV)
18.4ETV
+3238.2% vs 0.6 prior
Features share
36.8%
−19.6 pp vs prior window
Fixes share
5.7%
+5.7 pp vs prior window
Work mix
36.8% Features4% Maintenance30.4% Tests23% Docs5.7% Fixes
124 commits over 90 days, ending 2026-09-30.
Daily performance
Daily ETV, stacked by Features, Maintenance, Tests, Docs and Fixes.
Repository spread
Where this developer's commits land. Concentrated work (top1 > 80%) vs polymath spread (top1 < 30%).
Most impactful commits
Top 10 by ETV in the last 90 days.
- 2.1ETV`benchmark_ab.py`: multi-binary benchmark A/B reports Summary: `benchmark_ab.py` is a new tool to simplify iterating on changes that affect several benchmark binaries. It handles three recurring chores: - Finds benchmark binaries from Buck target patterns. - Aggregates repeated A/B runs into one report. - Prioritizes wins and regressions with absolute and relative thresholds. Read the file docblock for more. ```text $ folly/tool/benchmark_ab.py measure --before=bc56e16776 --after=0097974dca \ //folly/result/... ... High-priority regressions: 15.2+11.2ns (+73.5%): try_to_result_error fbcode//folly/result/test:result_bench 15.2+11.2, 15.2+11.3, 15.2+11.1, 15.2+11.2, 15.4+11.2 ``` Reviewed By: janondrusek Differential Revision: D112568029 fbshipit-source-id: 544b93f75660e791285743188150a3eec38b1e1egithub.com-facebook-folly · 206fc16e · 2026-07-28
- 1.1ETVAdd repeatable backtests for rule changes Summary: Reading a rule diff cannot show whether it improves the artifact an agent produces. Add a backtest runner that stages a fixed scenario and explicitly selected rules, then preserves the artifact, trace, errors, and run metadata. The runner refuses local changes to the source files that shape the run and records their checkout revision, so a later comparison can identify and reproduce those inputs. Model, reasoning effort, and executable paths remain explicit metadata. ___ Differential Revision: D119108233 fbshipit-source-id: db92be24e462f47d7c8a92b9bb31b17b8ba80707github.com-facebook-folly · 8883a5d3 · 2026-09-08
- 1.1ETVPreserve critic-iterate drafts and costs in backtests Summary: Backtests currently preserve only the final answer. Rule maintainers cannot tell what the author fixed unaided, what each external review changed, or what another round cost. Preserve the output after each critic-iterate phase and record its added time and tokens. One run can then show the initial draft, the result of self-review, and the change and cost of each later review. # Behavior details - `output.md` remains the final artifact. Scenarios that begin from an existing draft remain uncheckpointed. - Checkpoint instructions name only the next command; they do not enumerate the remaining review rounds. - Failed reviewer attempts remain in the phase record; a completed retry can still finish that phase. - Missing or out-of-order checkpoints and review records invalidate the run. - No-rules scenarios must explicitly select their prompt. ___ Differential Revision: D119374051 fbshipit-source-id: 1b628900b423f9fb9a05b2cd6d592aca10c6555bgithub.com-facebook-folly · f0fb8709 · 2026-09-10
- 0.8ETVAdd a coroutine-metadata backtest Summary: Add a backtest for writing a general-audience API contract from the middle of a real design conversation: - A partial implementation already exists but needs changes, much like a task resumed with truncated history. - Superseded discussion, an intermediate review, later decisions, and a frozen source snapshot must be reconciled. The resulting document must explain which coroutine work inherits metadata and where propagation stops, without exposing that history. ___ Reviewed By: yinglan98 Differential Revision: D119108234 fbshipit-source-id: 8ee2d0edc297074346f0da354c1124fca3838058github.com-facebook-folly · ce8fb3b0 · 2026-09-08
- 0.6ETVbenchmark_ab.py: flag benchmarks with high run-to-run spread Summary: +488/-47 nonblank lines, excluding recorded `testdata/`. Large run-to-run variation can make an otherwise reportable effect unreliable, or deserve attention even when the estimated effect is small. Measure each complete before/after series by its range relative to its median. Calibrate extreme spread against the measured benchmark set: - Require at least 20 eligible series so one benchmark cannot move the cutoff too much. - Use Tukey's outer fence for the relative cutoff. - Keep the low-priority nanosecond and percentage thresholds as practical floors. Report the result without duplicating benchmarks: - Annotate priority rows in place. - Put unclassified outliers in a separate section with TSV class `high-spread`. - Add before/after ranges in ns and as a percentage of median to TSV. Series missing any round or with median at or below 2ns remain available for effect classification but do not influence spread calibration. ___ Differential Revision: D114015996 fbshipit-source-id: dc6f5f97d85da3b32fa6b9df89717501f307f7cbgithub.com-facebook-folly · 60c56027 · 2026-08-27
- 0.5ETVCentralize Codex launch isolation Summary: Reviewer and backtest launchers need the same protection from ambient user and repository configuration. Introduce one OSS library that creates a private Git root, isolated engine state, and controlled task and working directories before launching Codex. Keep engine-specific startup behind an adapter so callers will not need to change when Claude support is added. ___ Differential Revision: D119918011 fbshipit-source-id: 363cab1f7a6c98a5aab22a01c8c94900a4eb8b3egithub.com-facebook-folly · 7bc0d0d7 · 2026-09-14
- 0.5ETVCompress stored backtest checkpoints Summary: Checkpointed samples can repeat nearly identical long outputs. That wastes repository space and makes each review's change harder to see. Add a deterministic post-processor for saved samples. It always keeps the final `output.md`, deduplicates identical phases, and keeps the earliest file for repeated non-final content. For each other earlier phase, it uses an adjacent reverse diff only when that diff is less than 60% of the full output. Before deleting any full output, the tool reconstructs every phase byte-for-byte through the reader-facing `apply_diffs` helper. It also removes unneeded runtime `artifact` fields from `checkpoints.json` and prints the phase map for the sample README. The next diff records the generated samples. ___ Differential Revision: D119576089 fbshipit-source-id: 2a1016b7518fa7a00f9437d316df766e3cbbd73bgithub.com-facebook-folly · 81353a08 · 2026-09-11
- 0.4ETVAdd a no-rules backtest baseline Summary: Existing backtests compare rule revisions but do not show what the same model produces without those rules. Add `--no-rules` so quality and cost can be compared with that baseline. The mode keeps the task and model settings while omitting injected rules and rule-only helpers. A scenario whose normal prompt mentions the rules can supply a bare prompt that removes only those instructions. Regular runs are unchanged. ___ Differential Revision: D119291040 fbshipit-source-id: f2cf3608e95c7e1ac57a87d9554ffeaa71cd504fgithub.com-facebook-folly · 43f86a10 · 2026-09-10
- 0.4ETVOmit metadata markers from signal-safe stack traces Summary: `getAsyncStackTraceSafe()` returns code addresses for symbolization. A metadata marker stores a discriminator, not code, in its return-address field, so returning it would create a bogus frame. Skip markers while retaining their visible wrapper frames. If a marker has no parent, continue into the outer stack through the async-stack root stored on the wrapper. Reviewed By: yfeldblum Differential Revision: D117316299 fbshipit-source-id: 1dfa11b657bfbb82bd453c4824d21f290333709cgithub.com-facebook-folly · eeac0003 · 2026-09-01
- 0.4ETVRecord document-coro-scoped-metadata sample (5-review checkpoint run) Summary: Record a run with up to five external-review rounds of the scenario that tests whether an agent can reconstruct a coroutine-metadata contract from conflicting evidence. The sample keeps the final output, phase costs, and evaluator findings; reverse diffs reconstruct earlier checkpoints. The run used all five reviews. ___ Differential Revision: D119576091 fbshipit-source-id: 44203a765a52c4b6d0ec20f6a0e6ad505f04cab8github.com-facebook-folly · 9c202ded · 2026-09-11