folly — Engineering Performance
59 engineers all time · Jan 2025 – Sep 2026 · built 2026-09-30 · GitHub
Performance snapshot
Today's rolling 90-day reading for folly, compared with the start of the series. Pick a window to move that comparison point.
Avg. perf / dev / mo
+218.2%
0.25 → 0.81 ETV
Active engineers
−20.0%
25.0 → 20.0
Features
−4.4pp
31.0% → 26.6%
vs. Meta
0.41x
0.36x → 0.41x · −59% below
folly vs. Meta
Per-engineer ETV for folly against Meta as a whole. Both lines are 90-day rolling averages scaled to a 30-day month, so they share one axis and can be read against each other at any point. Pick a window to zoom the chart to it.
Performance Composition
Each month's output split by type of work: Features (new value), Maintenance (sustaining systems), Tests, Docs, and Fixes (rework). The yellow line is output per engineer, so when it rises each engineer is delivering more, whatever the team size did. Unit: Engineering Throughput Value (ETV).
Engineering capacity
Effective engineers behind folly, against its pre-AI baseline. Each subject has its own: folly's is 0.40 ETV / dev / mo, its first reading in Q1 2025. Per-engineer ETV divided by that gives a capacity multiple, and that multiple applied to the engineers active in the trailing 90 days turns it into engineer-equivalents. The line is the real headcount, so the gap between line and area is what the leverage is worth. Because each baseline is its own, every subject opens at 1.0x on its first day: multiples measure improvement and are not comparable between subjects.
Knowledge concentration
How dependent is this repo on a small number of engineers? Higher top-1 share = higher key-person risk.
Yedidya Feldblum owns 26.9 % of commits.
Behind the numbers
Written summary of the work completed each month.
No monthly reports available yet.
Top engineers
Most impactful commits
Top 10 by ETV in the all-time window.
- 2.6ETVuint_divisor, lifted from thrift Fast64BitRemainderCalculator Summary: Lift `apache::thrift::frozen::detail::Fast64BitRemainderCalculator` into a new public header `folly/math/Division.h` as `folly::uint_divisor<Word>`, generalized to all unsigned integer types. Uses Lemire's constant-divisor technique (credited in doc-comments). Key changes vs. the thrift original: - Support all unsigned integer types, possibly falling back to integer division. - Support zero divisors since the optimized algorithm does not perform division. - Calculates composite division and remainder, division, or remainder. - `constexpr` construction and invocation. - Invocable-object usage with `operator()` in addition to `divrem`. Mathematical operator usage with `operator/` and `operator%` in addition to `div` and `rem`. - Thrift uses Lemire’s direct fast-remainder algorithm: multiply by a precomputed reciprocal, then perform another widened multiply to recover the remainder. uint_divisor uses the same reciprocal to compute the exact quotient and derives the remainder as dividend - quotient * divisor. For 64-bit words this reduces the hot path from roughly four multiplies to three while remaining branchless. Narrow types retain Lemire’s direct remainder path where it benchmarks faster. - Hold the original divisor in addition to the computed multiplier to support mathematical operator usage. This increases in-situ object size. - Add `uint_divisor<Word>::calc` that does not hold the original divisor; let both `uint_divisor` and `Fast64BitRemainderCalculator` delegate to it. Reviewed By: iahs Differential Revision: D115907214 fbshipit-source-id: de6d55b8dac71489f4f36b382eeb9e065380cad1Yedidya Feldblum · 653d0b3b · 2026-08-22
- 2.1ETV`benchmark_ab.py`: multi-binary benchmark A/B reports Summary: `benchmark_ab.py` is a new tool to simplify iterating on changes that affect several benchmark binaries. It handles three recurring chores: - Finds benchmark binaries from Buck target patterns. - Aggregates repeated A/B runs into one report. - Prioritizes wins and regressions with absolute and relative thresholds. Read the file docblock for more. ```text $ folly/tool/benchmark_ab.py measure --before=bc56e16776 --after=0097974dca \ //folly/result/... ... High-priority regressions: 15.2+11.2ns (+73.5%): try_to_result_error fbcode//folly/result/test:result_bench 15.2+11.2, 15.2+11.3, 15.2+11.1, 15.2+11.2, 15.4+11.2 ``` Reviewed By: janondrusek Differential Revision: D112568029 fbshipit-source-id: 544b93f75660e791285743188150a3eec38b1e1eAlexey Spiridonov · 206fc16e · 2026-07-28
- 2.0ETVcli_apply_args_files Summary: In the context of CLI args-parsing, a facility to apply args-files to an args vector. An arg beginning with `@` but not with `@@` is treated as a filename, relative to the current directory if relative, after the `@`. That arg is then replaced by parsing the file into a sequence of arguments, recursively. Reviewed By: hchokshi Differential Revision: D92167066 fbshipit-source-id: 9652db02b04202fa258f6589d468568d5f51d3d5Yedidya Feldblum · 56a01e21 · 2026-02-10
- 1.9ETVFix F14Table UBSAN insufficient-object-size error Summary: UBSAN fires `insufficient-object-size` when F14Table's single-chunk optimization allocates less than `sizeof(F14Chunk)` bytes, then calls member functions on the `Chunk*`. For example, `F14Chunk<unsigned int>` had `sizeof` = 64 bytes, but a 2-element table allocates only 24 bytes (16-byte header + 2×4-byte items). Fix by removing the `rawItems_` member array from `F14Chunk` so that `sizeof(F14Chunk)` equals 16 (just the header: `tags_[14]` + `control_` + `outboundOverflowCount_`). Items are accessed via pointer arithmetic from `this + kItemsOffset` instead of through the member array. Multi-chunk tables use an explicit stride constant (`kChunkStride`, equal to the old `sizeof`) for chunk-to-chunk navigation. This preserves the single-chunk memory optimization while making all member calls on `Chunk*` UBSAN-clean: - Minimum real allocation = 16 + 2*sizeof(Item) >= 24 > 16 = sizeof(Chunk) - Empty instance (F14EmptyTagVector) = 16 bytes = sizeof(Chunk) - Multi-chunk allocations use kChunkStride * N (same total size) Reviewed By: ilvokhin Differential Revision: D94696039 fbshipit-source-id: 06ef5480be9a2fdf3f0beb1e2c87916667e72abeYedidya Feldblum · 381d0747 · 2026-03-03
- 1.9ETVAsyncSocketTest: parametrize with IoUringBackend Summary: Validate native AsyncSocket support with IoUringBackend. Parametrize most tests (that make sense) with the backend, testing both the standard libevent backend and the io_uring backend. Reviewed By: vishwanath1306 Differential Revision: D90996725 fbshipit-source-id: a60b1702f1b6efe8d089ef1f2cb8e5b2d8207f42David Wei · 193b53f3 · 2026-01-21
- 1.5ETValgorithms for byte-sized find_first_of, find_first_not_of Summary: Includes a selection of scalar and vector algorithms all implementing `find_first_of` and `find_first_not_of`. May be useful to accelerate select parsers. Reviewed By: DenisYaroshevskiy Differential Revision: D61775260 fbshipit-source-id: 32fdd8e7c02c66a54764b0ae847aec1fa9564678Yedidya Feldblum · 8e49a654 · 2025-03-18
- 1.5ETVcstring_view, operator""_csv Summary: A class like `string_view`, but representing a pair of a C string and its precalculated size. Similar to [p3655r3](https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2025/p3655r3.html). Reviewed By: ilvokhin Differential Revision: D84381600 fbshipit-source-id: c793679ea6ebd95f2db80b8afd182384af476d36Yedidya Feldblum · 15947ac4 · 2025-10-15
- 1.3ETV`AsyncClosure.h` implement `async_closure` with async RAII Summary: Here's the short rationale for `async_closure()`. The header has a user-facing tl;dr, and `docs/` provide far more context. - It provides a robust solution to the "async cleanup" / "async RAII" problem. `co_cleanup_capture` allows the closure to await multiple cleanup steps on exit (like `co_scope_exit`), but the cleanup is actually enforced to be memory-safe (unlike `co_scope_exit`). - Async scopes owned by a closure can support "natural" cancellation, without incurring additional cost for it the way you would with `CancellableAsyncScope`. - It supports compile-time-checked "safe" reference-passing from ancestors to descendants. From a user's perspective, they just pass `capture`s into child closures, and it works. Or when it fails to compile, this means there may be a memory-safety bug. --- The integration of `MemberTask` with `async_closure` is a bit magic, in two ways: - A special branch lets us use the pre-existing `FOLLY_INVOKE_MEMBER` "overload set" convention with `capture` and `AsyncObjectPtr` wrappers that require dereferencing. We could avoid this by introducing a separate macro like `FOLLY_INVOKE_MEMBER_INDIRECT` or `FOLLY_INVOKE_MEMBER_TASK`, but the current UX feels comfortable, and I don't see a big risk. - (true as of the prior diff) As with `ClosureTask`, `async_closure` implicitly re-wraps `MemberTask` making it movable and upgrading its `SafeTask` safety level. PS We should eventually have a linter that enforces that `MemberTask` is only used for non-static member functions, but this seems low-risk for now. *The user-facing explanation of this is in `APIBestPractices.md` (D65408249).* Reviewed By: ispeters Differential Revision: D63299858 fbshipit-source-id: a3519b3873c7b0e9859db9f848a4cde916a950bfAlexey Spiridonov · 9b095871 · 2025-03-18
- 1.3ETVAdd --bm_mode=adaptive; change default --bm_min_usec=1000 Summary: # `--bm_mode=adaptive` I got frustrated with manually retrying benchmark runs to avoid system performance oscillations, and with manually aggregating these runs to get a reasonably precise number. Existing modes (regular or `--bm_estimate_time`) weren't working very well for me -- too noisy, too slow, or both. The new `--bm_mode=adaptive` tries to automate a "statistically sound measurement" for a system that is not short-term stationary. Think of it is "automatic noise cancellation" for the modern reality of working on multi-tenant VMs. If you're used to doing best-of-5 -- this is the same idea, but better. Read `docs/BenchmarkAdaptive.md` for a proper description, and the caveats. --- # `--bm_min_usec` default change: 100μs -> 1ms **This makes default runs 3-8x slower, but they'll stop being wrong.** Here are 5 back-to-back runs with `--bm_min_usec=100 -bm_max_secs=30` on a quiet system. I raise "max_secs" because the default under-samples with larger `min_usec`, giving even more noise. ``` LegacyCaseInsensitiveCheck 9.23us 108.33K CurrentCaseInsensitiveCheck 1.41us 711.35K LegacyCaseInsensitiveCheck 9.23us 108.33K CurrentCaseInsensitiveCheck 1.54us 650.59K LegacyCaseInsensitiveCheck 9.23us 108.33K CurrentCaseInsensitiveCheck 1.40us 712.26K LegacyCaseInsensitiveCheck 8.41us 118.87K CurrentCaseInsensitiveCheck 1.54us 650.26K LegacyCaseInsensitiveCheck 9.23us 108.33K CurrentCaseInsensitiveCheck 1.41us 710.83K ``` Wildly inconsistent! Each took 1.6 seconds, so 8 seconds total, and we're none the wiser about the true distribution. And it's much slower than the typical 1-2 sec for `adaptive` on the same 2 benchmarks, which *does* produce consistent results, fast. ``` $ time buck run @//mode/opt fbcode//folly/test:ascii_case_insensitive_benchmark -- --bm_mode=adaptive ... LegacyCaseInsensitiveCheck 8.54us 117.09K CurrentCaseInsensitiveCheck 1.40us 714.48K real 0m0.870s ``` With 100μs slices above, the regular runs shows cache interference from the benchmark harness code (I'm 80% sure that's the cause, from my experiments). Going to `--bm_min_usec=1000` hides that, and makes regular runs consistent too. With 1ms slices, regular-mode runs look like this, but they now take **~12 seconds** each (only 4 with default `max_secs`, but that's noisier). But hey, at least you get usable data! ``` LegacyCaseInsensitiveCheck 8.50us 117.60K CurrentCaseInsensitiveCheck 1.39us 718.93K LegacyCaseInsensitiveCheck 8.50us 117.61K CurrentCaseInsensitiveCheck 1.42us 704.37K ``` The timings differ slightly, since `adaptive` measures p33 by default, while regular always measures p0. If 1ms is deemed "too slow by default", 500μs is borderline -- you still see unacceptable interference, but not as much. I could imagine landing with that for the regular/legacy mode, and giving `adaptive` a better default. It'd be cool to redesign of the benchmark setup to skip timing the first few iterations of each slice to mitigate this more generally, but I'm not sure the cost-benefit is favorable. Reviewed By: yfeldblum Differential Revision: D92348138 fbshipit-source-id: 92dda73d6753a07aeac9b51cb4efb155cd1f6143Alexey Spiridonov · 2108e961 · 2026-02-25
- 1.2ETVIoUringProvidedBufferRing: dynamically grow buffer areas on demand Summary: Replace the single fixed-size buffer allocation with a pool of dynamically growable buffer areas. Previously the ring mapped one contiguous block holding exactly ringBufferCount_ buffers; once all of them were handed out to the user and still held, further reads returned -ENOBUFS as backpressure. This change decouples the kernel ring (still ringBufferCount_ entries) from the backing storage, which is now a set of BufferArea blocks, each holding a full ring's worth of buffers with its own BufferState[] and an `outstanding` counter (posted-in-ring plus held-by-user). The ring is refilled from these areas: - mapMemory() is split into mapRing() (the io_uring_buf ring) and allocateMemoryArea()/createArea()/appendAreas() for the buffer areas. - The pool starts with kInitialAreaCount (2) areas and grows by one each time there is no area available. - An area can only be reused once fully drained (outstanding == 0), preventing corruption of buffers still held by the user. - getIoBuf() now tracks the consume-side area (bufferConsumedArea_) so it follows refill's posted area order across area boundaries, including bundles that straddle areas, rather than assuming a naive (area + 1) % areaCount step. - getUtilPct() is reworked to count only user-held buffers across all areas. - returnBuffer() is replaced by ringRefill()/bufferConsumed() and the area `outstanding` accounting. Adds //folly:expected as a dependency (growAreas()/areaGetNextFree() return folly::Expected). Reviewed By: keithbusch Differential Revision: D116323996 fbshipit-source-id: ed6c1417797147935520a3bdd8e001a0be27795cDavid Wei · 7aa0b8b8 · 2026-08-21