Skip to content

Commit 44f9f3e

Browse files
Michael Norrisfacebook-github-bot
authored andcommitted
faiss HNSW: add opt-in deterministic lock-free graph build (faiss::hnsw_deterministic_build) (#5486)
Summary: TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind `faiss::hnsw_deterministic_build`. We plan to test internally, then make this the default flow and remove the existing flow. -- HOW TO ENABLE -- ``` C++: faiss::hnsw_deterministic_build = true; Python: faiss.cvar.hnsw_deterministic_build = True ``` Nothing else is needed, and nothing changes until you do it: the flag defaults to false, so this diff is a no-op on its own. Meta-internally the default can instead come from the JustKnob `faiss/hnsw:enable_deterministic_build`, so a customer ramps and rolls back by config with no rebuild: - **C++ consumers** add `//faiss/fb/runtime_config:runtime_config` to their deps -- the hook installs itself, no code change. Targeting is by JustKnobs condition rule (e.g. `TW_JOB_HANDLE REGEX ...`), since the hook passes no switchval. - **Python consumers** cannot use that target (a `cpp_library` does not link into the pyfaiss extension, and the hook's process-phase guard is never satisfied under CPython), so they read the knob themselves, with a per-usecase switchval: ``` faiss.cvar.hnsw_deterministic_build = justknobs.check( "faiss/hnsw:enable_deterministic_build", switchval="<usecase>") ``` Read it once per job, not per index: a mid-job flip would otherwise leave two indexes built by different algorithms. `python/__init__.pyi` now declares `faiss.cvar`. It was missing, so any Python consumer assigning a faiss global failed Pyre with "No attribute `cvar` in module `faiss`" -- fixed here rather than with a `pyre-ignore` in each consumer. Consumer diffs: D114448906 (Laser), D114448902 (VeST), D114448905 (OVIS), D114448904 (Unicorn). Ads Vector DB needs nothing -- it never constructs HNSW. WHEN IT IS WORTH ENABLING -- Only for builds that run multi-threaded. At one thread the existing lock-based build is *already* deterministic, so the flag buys nothing beyond ~6% build time. Verified: lock-based at OMP=1 is byte-identical across runs, while at OMP=8 46% of neighbor slots differ. A consumer that pins `omp_set_num_threads(1)` should not bother. similarities to parlayANN: - add vertices in doubling batches against frozen snapshot - defer adding reciprocal edges immediately, add them later after parallel phase differences from ParlayANN: - re-uses Faiss HNSW pruning in `shrink_neighbor_list` --- AI (with a bunch of edits) explanation in more detail: -- What changed - `IndexHNSW::add` and `IndexBinaryHNSW::add` pick the build at runtime. Both paths are compiled in and **the default is unchanged** (lock-based); nothing switches until the flag or the hook says so. - Turning it on: ``` C++: faiss::hnsw_deterministic_build = true; Python: faiss.cvar.hnsw_deterministic_build = True ``` - Internally the default can also come from the JustKnob `faiss/hnsw:enable_deterministic_build`, via the new `faiss/fb/runtime_config` target (stripped from the OSS export). It installs `faiss::hnsw_deterministic_build_hook`, and `add()` uses the deterministic build when the flag is set OR the hook returns true -- so setting the flag explicitly always wins, and clearing the hook forces the legacy build. - **That target is opt-in per consumer.** `//faiss:faiss` deliberately does NOT depend on it: core faiss has no Meta-internal dependencies today, and adding `//justknobs` there would push one onto every faiss consumer. A C++ service opts in by adding `//faiss/fb/runtime_config:runtime_config` to its own deps -- no code change, the hook installs itself. Python consumers cannot link a `cpp_library` into the pyfaiss extension, so they assign the global directly: ``` import faiss, pyjk faiss.cvar.hnsw_deterministic_build = pyjk.check("faiss/hnsw:enable_deterministic_build") ``` - Keeping the hook out of `//faiss:faiss` also keeps faiss's own tests independent of a production config value. - Caveat of the hook being live: it is re-evaluated on **every** `add()`, so a knob flip reaches running processes without a restart, but two `add()` calls in one process can disagree if the knob flips between them -- an index built in several batches during a ramp could be built partly one way and partly the other. - The hook exists because the knob cannot simply be written to the global at static-init time: a JustKnobs read before `initFacebook()` aborts the process (folly "unregistered singleton"), which is not catchable. - **This flag is transitional.** Once the deterministic build is validated internally, the follow-up makes it the only path and deletes the lock-based build, the flag, the hook, and `faiss/fb/runtime_config` entirely. - The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path. - `IndexBinaryHNSW::add` is gated by the same flag and shares the same deterministic implementation: `hnsw_add_vertices_deterministic` takes `make_distance_computer` / `set_query` callbacks, so the binary index just supplies its own Hamming distance computer. It keeps its own lock-based `hnsw_add_vertices` for the default path. Background -- HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race. Algorithm (adapted from ParlayANN to Faiss's level-batched structure): - Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index). - Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free. - Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning). Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch. ## Performance: build time (40M, d=128, M=32, efC=64, 166 threads) 10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions): deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29 lock-based per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92 deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53 lock-based: min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12 det/lock: mean=0.860 (deterministic ~14% faster), median=0.867 The deterministic build is ~14% faster than the lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds). ## Performance: search time Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores), search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is purely graph structure + measurement noise. Search QPS: deterministic vs lock-based (40M synthetic, repeat=100) HNSW16 ``` ┌──────────┬─────────────────┬─────────┬──────────┬───────┐ │ efSearch │ recall det/lock │ QPS det │ QPS lock │ Δ │ ├──────────┼─────────────────┼─────────┼──────────┼───────┤ │ 64 │ 0.828/0.820 │ 170,329 │ 177,995 │ −4.3% │ ├──────────┼─────────────────┼─────────┼──────────┼───────┤ │ 128 │ 0.866/0.862 │ 112,727 │ 110,727 │ +1.8% │ ├──────────┼─────────────────┼─────────┼──────────┼───────┤ │ 256 │ 0.886/0.888 │ 59,815 │ 56,784 │ +5.3% │ └──────────┴─────────────────┴─────────┴──────────┴───────┘ ``` HNSW32 ``` ┌──────────┬─────────────────┬─────────┬──────────┬───────┐ │ efSearch │ recall det/lock │ QPS det │ QPS lock │ Δ │ ├──────────┼─────────────────┼─────────┼──────────┼───────┤ │ 64 │ 0.935/0.930 │ 110,186 │ 108,411 │ +1.6% │ ├──────────┼─────────────────┼─────────┼──────────┼───────┤ │ 128 │ 0.958/0.953 │ 68,019 │ 66,308 │ +2.6% │ ├──────────┼─────────────────┼─────────┼──────────┼───────┤ │ 256 │ 0.965/0.960 │ 37,624 │ 35,828 │ +5.0% │ └──────────┴─────────────────┴─────────┴──────────┴───────┘ ``` HNSW32,SQ8 ``` ┌──────────┬─────────────────┬─────────┬──────────┬────────┐ │ efSearch │ recall det/lock │ QPS det │ QPS lock │ Δ │ ├──────────┼─────────────────┼─────────┼──────────┼────────┤ │ 64 │ 0.926/0.934 │ 220,713 │ 198,325 │ +11.3% │ ├──────────┼─────────────────┼─────────┼──────────┼────────┤ │ 128 │ 0.948/0.953 │ 117,504 │ 129,173 │ −9.0% │ ├──────────┼─────────────────┼─────────┼──────────┼────────┤ │ 256 │ 0.961/0.963 │ 58,582 │ 65,551 │ −10.6% │ └──────────┴─────────────────┴─────────┴──────────┴────────┘ ``` (Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.) Verdict: no search-QPS regression - Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with equal-or-better recall. - HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963), i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated) comparison would remove the operating-point confound. - Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with all prior results. ## Single-threaded (OMP_NUM_THREADS=1) Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded: build time: lock-based 188.36s vs deterministic 176.40s (0.94x -> deterministic ~6% FASTER) peak RSS: 1.6 GB (both, identical) recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798 No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup. ## Serialization compatibility No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`: - `hnsw_deterministic_build` is a namespace-scope global, not a field on `HNSW` or `IndexHNSW`, so it is not part of any serialized struct. Like `retain_locks`, it is a pure runtime build flag. - `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before. - `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized). - The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss. - Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass). - Measured on a 1M index: written with the flag off, flag flipped on, read back -> `neighbors`, `offsets`, `entry_point`, `max_level` byte-identical and search ids+distances identical. The flag is read only inside `add()`. - Mixed-mode append (lock-built graph extended by a deterministic `add()`) was measured against brute-force ground truth at 220k and is indistinguishable from either pure build -- recall@10 within +-0.02 of both at ef 32/64/128. No existing test covers this, since every test builds one way in one process. ## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression): all_neighbors build: 367.4s optimize: 231.7s copyTo: 18.4s serialize: 28.9s (66 GB) INDEX build -> serialize total: 661.7s (11.0 min) [D106837134 baseline: 721s] recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429 - This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add(). - The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python). ## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note) One deliberate difference from the lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations. Differential Revision: D112025877
1 parent da3191e commit 44f9f3e

8 files changed

Lines changed: 1052 additions & 74 deletions

File tree

benchs/bench_hnsw.py

Lines changed: 0 additions & 41 deletions
Original file line numberDiff line numberDiff line change
@@ -191,44 +191,3 @@ def evaluate(index):
191191
print("search_L", search_L, end=" ")
192192
index.nsg.search_L = search_L
193193
evaluate(index)
194-
195-
196-
if "hnsw_locks" in todo:
197-
198-
ntotal, _ = xb.shape
199-
batch_size = ntotal // 100
200-
print(
201-
f"Testing HNSW Flat: add with {batch_size=}, "
202-
"with and without retaining locks"
203-
)
204-
205-
# Unbatched
206-
t0 = time.time()
207-
index = faiss.IndexHNSWFlat(d, 32)
208-
index.add(xb)
209-
t1 = time.time()
210-
print(
211-
f"\t single bulk add(): {index.ntotal} added in {t1 - t0:6.3f}s"
212-
f" = {index.ntotal / (t1 - t0):.0f}/s"
213-
)
214-
215-
for retain_locks in [False, True]:
216-
index = faiss.IndexHNSWFlat(d, 32)
217-
index.retain_locks = retain_locks
218-
219-
t0 = time.time()
220-
t1 = None
221-
t2 = None
222-
for i in range(0, len(xb), batch_size):
223-
t1 = time.time()
224-
index.add(xb[i : i + batch_size])
225-
t2 = time.time()
226-
if i > 2 and t2 - t0 > 2:
227-
break
228-
229-
assert t1 and t2
230-
dt = t2 - t0
231-
print(
232-
f"\t {retain_locks=:1}: {index.ntotal} added in {t2 - t0:6.3f}s"
233-
f" = {index.ntotal / (t2 - t0):.0f}/s"
234-
)

faiss/IndexBinaryHNSW.cpp

Lines changed: 19 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -270,13 +270,25 @@ void IndexBinaryHNSW::add(idx_t n, const uint8_t* x) {
270270
storage->add(n, x);
271271
ntotal = storage->ntotal;
272272

273-
hnsw_add_vertices(
274-
*this,
275-
n0,
276-
n,
277-
x,
278-
verbose,
279-
hnsw.levels.size() == static_cast<size_t>(ntotal));
273+
bool preset_levels = hnsw.levels.size() == static_cast<size_t>(ntotal);
274+
275+
if (hnsw_use_deterministic_build()) {
276+
hnsw_add_vertices_deterministic(
277+
hnsw,
278+
n0,
279+
n,
280+
d,
281+
init_level0,
282+
keep_max_size_level0,
283+
preset_levels,
284+
verbose,
285+
[this] { return get_distance_computer(); },
286+
[this, x, n0](DistanceComputer& dc, HNSW::storage_idx_t pt_id) {
287+
dc.set_query((const float*)(x + (pt_id - n0) * code_size));
288+
});
289+
} else {
290+
hnsw_add_vertices(*this, n0, n, x, verbose, preset_levels);
291+
}
280292
}
281293

282294
void IndexBinaryHNSW::reset() {

0 commit comments

Comments
 (0)