Commit c71517c
faiss HNSW: add opt-in deterministic lock-free graph build (faiss::hnsw_deterministic_build) (#5486)
Summary:
TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind `faiss::hnsw_deterministic_build`. We plan to test internally, then make this the default flow and remove the existing flow.
--
HOW TO ENABLE
--
```
C++: faiss::hnsw_deterministic_build = true;
Python: faiss.cvar.hnsw_deterministic_build = True
```
Nothing else is needed, and nothing changes until you do it: the flag defaults
to false, so this diff is a no-op on its own.
An embedder driving this from its own configuration system should resolve the
value once at startup and assign the global, rather than per index: changing it
between two `add()` calls would leave one index built by each algorithm.
`python/__init__.pyi` now declares `faiss.cvar`. It was missing, so any Python
consumer assigning a faiss global failed Pyre with "No attribute `cvar` in
module `faiss`" -- fixed here rather than with a `pyre-ignore` in each consumer.
WHEN IT IS WORTH ENABLING
--
Only for builds that run multi-threaded. At one thread the existing lock-based
build is *already* deterministic, so the flag buys nothing beyond ~6% build
time. Verified: lock-based at OMP=1 is byte-identical across runs, while at
OMP=8 46% of neighbor slots differ. A consumer that pins
`omp_set_num_threads(1)` should not bother.
similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase
differences from ParlayANN:
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`
---
AI (with a bunch of edits) explanation in more detail:
--
What changed
- `IndexHNSW::add` and `IndexBinaryHNSW::add` pick the build at runtime. Both paths are compiled in and **the default is unchanged** (lock-based); nothing switches until the flag or the hook says so.
- Turning it on:
```
C++: faiss::hnsw_deterministic_build = true;
Python: faiss.cvar.hnsw_deterministic_build = True
```
- A single plain `bool` is the whole API: libfaiss takes no dependency on any configuration system, and an embedder that wants one resolves it itself and assigns the global:
```
faiss.cvar.hnsw_deterministic_build = <your config lookup>
```
- Resolve it once at startup, not per index. `add()` reads the flag each call, so changing it between two `add()` calls on one index would build part of the graph with each algorithm -- reproducible in neither.
- **This flag is transitional.** Once the deterministic build is validated in production, the follow-up makes it the only path and deletes the lock-based build and the flag.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- `IndexBinaryHNSW::add` is gated by the same flag and shares the same deterministic implementation: `hnsw_add_vertices_deterministic` takes `make_distance_computer` / `set_query` callbacks, so the binary index just supplies its own Hamming distance computer. It keeps its own lock-based `hnsw_add_vertices` for the default path.
Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.
Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).
Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.
## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
lock-based per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
lock-based: min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).
## Performance: search time
Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
purely graph structure + measurement noise.
Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)
HNSW16
```
┌──────────┬─────────────────┬─────────┬──────────┬───────┐
│ efSearch │ recall det/lock │ QPS det │ QPS lock │ Δ │
├──────────┼─────────────────┼─────────┼──────────┼───────┤
│ 64 │ 0.828/0.820 │ 170,329 │ 177,995 │ −4.3% │
├──────────┼─────────────────┼─────────┼──────────┼───────┤
│ 128 │ 0.866/0.862 │ 112,727 │ 110,727 │ +1.8% │
├──────────┼─────────────────┼─────────┼──────────┼───────┤
│ 256 │ 0.886/0.888 │ 59,815 │ 56,784 │ +5.3% │
└──────────┴─────────────────┴─────────┴──────────┴───────┘
```
HNSW32
```
┌──────────┬─────────────────┬─────────┬──────────┬───────┐
│ efSearch │ recall det/lock │ QPS det │ QPS lock │ Δ │
├──────────┼─────────────────┼─────────┼──────────┼───────┤
│ 64 │ 0.935/0.930 │ 110,186 │ 108,411 │ +1.6% │
├──────────┼─────────────────┼─────────┼──────────┼───────┤
│ 128 │ 0.958/0.953 │ 68,019 │ 66,308 │ +2.6% │
├──────────┼─────────────────┼─────────┼──────────┼───────┤
│ 256 │ 0.965/0.960 │ 37,624 │ 35,828 │ +5.0% │
└──────────┴─────────────────┴─────────┴──────────┴───────┘
```
HNSW32,SQ8
```
┌──────────┬─────────────────┬─────────┬──────────┬────────┐
│ efSearch │ recall det/lock │ QPS det │ QPS lock │ Δ │
├──────────┼─────────────────┼─────────┼──────────┼────────┤
│ 64 │ 0.926/0.934 │ 220,713 │ 198,325 │ +11.3% │
├──────────┼─────────────────┼─────────┼──────────┼────────┤
│ 128 │ 0.948/0.953 │ 117,504 │ 129,173 │ −9.0% │
├──────────┼─────────────────┼─────────┼──────────┼────────┤
│ 256 │ 0.961/0.963 │ 58,582 │ 65,551 │ −10.6% │
└──────────┴─────────────────┴─────────┴──────────┴────────┘
```
(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)
Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
all prior results.
## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
build time: lock-based 188.36s vs deterministic 176.40s (0.94x -> deterministic ~6% FASTER)
peak RSS: 1.6 GB (both, identical)
recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.
## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `hnsw_deterministic_build` is a namespace-scope global, not a field on `HNSW` or `IndexHNSW`, so it is not part of any serialized struct. Like `retain_locks`, it is a pure runtime build flag.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).
- Measured on a 1M index: written with the flag off, flag flipped on, read back
-> `neighbors`, `offsets`, `entry_point`, `max_level` byte-identical and
search ids+distances identical. The flag is read only inside `add()`.
- Mixed-mode append (lock-built graph extended by a deterministic `add()`) was
measured against brute-force ground truth at 220k and is indistinguishable
from either pure build -- recall@10 within +-0.02 of both at ef 32/64/128. No
existing test covers this, since every test builds one way in one process.
## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
all_neighbors build: 367.4s optimize: 231.7s copyTo: 18.4s serialize: 28.9s (66 GB)
INDEX build -> serialize total: 661.7s (11.0 min) [D106837134 baseline: 721s]
recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).
## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.
Differential Revision: D1120258771 parent da3191e commit c71517c
7 files changed
Lines changed: 984 additions & 33 deletions
File tree
- faiss
- impl
- python
- tests
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
270 | 270 | | |
271 | 271 | | |
272 | 272 | | |
273 | | - | |
274 | | - | |
275 | | - | |
276 | | - | |
277 | | - | |
278 | | - | |
279 | | - | |
| 273 | + | |
| 274 | + | |
| 275 | + | |
| 276 | + | |
| 277 | + | |
| 278 | + | |
| 279 | + | |
| 280 | + | |
| 281 | + | |
| 282 | + | |
| 283 | + | |
| 284 | + | |
| 285 | + | |
| 286 | + | |
| 287 | + | |
| 288 | + | |
| 289 | + | |
| 290 | + | |
| 291 | + | |
280 | 292 | | |
281 | 293 | | |
282 | 294 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
41 | 41 | | |
42 | 42 | | |
43 | 43 | | |
| 44 | + | |
| 45 | + | |
44 | 46 | | |
45 | 47 | | |
46 | 48 | | |
| |||
214 | 216 | | |
215 | 217 | | |
216 | 218 | | |
| 219 | + | |
| 220 | + | |
| 221 | + | |
| 222 | + | |
| 223 | + | |
| 224 | + | |
| 225 | + | |
| 226 | + | |
| 227 | + | |
| 228 | + | |
| 229 | + | |
| 230 | + | |
| 231 | + | |
| 232 | + | |
| 233 | + | |
| 234 | + | |
| 235 | + | |
| 236 | + | |
| 237 | + | |
| 238 | + | |
| 239 | + | |
| 240 | + | |
| 241 | + | |
| 242 | + | |
| 243 | + | |
| 244 | + | |
| 245 | + | |
| 246 | + | |
| 247 | + | |
| 248 | + | |
| 249 | + | |
| 250 | + | |
| 251 | + | |
| 252 | + | |
| 253 | + | |
| 254 | + | |
| 255 | + | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
| 259 | + | |
| 260 | + | |
| 261 | + | |
| 262 | + | |
| 263 | + | |
| 264 | + | |
| 265 | + | |
| 266 | + | |
| 267 | + | |
| 268 | + | |
| 269 | + | |
| 270 | + | |
| 271 | + | |
| 272 | + | |
| 273 | + | |
| 274 | + | |
| 275 | + | |
| 276 | + | |
| 277 | + | |
| 278 | + | |
| 279 | + | |
| 280 | + | |
| 281 | + | |
| 282 | + | |
| 283 | + | |
| 284 | + | |
| 285 | + | |
| 286 | + | |
| 287 | + | |
| 288 | + | |
| 289 | + | |
| 290 | + | |
| 291 | + | |
| 292 | + | |
| 293 | + | |
| 294 | + | |
| 295 | + | |
| 296 | + | |
| 297 | + | |
| 298 | + | |
| 299 | + | |
| 300 | + | |
| 301 | + | |
| 302 | + | |
| 303 | + | |
| 304 | + | |
| 305 | + | |
| 306 | + | |
| 307 | + | |
| 308 | + | |
| 309 | + | |
| 310 | + | |
| 311 | + | |
| 312 | + | |
| 313 | + | |
| 314 | + | |
| 315 | + | |
| 316 | + | |
| 317 | + | |
| 318 | + | |
| 319 | + | |
| 320 | + | |
| 321 | + | |
| 322 | + | |
| 323 | + | |
| 324 | + | |
| 325 | + | |
| 326 | + | |
| 327 | + | |
| 328 | + | |
| 329 | + | |
| 330 | + | |
| 331 | + | |
| 332 | + | |
| 333 | + | |
| 334 | + | |
| 335 | + | |
| 336 | + | |
| 337 | + | |
| 338 | + | |
| 339 | + | |
| 340 | + | |
| 341 | + | |
| 342 | + | |
| 343 | + | |
| 344 | + | |
| 345 | + | |
| 346 | + | |
| 347 | + | |
| 348 | + | |
| 349 | + | |
| 350 | + | |
| 351 | + | |
| 352 | + | |
| 353 | + | |
| 354 | + | |
| 355 | + | |
| 356 | + | |
| 357 | + | |
| 358 | + | |
| 359 | + | |
| 360 | + | |
| 361 | + | |
| 362 | + | |
| 363 | + | |
| 364 | + | |
| 365 | + | |
| 366 | + | |
| 367 | + | |
| 368 | + | |
| 369 | + | |
| 370 | + | |
| 371 | + | |
| 372 | + | |
| 373 | + | |
| 374 | + | |
| 375 | + | |
| 376 | + | |
| 377 | + | |
| 378 | + | |
| 379 | + | |
| 380 | + | |
| 381 | + | |
| 382 | + | |
| 383 | + | |
| 384 | + | |
| 385 | + | |
| 386 | + | |
| 387 | + | |
| 388 | + | |
| 389 | + | |
| 390 | + | |
| 391 | + | |
| 392 | + | |
| 393 | + | |
| 394 | + | |
| 395 | + | |
| 396 | + | |
| 397 | + | |
| 398 | + | |
| 399 | + | |
| 400 | + | |
| 401 | + | |
| 402 | + | |
| 403 | + | |
| 404 | + | |
| 405 | + | |
| 406 | + | |
| 407 | + | |
| 408 | + | |
| 409 | + | |
| 410 | + | |
| 411 | + | |
| 412 | + | |
| 413 | + | |
| 414 | + | |
| 415 | + | |
| 416 | + | |
| 417 | + | |
| 418 | + | |
| 419 | + | |
| 420 | + | |
| 421 | + | |
| 422 | + | |
| 423 | + | |
| 424 | + | |
| 425 | + | |
| 426 | + | |
| 427 | + | |
| 428 | + | |
| 429 | + | |
| 430 | + | |
| 431 | + | |
| 432 | + | |
| 433 | + | |
| 434 | + | |
| 435 | + | |
| 436 | + | |
| 437 | + | |
| 438 | + | |
| 439 | + | |
| 440 | + | |
| 441 | + | |
| 442 | + | |
| 443 | + | |
| 444 | + | |
| 445 | + | |
| 446 | + | |
| 447 | + | |
| 448 | + | |
| 449 | + | |
| 450 | + | |
| 451 | + | |
| 452 | + | |
| 453 | + | |
| 454 | + | |
| 455 | + | |
| 456 | + | |
| 457 | + | |
| 458 | + | |
| 459 | + | |
| 460 | + | |
| 461 | + | |
| 462 | + | |
| 463 | + | |
| 464 | + | |
| 465 | + | |
| 466 | + | |
| 467 | + | |
| 468 | + | |
| 469 | + | |
| 470 | + | |
| 471 | + | |
| 472 | + | |
| 473 | + | |
| 474 | + | |
| 475 | + | |
| 476 | + | |
| 477 | + | |
| 478 | + | |
| 479 | + | |
| 480 | + | |
| 481 | + | |
| 482 | + | |
| 483 | + | |
| 484 | + | |
| 485 | + | |
| 486 | + | |
| 487 | + | |
| 488 | + | |
| 489 | + | |
| 490 | + | |
| 491 | + | |
| 492 | + | |
| 493 | + | |
| 494 | + | |
| 495 | + | |
| 496 | + | |
| 497 | + | |
| 498 | + | |
| 499 | + | |
| 500 | + | |
| 501 | + | |
| 502 | + | |
| 503 | + | |
| 504 | + | |
| 505 | + | |
| 506 | + | |
| 507 | + | |
| 508 | + | |
| 509 | + | |
| 510 | + | |
| 511 | + | |
| 512 | + | |
217 | 513 | | |
218 | 514 | | |
219 | 515 | | |
| |||
384 | 680 | | |
385 | 681 | | |
386 | 682 | | |
387 | | - | |
388 | | - | |
389 | | - | |
390 | | - | |
391 | | - | |
392 | | - | |
393 | | - | |
| 683 | + | |
| 684 | + | |
| 685 | + | |
| 686 | + | |
| 687 | + | |
| 688 | + | |
| 689 | + | |
| 690 | + | |
| 691 | + | |
| 692 | + | |
| 693 | + | |
| 694 | + | |
| 695 | + | |
| 696 | + | |
| 697 | + | |
| 698 | + | |
| 699 | + | |
| 700 | + | |
| 701 | + | |
394 | 702 | | |
395 | 703 | | |
396 | 704 | | |
| |||
0 commit comments