Skip to content

faiss HNSW: add opt-in deterministic lock-free graph build (faiss::hnsw_deterministic_build) (#5486) - #5486

Closed
mnorris11 wants to merge 1 commit into
facebookresearch:mainfrom
mnorris11:export-D112025877
Closed

faiss HNSW: add opt-in deterministic lock-free graph build (faiss::hnsw_deterministic_build) (#5486)#5486
mnorris11 wants to merge 1 commit into
facebookresearch:mainfrom
mnorris11:export-D112025877

Conversation

@mnorris11

@mnorris11 mnorris11 commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Summary:

TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind faiss::hnsw_deterministic_build. We plan to test internally, then make this the default flow and remove the existing flow.

https://arxiv.org/abs/2305.04359

HOW TO ENABLE

C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True

Nothing else is needed, and nothing changes until you do it: the flag defaults
to false, so this diff is a no-op on its own.

An embedder driving this from its own configuration system should resolve the
value once at startup and assign the global, rather than per index: changing it
between two add() calls would leave one index built by each algorithm.

python/__init__.pyi now declares faiss.cvar. It was missing, so any Python
consumer assigning a faiss global failed Pyre with "No attribute cvar in
module faiss" -- fixed here rather than with a pyre-ignore in each consumer.

WHEN IT IS WORTH ENABLING

Only for builds that run multi-threaded. At one thread the existing lock-based
build is already deterministic, so the flag buys nothing beyond ~6% build
time. Verified: lock-based at OMP=1 is byte-identical across runs, while at
OMP=8 46% of neighbor slots differ. A consumer that pins
omp_set_num_threads(1) should not bother.

similarities to parlayANN:

  • add vertices in doubling batches against frozen snapshot
  • defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:

  • re-uses Faiss HNSW pruning in shrink_neighbor_list

AI (with a bunch of edits) explanation in more detail:

What changed

  • IndexHNSW::add and IndexBinaryHNSW::add pick the build at runtime. Both paths are compiled in and the default is unchanged (lock-based); nothing switches until the flag or the hook says so.
  • Turning it on:
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
  • A single plain bool is the whole API: libfaiss takes no dependency on any configuration system, and an embedder that wants one resolves it itself and assigns the global:
faiss.cvar.hnsw_deterministic_build = <your config lookup>
  • Resolve it once at startup, not per index. add() reads the flag each call, so changing it between two add() calls on one index would build part of the graph with each algorithm -- reproducible in neither.
  • This flag is transitional. Once the deterministic build is validated in production, the follow-up makes it the only path and deletes the lock-based build and the flag.
  • The deterministic path now supports the CAGRA level-0 import configuration: init_level0=false skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and keep_max_size_level0 fills the base layer to 2*M. So IndexHNSWCagra (CPU) and GpuIndexCagra::copyTo(IndexHNSWCagra*) build through the deterministic path.
  • IndexBinaryHNSW::add is gated by the same flag and shares the same deterministic implementation: hnsw_add_vertices_deterministic takes make_distance_computer / set_query callbacks, so the binary index just supplies its own Hamming distance computer. It keeps its own lock-based hnsw_add_vertices for the default path.

Background

HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of IndexHNSW::add with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in add_links_starting_from_impl; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap #pragma omp critical race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):

  • Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
  • Phase A (HNSW::compute_forward_links_deterministic, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
  • Phase B (HNSW::merge_reverse_links_deterministic, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of dest (a small constant bucket count, independent of ntotal and thread count, so grouping stays O(edges) in memory), each bucket sorted by (level, dest) and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every dest maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses schedule(static) — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

Performance: build time (40M, d=128, M=32, efC=64, 166 threads)

10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
lock-based per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
lock-based: min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
purely graph structure + measurement noise.

Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression

  • Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
    equal-or-better recall.
  • HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
    i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
    identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
    comparison would remove the operating-point confound.
  • Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
    all prior results.

Single-threaded (OMP_NUM_THREADS=1)

Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
build time: lock-based 188.36s vs deterministic 176.40s (0.94x -> deterministic ~6% FASTER)
peak RSS: 1.6 GB (both, identical)
recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

Serialization compatibility

No on-disk format change, verified in index_read.cpp / index_write.cpp:

  • hnsw_deterministic_build is a namespace-scope global, not a field on HNSW or IndexHNSW, so it is not part of any serialized struct. Like retain_locks, it is a pure runtime build flag.
  • write_HNSW / read_HNSW and the IndexHNSW field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
  • keep_max_size_level0 is still serialized only for the CAGRA subtype (IHc2/IHNc); init_level0 is build-only (not serialized).
  • The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
  • Verified by the io_and_retest serialize -> deserialize -> re-search round-trips in test_graph_based.py / test_hnsw.cpp (all pass).
  • Measured on a 1M index: written with the flag off, flag flipped on, read back
    -> neighbors, offsets, entry_point, max_level byte-identical and
    search ids+distances identical. The flag is read only inside add().
  • Mixed-mode append (lock-built graph extended by a deterministic add()) was
    measured against brute-force ground truth at 220k and is indistinguishable
    from either pure build -- recall@10 within +-0.02 of both at ef 32/64/128. No
    existing test covers this, since every test builds one way in one process.

CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification

Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
all_neighbors build: 367.4s optimize: 231.7s copyTo: 18.4s serialize: 28.9s (66 GB)
INDEX build -> serialize total: 661.7s (11.0 min) [D106837134 baseline: 721s]
recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429

  • This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
  • The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch (C++), test_hnsw_no_init_level0, and test_hnsw_cagra_IP / _base_level_only (Python).

Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)

One deliberate difference from the lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2M" behavior on the inserted point's OWN top level (keep_max_size_level0 && pt_level == 0), so a level>=1 node's level-0 list could be pruned below 2M. The deterministic build gates on the LINK level (keep_max_size_level0 && level == 0), so EVERY node's level-0 list is filled to 2M when keep_max_size_level0 is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2M slot capacity (no overflow) and yields a fuller/denser base layer for CPU IndexHNSWCagra, which is what GpuIndexCagra::copyFrom(IndexHNSWCagra*) reads back. It is INERT for the default build (keep_max_size_level0 defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU IndexHNSWCagra base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU copyFrom expectations.

Reviewed By: pankajsingh88

Differential Revision: D112025877

@meta-cla meta-cla Bot added the CLA Signed label Jul 29, 2026
@meta-codesync

meta-codesync Bot commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

@mnorris11 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D112025877.

@meta-codesync meta-codesync Bot changed the title faiss HNSW: make graph construction deterministic by default (remove lock-based build) faiss HNSW: make graph construction deterministic by default (remove lock-based build) (#5486) Jul 30, 2026
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from 6fb5a07 to 3dd2f60 Compare July 30, 2026 03:05
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from 3dd2f60 to c67411b Compare July 30, 2026 04:59
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from c67411b to 615a20e Compare July 30, 2026 15:27
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from 615a20e to 2a2b181 Compare July 30, 2026 16:15
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from 2a2b181 to a5c5eef Compare July 30, 2026 17:43
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from a5c5eef to 4dd0bde Compare July 30, 2026 19:34
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 30, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch from 4dd0bde to 32f2100 Compare July 30, 2026 23:22
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 31, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32

  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘

  HNSW32,SQ8

  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘

(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@mnorris11
mnorris11 force-pushed the export-D112025877 branch 2 times, most recently from 704d7f4 to ad98bbe Compare July 31, 2026 16:30
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Jul 31, 2026
…lock-based build) (facebookresearch#5486)

Summary:

TLDR: makes the deterministic HNSW graph build the default (and only) float build path, and removes the legacy lock-based one. Inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- original ParlanANN targets flat graphs like Vamana
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32,SQ8
```
  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘
```
(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@meta-codesync meta-codesync Bot changed the title faiss HNSW: make graph construction deterministic by default (remove lock-based build) (#5486) faiss HNSW: add opt-in deterministic lock-free graph build (faiss.deterministic_hnsw) (#5486) Aug 2, 2026
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Aug 2, 2026
…erministic_hnsw) (facebookresearch#5486)

Summary:

TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind FAISS_DETERMINISTIC_HNSW. We plan to test internally, then make this the default flow and remove the existing flow.
--

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` now always uses the deterministic, lock-free build. The lock-based `hnsw_add_vertices` (float) and the opt-in `deterministic_build` flag are removed.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- The binary `IndexBinaryHNSW` keeps its own independent lock-based build (it has no deterministic variant).

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the removed lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32,SQ8
```
  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘
```
(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `deterministic_build` was never serialized (zero references), so removing it is format-neutral. It was a runtime build flag, like `retain_locks`.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the removed lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
@meta-codesync meta-codesync Bot changed the title faiss HNSW: add opt-in deterministic lock-free graph build (faiss.deterministic_hnsw) (#5486) faiss HNSW: add opt-in deterministic lock-free graph build (faiss::hnsw_deterministic_build) (#5486) Aug 5, 2026
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Aug 5, 2026
…sw_deterministic_build) (facebookresearch#5486)

Summary:

TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind `faiss::hnsw_deterministic_build`. We plan to test internally, then make this the default flow and remove the existing flow.
--

HOW TO ENABLE
--
```
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
```
Nothing else is needed, and nothing changes until you do it: the flag defaults
to false, so this diff is a no-op on its own.

Meta-internally the default can instead come from the JustKnob
`faiss/hnsw:enable_deterministic_build`, so a customer ramps and rolls back by
config with no rebuild:
- **C++ consumers** add `//faiss/fb/runtime_config:runtime_config` to their deps
  -- the hook installs itself, no code change. Targeting is by JustKnobs
  condition rule (e.g. `TW_JOB_HANDLE REGEX ...`), since the hook passes no
  switchval.
- **Python consumers** cannot use that target (a `cpp_library` does not link
  into the pyfaiss extension, and the hook's process-phase guard is never
  satisfied under CPython), so they read the knob themselves, with a per-usecase
  switchval:
  ```
  faiss.cvar.hnsw_deterministic_build = justknobs.check(
      "faiss/hnsw:enable_deterministic_build", switchval="<usecase>")
  ```
  Read it once per job, not per index: a mid-job flip would otherwise leave two
  indexes built by different algorithms.

`python/__init__.pyi` now declares `faiss.cvar`. It was missing, so any Python
consumer assigning a faiss global failed Pyre with "No attribute `cvar` in
module `faiss`" -- fixed here rather than with a `pyre-ignore` in each consumer.

Consumer diffs: D114448906 (Laser), D114448902 (VeST), D114448905 (OVIS),
D114448904 (Unicorn). Ads Vector DB needs nothing -- it never constructs HNSW.

WHEN IT IS WORTH ENABLING
--
Only for builds that run multi-threaded. At one thread the existing lock-based
build is *already* deterministic, so the flag buys nothing beyond ~6% build
time. Verified: lock-based at OMP=1 is byte-identical across runs, while at
OMP=8 46% of neighbor slots differ. A consumer that pins
`omp_set_num_threads(1)` should not bother.

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` and `IndexBinaryHNSW::add` pick the build at runtime. Both paths are compiled in and **the default is unchanged** (lock-based); nothing switches until the flag or the hook says so.
- Turning it on:
```
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
```
- Internally the default can also come from the JustKnob `faiss/hnsw:enable_deterministic_build`, via the new `faiss/fb/runtime_config` target (stripped from the OSS export). It installs `faiss::hnsw_deterministic_build_hook`, and `add()` uses the deterministic build when the flag is set OR the hook returns true -- so setting the flag explicitly always wins, and clearing the hook forces the legacy build.
- **That target is opt-in per consumer.** `//faiss:faiss` deliberately does NOT depend on it: core faiss has no Meta-internal dependencies today, and adding `//justknobs` there would push one onto every faiss consumer. A C++ service opts in by adding `//faiss/fb/runtime_config:runtime_config` to its own deps -- no code change, the hook installs itself. Python consumers cannot link a `cpp_library` into the pyfaiss extension, so they assign the global directly:
```
import faiss, pyjk
faiss.cvar.hnsw_deterministic_build = pyjk.check("faiss/hnsw:enable_deterministic_build")
```
- Keeping the hook out of `//faiss:faiss` also keeps faiss's own tests independent of a production config value.
- Caveat of the hook being live: it is re-evaluated on **every** `add()`, so a knob flip reaches running processes without a restart, but two `add()` calls in one process can disagree if the knob flips between them -- an index built in several batches during a ramp could be built partly one way and partly the other.
- The hook exists because the knob cannot simply be written to the global at static-init time: a JustKnobs read before `initFacebook()` aborts the process (folly "unregistered singleton"), which is not catchable.
- **This flag is transitional.** Once the deterministic build is validated internally, the follow-up makes it the only path and deletes the lock-based build, the flag, the hook, and `faiss/fb/runtime_config` entirely.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- `IndexBinaryHNSW::add` is gated by the same flag and shares the same deterministic implementation: `hnsw_add_vertices_deterministic` takes `make_distance_computer` / `set_query` callbacks, so the binary index just supplies its own Hamming distance computer. It keeps its own lock-based `hnsw_add_vertices` for the default path.

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32,SQ8
```
  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘
```
(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `hnsw_deterministic_build` is a namespace-scope global, not a field on `HNSW` or `IndexHNSW`, so it is not part of any serialized struct. Like `retain_locks`, it is a pure runtime build flag.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).
- Measured on a 1M index: written with the flag off, flag flipped on, read back
  -> `neighbors`, `offsets`, `entry_point`, `max_level` byte-identical and
  search ids+distances identical. The flag is read only inside `add()`.
- Mixed-mode append (lock-built graph extended by a deterministic `add()`) was
  measured against brute-force ground truth at 220k and is indistinguishable
  from either pure build -- recall@10 within +-0.02 of both at ef 32/64/128. No
  existing test covers this, since every test builds one way in one process.

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
mnorris11 pushed a commit to mnorris11/faiss that referenced this pull request Aug 6, 2026
…sw_deterministic_build) (facebookresearch#5486)

Summary:

TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind `faiss::hnsw_deterministic_build`. We plan to test internally, then make this the default flow and remove the existing flow.
--

HOW TO ENABLE
--
```
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
```
Nothing else is needed, and nothing changes until you do it: the flag defaults
to false, so this diff is a no-op on its own.

An embedder driving this from its own configuration system should resolve the
value once at startup and assign the global, rather than per index: changing it
between two `add()` calls would leave one index built by each algorithm.

`python/__init__.pyi` now declares `faiss.cvar`. It was missing, so any Python
consumer assigning a faiss global failed Pyre with "No attribute `cvar` in
module `faiss`" -- fixed here rather than with a `pyre-ignore` in each consumer.

WHEN IT IS WORTH ENABLING
--
Only for builds that run multi-threaded. At one thread the existing lock-based
build is *already* deterministic, so the flag buys nothing beyond ~6% build
time. Verified: lock-based at OMP=1 is byte-identical across runs, while at
OMP=8 46% of neighbor slots differ. A consumer that pins
`omp_set_num_threads(1)` should not bother.

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` and `IndexBinaryHNSW::add` pick the build at runtime. Both paths are compiled in and **the default is unchanged** (lock-based); nothing switches until the flag or the hook says so.
- Turning it on:
```
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
```
- A single plain `bool` is the whole API: libfaiss takes no dependency on any configuration system, and an embedder that wants one resolves it itself and assigns the global:
```
faiss.cvar.hnsw_deterministic_build = <your config lookup>
```
- Resolve it once at startup, not per index. `add()` reads the flag each call, so changing it between two `add()` calls on one index would build part of the graph with each algorithm -- reproducible in neither.
- **This flag is transitional.** Once the deterministic build is validated in production, the follow-up makes it the only path and deletes the lock-based build and the flag.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- `IndexBinaryHNSW::add` is gated by the same flag and shares the same deterministic implementation: `hnsw_add_vertices_deterministic` takes `make_distance_computer` / `set_query` callbacks, so the binary index just supplies its own Hamming distance computer. It keeps its own lock-based `hnsw_add_vertices` for the default path.

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32,SQ8
```
  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘
```
(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `hnsw_deterministic_build` is a namespace-scope global, not a field on `HNSW` or `IndexHNSW`, so it is not part of any serialized struct. Like `retain_locks`, it is a pure runtime build flag.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).
- Measured on a 1M index: written with the flag off, flag flipped on, read back
  -> `neighbors`, `offsets`, `entry_point`, `max_level` byte-identical and
  search ids+distances identical. The flag is read only inside `add()`.
- Mixed-mode append (lock-built graph extended by a deterministic `add()`) was
  measured against brute-force ground truth at 220k and is indistinguishable
  from either pure build -- recall@10 within +-0.02 of both at ef 32/64/128. No
  existing test covers this, since every test builds one way in one process.

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Differential Revision: D112025877
…sw_deterministic_build) (facebookresearch#5486)

Summary:

TLDR: adds deterministic HNSW build inspired by ParlayANN. The deterministic build is reproducible AND faster than the lock-based build at every scale and thread count we measured. This is gated behind `faiss::hnsw_deterministic_build`. We plan to test internally, then make this the default flow and remove the existing flow.
--

https://arxiv.org/abs/2305.04359
--

HOW TO ENABLE
--
```
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
```
Nothing else is needed, and nothing changes until you do it: the flag defaults
to false, so this diff is a no-op on its own.

An embedder driving this from its own configuration system should resolve the
value once at startup and assign the global, rather than per index: changing it
between two `add()` calls would leave one index built by each algorithm.

`python/__init__.pyi` now declares `faiss.cvar`. It was missing, so any Python
consumer assigning a faiss global failed Pyre with "No attribute `cvar` in
module `faiss`" -- fixed here rather than with a `pyre-ignore` in each consumer.

WHEN IT IS WORTH ENABLING
--
Only for builds that run multi-threaded. At one thread the existing lock-based
build is *already* deterministic, so the flag buys nothing beyond ~6% build
time. Verified: lock-based at OMP=1 is byte-identical across runs, while at
OMP=8 46% of neighbor slots differ. A consumer that pins
`omp_set_num_threads(1)` should not bother.

similarities to parlayANN:
- add vertices in doubling batches against frozen snapshot
- defer adding reciprocal edges immediately, add them later after parallel phase

differences from ParlayANN:
- re-uses Faiss HNSW pruning in `shrink_neighbor_list`

---

AI (with a bunch of edits) explanation in more detail:
--

What changed
- `IndexHNSW::add` and `IndexBinaryHNSW::add` pick the build at runtime. Both paths are compiled in and **the default is unchanged** (lock-based); nothing switches until the flag or the hook says so.
- Turning it on:
```
C++:     faiss::hnsw_deterministic_build = true;
Python:  faiss.cvar.hnsw_deterministic_build = True
```
- A single plain `bool` is the whole API: libfaiss takes no dependency on any configuration system, and an embedder that wants one resolves it itself and assigns the global:
```
faiss.cvar.hnsw_deterministic_build = <your config lookup>
```
- Resolve it once at startup, not per index. `add()` reads the flag each call, so changing it between two `add()` calls on one index would build part of the graph with each algorithm -- reproducible in neither.
- **This flag is transitional.** Once the deterministic build is validated in production, the follow-up makes it the only path and deletes the lock-based build and the flag.
- The deterministic path now supports the CAGRA level-0 import configuration: `init_level0=false` skips the level-0-only bucket (level 0 is supplied by the imported CAGRA graph), and `keep_max_size_level0` fills the base layer to 2*M. So `IndexHNSWCagra` (CPU) and `GpuIndexCagra::copyTo(IndexHNSWCagra*)` build through the deterministic path.
- `IndexBinaryHNSW::add` is gated by the same flag and shares the same deterministic implementation: `hnsw_add_vertices_deterministic` takes `make_distance_computer` / `set_query` callbacks, so the binary index just supplies its own Hamming distance computer. It keeps its own lock-based `hnsw_add_vertices` for the default path.

Background
--
HNSW construction in Faiss was non-deterministic under parallel builds: multiple runs of `IndexHNSW::add` with the same data and seeds could produce different graphs, a problem for persistence, crash recovery, and replication (the ParlayANN motivation, https://arxiv.org/abs/2305.04359). Sources of non-determinism were: (1) the reciprocal-link write race in `add_links_starting_from_impl`; (2) floating-point distance ties resolved in heap/visitation order; (3) the entry-point bootstrap `#pragma omp critical` race.

Algorithm (adapted from ParlayANN to Faiss's level-batched structure):
- Per level bucket (highest first, deterministic shuffle), points are inserted in prefix-doubling sub-batches (batch sizes 1, 2, 4, ... capped at 2% of the index).
- Phase A (`HNSW::compute_forward_links_deterministic`, parallel): each point greedily descends and computes its forward links against the immutable snapshot from the end of the previous sub-batch, writing only its own neighbor slots. Reciprocal-edge requests are collected, not applied, so this phase is race-free.
- Phase B (`HNSW::merge_reverse_links_deterministic`, parallel): reverse edges are grouped by destination with a fixed-size 256-bucket radix partition on the low bits of `dest` (a small constant bucket count, independent of `ntotal` and thread count, so grouping stays O(edges) in memory), each bucket sorted by `(level, dest)` and merged in parallel. Every affected node is merged exactly once in a total order (distance, ties by id) and re-pruned with the same RNG heuristic. Because every `dest` maps to exactly one bucket, distinct nodes touch disjoint slots (no locks) and the merge is order- and thread-count-independent. The Phase-B parallel-for uses `schedule(static)` — the libomp dynamic dispatcher segfaults in some build configs (the pre-existing lock-based build carried the same warning).

Guarantee: the resulting graph is reproducible across runs at a fixed thread count and, in practice, across thread counts (the merge is fully order-independent). Recall matches the previous default at every efSearch.

## Performance: build time (40M, d=128, M=32, efC=64, 166 threads)
10-round interleaved timing study (one deterministic + one lock-based build per round, so both see identical host conditions):
  deterministic per-round s: 285.58 275.03 280.84 272.62 273.73 272.61 272.05 272.90 269.87 272.29
  lock-based    per-round s: 306.23 352.97 322.76 294.23 339.32 303.04 291.46 341.96 359.31 282.92
  deterministic: min=269.87 mean=274.75 median=272.76 max=285.58 std=4.53
  lock-based:    min=282.92 mean=319.42 median=314.50 max=359.31 std=26.12
  det/lock: mean=0.860 (deterministic ~14% faster), median=0.867
The deterministic build is ~14% faster than the lock-based build at 40M and ~6x more stable run-to-run (std 4.53s vs 26.12s), since it does not depend on lock-contention timing. Peak RSS ~66GB vs ~56GB. Recall matches at every efSearch (byte-identical graph across builds).

## Performance: search time

Back on the deterministic HEAD, tree clean. Here's the matched A/B — same 40M synthetic data, same machine (AMD Genoa, 166 cores),
  search_repeat=100, deterministic (my HEAD) vs lock-based (parent commit). Since my diff doesn't touch search() at all, any difference is
  purely graph structure + measurement noise.

  Search QPS: deterministic vs lock-based (40M synthetic, repeat=100)

  HNSW16
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.828/0.820     │ 170,329 │  177,995 │ −4.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.866/0.862     │ 112,727 │  110,727 │ +1.8% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.886/0.888     │  59,815 │   56,784 │ +5.3% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32
```
  ┌──────────┬─────────────────┬─────────┬──────────┬───────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ   │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │       64 │ 0.935/0.930     │ 110,186 │  108,411 │ +1.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      128 │ 0.958/0.953     │  68,019 │   66,308 │ +2.6% │
  ├──────────┼─────────────────┼─────────┼──────────┼───────┤
  │      256 │ 0.965/0.960     │  37,624 │   35,828 │ +5.0% │
  └──────────┴─────────────────┴─────────┴──────────┴───────┘
```
  HNSW32,SQ8
```
  ┌──────────┬─────────────────┬─────────┬──────────┬────────┐
  │ efSearch │ recall det/lock │ QPS det │ QPS lock │   Δ    │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │       64 │ 0.926/0.934     │ 220,713 │  198,325 │ +11.3% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      128 │ 0.948/0.953     │ 117,504 │  129,173 │  −9.0% │
  ├──────────┼─────────────────┼─────────┼──────────┼────────┤
  │      256 │ 0.961/0.963     │  58,582 │   65,551 │ −10.6% │
  └──────────┴─────────────────┴─────────┴──────────┴────────┘
```
(Low-ef points ef16/32 omitted from the verdict — even at 100 repeats their std is ~8–20%, too noisy; ef128/256 std is ~3–5%.)

Verdict: no search-QPS regression
- Pure HNSW (16, 32): QPS at parity — within ±5%, and actually slightly faster deterministic at the high-recall points (ef128/256), with
  equal-or-better recall.
- HNSW32,SQ8: more scatter (±10%, mixed direction) — but it tracks small correlated recall differences (det ef256 is 0.961 vs 0.963),
  i.e. the two different graphs sit at slightly different recall/QPS operating points, not a systematic slowdown. Search code is
  identical, so this is graph-structure + noise, not a code regression. If you want it pinned down, a recall-matched (interpolated)
  comparison would remove the operating-point confound.
- Bonus: the deterministic build was 2–3× faster in every case (e.g. HNSW32: 277 s vs 527 s; HNSW16: 164 s vs 429 s) — consistent with
  all prior results.

## Single-threaded (OMP_NUM_THREADS=1)
Customers frequently build with OMP=1 or OpenMP disabled, so this case matters. Measured at 1M / d=128 / M=32 / efC=64, single-threaded:
  build time: lock-based 188.36s vs deterministic 176.40s  (0.94x -> deterministic ~6% FASTER)
  peak RSS:   1.6 GB (both, identical)
  recall@10 ef 16/32/64/128: lock-based .8830/.9387/.9676/.9853 vs deterministic .8832/.9381/.9625/.9798
No single-threaded regression: the deterministic build is slightly faster (it avoids the per-node OpenMP lock ops), uses the same memory, and matches recall within noise. Note the lock-based build was already deterministic at a single thread, so single-threaded users lose nothing and gain a small speedup.

## Serialization compatibility
No on-disk format change, verified in `index_read.cpp` / `index_write.cpp`:
- `hnsw_deterministic_build` is a namespace-scope global, not a field on `HNSW` or `IndexHNSW`, so it is not part of any serialized struct. Like `retain_locks`, it is a pure runtime build flag.
- `write_HNSW` / `read_HNSW` and the `IndexHNSW` field layout are unchanged. The subtype fourcc tags, header, CAGRA block, graph CSR (entry_point / max_level / levels / offsets / neighbors / efC / efS), and storage are all as before.
- `keep_max_size_level0` is still serialized only for the CAGRA subtype (`IHc2`/`IHNc`); `init_level0` is build-only (not serialized).
- The deterministic build emits the same HNSW CSR structure (only neighbor content differs), so old indexes read unchanged and new indexes remain readable by older Faiss.
- Verified by the `io_and_retest` serialize -> deserialize -> re-search round-trips in `test_graph_based.py` / `test_hnsw.cpp` (all pass).
- Measured on a 1M index: written with the flag off, flag flipped on, read back
  -> `neighbors`, `offsets`, `entry_point`, `max_level` byte-identical and
  search ids+distances identical. The flag is read only inside `add()`.
- Mixed-mode append (lock-built graph extended by a deterministic `add()`) was
  measured against brute-force ground truth at 220k and is indistinguishable
  from either pure build -- recall@10 within +-0.02 of both at ef 32/64/128. No
  existing test covers this, since every test builds one way in one process.

## CAGRA API for HNSW build on multi-GPU (aka D106837134) — MAST verification
Verified end-to-end on MAST (8x H100 Grand Teton, Approach D, 100M vectors) with this change in the build — the multi-GPU CAGRA -> HNSW graph-build time is comparable to the D106837134 baseline (no regression):
  all_neighbors build: 367.4s   optimize: 231.7s   copyTo: 18.4s   serialize: 28.9s (66 GB)
  INDEX build -> serialize total: 661.7s (11.0 min)   [D106837134 baseline: 721s]
  recall@10 (tiled 100M): ef64 0.7746, ef128 0.8830, ef256 0.9429
- This confirms this CPU-side change builds, links, and runs in the GPU CAGRA binary at scale and does not regress the pipeline. Note the Approach-D run uses copyTo(base_level_only=True), which imports the CAGRA graph directly as HNSW level 0 and skips add(), so it does not itself route through the deterministic add().
- The deterministic CAGRA level-0 import this change adds (the copyTo path with base_level_only=False: init_level0=false skips the level-0 bucket; keep_max_size_level0 fills the base layer) is covered by passing unit tests: `Test_IndexHNSWCagra_BaseLevelOnly_RangeSearch` (C++), `test_hnsw_no_init_level0`, and `test_hnsw_cagra_IP` / `_base_level_only` (Python).


## Behavioral note: level-0 base layer under keep_max_size_level0 (reviewers, please note)
One deliberate difference from the lock-based build, in the CAGRA base-layer case only: the old build gated the "fill the level-0 list up to 2*M" behavior on the inserted point's OWN top level (`keep_max_size_level0 && pt_level == 0`), so a level>=1 node's level-0 list could be pruned below 2*M. The deterministic build gates on the LINK level (`keep_max_size_level0 && level == 0`), so EVERY node's level-0 list is filled to 2*M when `keep_max_size_level0` is set (not only the level-0-only points). This is a strict superset of the old coverage -- it fills exactly to the 2*M slot capacity (no overflow) and yields a fuller/denser base layer for CPU `IndexHNSWCagra`, which is what `GpuIndexCagra::copyFrom(IndexHNSWCagra*)` reads back. It is INERT for the default build (`keep_max_size_level0` defaults to false, so the gate is never true) and never affects a non-CAGRA graph. Called out explicitly so reviewers know the CPU `IndexHNSWCagra` base-layer graph is intentionally denser than the pre-diff build; worth a sanity check against GPU `copyFrom` expectations.

Reviewed By: pankajsingh88

Differential Revision: D112025877
@meta-codesync

meta-codesync Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

This pull request has been merged in a424dcb.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant