Commit 8da5b57
Fold multi-GPU CAGRA build into train(), delete trainMultiGpu (facebookresearch#5500)
Summary:
`GpuIndexCagra` had two multi-GPU build entry points totalling 16 positional
arguments, neither of which fits the `Index` API. This removes one, folds the
other into `train()`, and speeds up the build.
**Usage.** Listing more than one device in the config selects the multi-GPU
build; `train()` routes to it:
```python
devices = faiss.Int32Vector()
for i in range(8):
devices.push_back(i)
an = faiss.AllNeighborsCagraConfig()
an.n_clusters = 16 # 0 = auto: max(2 * n_devices, 4)
an.overlap_factor = 2 # must be >= 2
an.ivf_pq_search_batch_size = 8192 # 0 = cuVS default; caps IVF-PQ memory
config = faiss.GpuIndexCagraConfig()
config.graph_degree = 32
config.intermediate_graph_degree = 32
config.build_algo = faiss.graph_build_algo_IVF_PQ
config.devices = devices # >1 device selects the multi-GPU build
config.all_neighbors_params = an
index = faiss.GpuIndexCagra(res, d, faiss.METRIC_L2, config)
index.train(xb) # xb must stay alive until copyTo() completes
cpu_index = faiss.IndexHNSWCagra()
cpu_index.base_level_only = True
index.copyTo(cpu_index) # required: the GPU index is not searchable on this path
```
Float32-only, does not copy `x`, and leaves `index_` empty so `copyTo()` is the
only valid follow-up. Single-GPU behaviour is unchanged when `devices` is empty.
**Deleted `trainMultiGpu`.** Sharded SNMG CAGRA produces per-shard graphs with
zero cross-shard edges by construction, so it needs post-hoc stitching just to
be usable, and it lost to the `all_neighbors` path on both build time and
recall. It had no callers outside the benchmark and one test.
**Folded `trainAllNeighbors` into `train()`.** Its 6 trailing bare scalars now
live in the constructor-time config struct, which is the established Faiss GPU
convention:
| Old argument | Now |
| --- | --- |
| `devices` | `GpuIndexCagraConfig::devices`, also the dispatch predicate |
| `build_algo` (0/1/2) | `GpuIndexCagraConfig::build_algo` |
| `refinement_rate` | `AllNeighborsCagraConfig::refinement_rate` |
| `n_clusters`, `overlap_factor`, `ivfpq_search_batch` | new `AllNeighborsCagraConfig` |
This also kills a live footgun: the old `int build_algo` used an encoding
(0=NN-descent, 1=brute-force, 2=IVF-PQ) that disagreed with the
`graph_build_algo` enum in the same header (0=IVF_PQ, 1=NN_DESCENT). Both
callers set `config.build_algo` and then passed an unrelated int, and the config
field was silently dead. `BRUTE_FORCE` is appended to `graph_build_algo` (at the
end, so existing values do not renumber) and the config field is now the single
source of truth.
**Graph pruning no longer copies the graph off the device.** Previously only
detour counting ran on the GPU and pruning plus reverse-graph construction ran
on the host, which was 35% of total build time.
This part is **temporary and self-removing**. cuVS has no public API to prune a
graph that is already on the device: its exported `helpers::optimize` takes host
matrices, and the dispatch behind it erases the mdspan accessor to host memory,
so cuVS's own device code path is unreachable from outside. Until that is fixed
upstream, this reaches into cuVS's internal headers when the build has them
available, and otherwise falls back to the public host API -- correct either
way, just slower on the fallback. **Open-source builds get the fallback**, since
cuVS does not install those headers; the `train()` docs say so.
An upstream patch adding a `device_matrix_view` overload to
`cuvs::neighbors::cagra::helpers::optimize` is prepared and will be submitted to
rapidsai/cuvs. When it ships, the internal include, the build flag guarding it,
and the fallback branch all get deleted and every build gets the fast path.
At 50M vectors on 8 GPUs the optimize step goes from 119.1s to 1.65s (72x) and
end-to-end build->serialize from 8.0 to 5.4 minutes, recall unchanged.
**Collapsed the benchmark to one path.** With the stitching approaches gone,
`bench_approaches.py` is a single-path tool for validating and tuning the
production build, and reports a per-phase build breakdown plus an efSearch sweep
of recall and QPS.
Deliberately *not* merged into the config:
- `ivf_pq_params` / `ivf_pq_search_params` are still not consulted on this path.
cuVS derives them from the dataset shape and those derived values beat the
static defaults in the faiss structs, so only the knobs cuVS cannot infer are
overridden.
- `faiss::cagra_build_algo` only has `{IVF_PQ, NN_DESCENT}`, so a `BRUTE_FORCE`
config would silently degrade to NN-descent on the single-GPU path;
`train_ex()` now rejects it there. `GpuIndexBinaryCagra` shares this config
struct and has no multi-GPU build, so it rejects `devices.size() > 1` rather
than silently building on one device.
Behaviour changes: the default `build_algo` on the multi-GPU path is now
`IVF_PQ` (the config default) rather than NN-descent (the old argument default),
and `IndexHNSW::num_base_level_search_entrypoints` goes 32 -> 256, which
measured better on both recall and QPS at every efSearch.
Differential Revision: D1146857551 parent da3191e commit 8da5b57
7 files changed
Lines changed: 660 additions & 1284 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1123 | 1123 | | |
1124 | 1124 | | |
1125 | 1125 | | |
1126 | | - | |
1127 | | - | |
| 1126 | + | |
| 1127 | + | |
| 1128 | + | |
1128 | 1129 | | |
1129 | 1130 | | |
1130 | 1131 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
253 | 253 | | |
254 | 254 | | |
255 | 255 | | |
256 | | - | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
| 259 | + | |
| 260 | + | |
257 | 261 | | |
258 | 262 | | |
259 | 263 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
78 | 78 | | |
79 | 79 | | |
80 | 80 | | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
81 | 88 | | |
82 | 89 | | |
83 | 90 | | |
| |||
0 commit comments