Vectorize RVV FP16 encoding with scalar-compatible rounding - #5603
Vectorize RVV FP16 encoding with scalar-compatible rounding#5603lyd1992 wants to merge 1 commit into
Conversation
Co-authored-by: ww8191201-coder <wanghongyan2025@iscas.ac.cn> Co-authored-by: ihb2032 <hebome@foxmail.com> Co-authored-by: YuanSheng <yuansheng@isrc.iscas.ac.cn>
|
Hi @lyd1992! Thank you for your pull request and welcome to our community. Action RequiredIn order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you. ProcessIn order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA. Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks! |
|
Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks! |
|
Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks! |
RVV
QT_fp16currently inherits scalar encoding. This adds an RVV encoder while preserving the generic codec's byte representation, including its rounding behavior.The generic codec uses ryg's scale-and-round conversion, whose halfway behavior differs from a direct IEEE round-to-nearest-even FP16 conversion. The vector path mirrors its mask, scale, clamp, bias, and sign operations. Chunks containing NaN or infinity use the scalar codec. The generic encoder becomes overridable so the RVV specialization can provide this implementation.
The new regression tests compare encoded bytes with the generic codec for finite values and, separately, for NaN/Inf inputs. Keeping these corpora separate exercises the vector path without allowing special-value fallback to hide an error. Tests include ties, subnormals, overflow boundaries, short tails, and output guards. This PR is independent of the other RVV encoding/distance changes.
Performance
Measured on a native SG2044 RISC-V host (VLEN=128), GCC 15.1, Release
-O3, dynamic dispatch (FAISS_OPT_LEVEL=dd),rv64gcv_zvfhmin/lp64d, one thread pinned to CPU 2. The baseline is official commit80a16564f86530dbf0bfaf96c2b71feffeb5093f; the candidate is that same baseline plus only this kernel change. These measurements were not collected on the newer PR base2ed4c106e9fb9686e7727e5daf8ad6ad1e164109. The affected scalar-quantizer source files are unchanged between those bases, and the submitted kernel differs from the measured one only in comments/formatting.Geometric mean of the dimension-specific speedups at d=32/128/768: FP16 encode 2.434x.
There are three consecutive sessions, each with four alternating ABBA/BAAB blocks per dimension/path: 48 paired blocks for this candidate. A is the baseline and B is the candidate. Each call is calibrated to at least 0.1 s (the shortest formal call in the full campaign was 0.192 s), with three warm-up batches and
n = max(32, floor(32768/d)). Each block uses the ratio of the two-call geometric mean times. Reported speedup is the median of the three session medians; intervals use 5,000 hierarchical bootstrap resamples, resampling sessions and then blocks within each ABBA/BAAB order stratum. The timing columns are separate medians, so their quotient need not equal the paired speedup. All 48 paired blocks favored this candidate.The timed operation is public
ScalarQuantizer::compute_codes, using fixed-seed (718) finite input batches. FP16 input values arefloat(int(rng()%100000)-50000)/113.f.Variability is material at d=32/128: paired-ratio CV is 13.25%/13.47%, and ABBA-vs-BAAB order gaps are 13.46%/4.80%. The d=32/128/768 geometric mean for the individual sessions is 2.419x, 2.040x, and 2.463x. The aggregate 2.434x should therefore not be interpreted as a stable per-run guarantee.
These results describe public encoding/distance throughput on one non-exclusive host, not end-to-end ANN search speedup. The intervals describe these sessions only, with no multiple-comparison correction; they do not establish portability across machines or vector lengths.
Validation
2ed4c106e9fb9686e7727e5daf8ad6ad1e164109: all three independent candidate builds succeeded; this PR's focused C++ suite (NONE: 11 passed, 6 skipped, 0 failed;RISCV_RVV: 13 passed, 4 skipped, 0 failed) and the public-path oracle above pass on native RISC-V. The newly added RVV regression tests are executed and passed in the RVV run, and skipped when RVV is disabled in the NONE run.git diff --checkclean.