On arm64-v8a, setting use_fp16_arithmetic = true makes a ViT-tiny (patch16, 224×224, 256-d output) return 256/256 non-finite values. use_fp16_storage and fp32 are numerically clean on the identical setup (cosine ≥0.9994 vs the ONNX reference).
Environment: ncnn 20260526, arm64-v8a static build, Samsung Galaxy A35 (Exynos 1380, Cortex-A78+A55, asimdhp present, no i8mm), Android 14.
Repro: load the model from file into a raw ncnn::Net, single-threaded, no app framework; set opt.use_fp16_arithmetic = true; run one 1×3×224×224 float input. Output blob is entirely NaN. Set it to false → correct output.
The model was converted from ONNX with pnnx. x86 is unaffected. Happy to provide a reproducing .param/.bin if useful.
On arm64-v8a, setting use_fp16_arithmetic = true makes a ViT-tiny (patch16, 224×224, 256-d output) return 256/256 non-finite values. use_fp16_storage and fp32 are numerically clean on the identical setup (cosine ≥0.9994 vs the ONNX reference).
Environment: ncnn 20260526, arm64-v8a static build, Samsung Galaxy A35 (Exynos 1380, Cortex-A78+A55, asimdhp present, no i8mm), Android 14.
Repro: load the model from file into a raw ncnn::Net, single-threaded, no app framework; set opt.use_fp16_arithmetic = true; run one 1×3×224×224 float input. Output blob is entirely NaN. Set it to false → correct output.
The model was converted from ONNX with pnnx. x86 is unaffected. Happy to provide a reproducing .param/.bin if useful.