MViTv2_base is an efficient multi-scale vision Transformer model designed specifically for image classification tasks. By employing a multi-scale structure and hierarchical representation, it effectively captures both global and local image features while maintaining computational efficiency. The MViTv2_base has demonstrated excellent performance on multiple standard datasets and is suitable for a variety of visual recognition tasks.
| GPU | IXUCA SDK | Release | Branch |
|---|---|---|---|
| MR-V100 | 4.4.0 | 26.03 | release/26.03 |
| MR-V100 | 4.3.0 | 25.12 | release/25.12 |
Note: 请切换到与您的 SDK 版本对应的 Release 分支进行测试。请勿直接在 master 分支上运行测试,因为 master 分支可能包含与您的本地 SDK 版本不兼容的最新更改。
切换分支命令示例:
git checkout release/26.03
Pretrained model: https://download.openmmlab.com/mmclassification/v0/mvit/mvitv2-base_3rdparty_in1k_20220722-9c4f0a17.pth
Dataset: https://www.image-net.org/download.php to download the validation dataset.
pip3 install -r ../../igie_common/requirements.txt
pip3 install --no-build-isolation mmcv==1.5.3 mmcls==0.24.0# git clone mmpretrain
git clone -b v0.24.0 https://github.com/open-mmlab/mmpretrain.git
# export onnx model
python3 ../../igie_common/export_mmcls.py --cfg mmpretrain/configs/mvit/mvitv2-base_8xb256_in1k.py --weight mvitv2-base_3rdparty_in1k_20220722-9c4f0a17.pth --output mvitv2_base.onnx
# Use onnxsim optimize onnx model
onnxsim mvitv2_base.onnx mvitv2_base_opt.onnxexport DATASETS_DIR=/Path/to/imagenet-val/
export RUN_DIR=../../igie_common/# Accuracy
bash scripts/infer_mvitv2_base_fp16_accuracy.sh
# Performance
bash scripts/infer_mvitv2_base_fp16_performance.sh| Model | BatchSize | Precision | FPS | Top-1(%) | Top-5(%) |
|---|---|---|---|---|---|
| MViTv2-base | 16 | FP16 | 58.76 | 84.226 | 96.848 |