The MobileViT-S model is a light-weight, general-purpose vision transformer designed specifically for mobile devices. It introduces a novel perspective by treating Transformers as convolutions, effectively combining the local processing strengths of CNNs with the global representation capabilities of Transformers.
| GPU | IXUCA SDK | Release | Branch |
|---|---|---|---|
| MR-V100 | 4.4.0 | 26.03 | release/26.03 |
| MR-V100 | 4.3.0 | 25.12 | release/25.12 |
Note: 请切换到与您的 SDK 版本对应的 Release 分支进行测试。请勿直接在 master 分支上运行测试,因为 master 分支可能包含与您的本地 SDK 版本不兼容的最新更改。
切换分支命令示例:
git checkout release/26.03
Pretrained model: https://huggingface.co/timm/mobilevit_s.cvnets_in1k
Dataset: https://www.image-net.org/download.php to download the validation dataset.
pip3 install -r ../../igie_common/requirements.txt
pip3 install timm# downloand mobilevit_s.cvnets_in1k from huggingface into ./mobilevit_s.cvnets_in1k
# export onnxmodel from timm
python3 export.py --model-name mobilevit_s.cvnets_in1k --output mobilevit_s.onnx
# use onnxsim optimize onnx model
onnxsim mobilevit_s.onnx mobilevit_s_opt.onnxexport DATASETS_DIR=/Path/to/imagenet_val/
export RUN_DIR=../../igie_common/# Accuracy
bash scripts/infer_mobilevit_s_fp16_accuracy.sh
# Performance
bash scripts/infer_mobilevit_s_fp16_performance.sh| Model | BatchSize | Precision | FPS | Top-1(%) | Top-5(%) |
|---|---|---|---|---|---|
| mobilevit_s | 32 | FP16 | 1827.75 | 77.127 | 93.546 |