DINOv2 is a self-supervised vision Transformer model designed for general-purpose image representation learning. This example uses the ViT-S/14 pretrained backbone for ImageNet linear evaluation: the frozen backbone extracts intermediate features, and a linear classifier is trained and evaluated on top of them.
| GPU | IXUCA SDK | Release |
|---|---|---|
| MR-V100 | 4.4.0 | 26.06 |
visit https://github.com/facebookresearch/dinov2 to see the official repo.
Pretrained model: https://dl.fbaipublicfiles.com/dinov2/dinov2_vits14/dinov2_vits14_pretrain.pth
Dataset: https://www.image-net.org/download.php to download the ImageNet-1K dataset.
git clone https://github.com/facebookresearch/dinov2.git
cp -r eval dinov2/dinov2
cp dinov2-patch/* dinov2/
pip3 install omegaconf==2.3.0 fvcore iopath submitit==1.5.4 torchmetricscd dinov2
wget https://github.com/facebookresearch/dinov2/files/13427966/labels.txt -O /mnt/deepspark/data/datasets/imagenet
python3 prepare_datasets.py
python3 export.py \
--config-file dinov2/configs/eval/vits14_pretrain.yaml \
--pretrained-weights dinov2_vits14_pretrain.pth \
--onnx-path dinov2_vits14_pretrain.onnx \
--input-size 224 \
--n-last-blocks 4 \
--device cuda
export batchsize=64
python3 build_engine.py \
--model_path dinov2_vits14_pretrain.onnx \
--input input:${batchsize},3,224,224 \
--precision fp16 \
--engine_path dinov2_vits14_pretrain_bs_${batchsize}_fp16.soexport PYTHONPATH=/path/to/dinov2:${PYTHONPATH}
export IMAGENET_1K=/path/to/imagenet
# ACC
bash scripts/infer_dinov2_fp16_accuracy.sh
# Performence
bash scripts/infer_dinov2_fp16_performance.sh| Model | Task | BatchSize | Precision | FPS | Accuracy |
|---|---|---|---|---|---|
| DINOv2 ViT-S/14 | Linear Eval | 64 | FP16 | 889.542 | 81.1% |