Embodied Intelligence · World Models · Representation Learning
📍 Shenzhen, China | 🎓 Shenzhen University | 📧 Email
CVPR 2026
This work addresses the challenge of open-vocabulary camouflaged object detection, where existing models struggle to identify objects that blend into their surroundings. We construct OVCOD-D, a benchmark containing 10k+ samples with 87 fine-grained semantic categories. The proposed method introduces a sub-description principal contrastive fusion strategy using SVD to eliminate text redundancy and extract discriminative features, along with a Spatial Focusing Gated Linear Unit (SF-GLU) for dynamic spatial feature enhancement. The model achieves 56.4 AP on open-set evaluation, significantly outperforming YOLO-World and other baselines.
aka @Zh1fen
