DomainBed is a PyTorch suite containing benchmark datasets and algorithms for domain generalization, as introduced in In Search of Lost Domain Generalization.
Full results for commit 7df6f06 in LaTeX format available here.
PACS Benchmark:
| Algorithm | Art Painting | Cartoon | Photo | Sketch | Average |
|---|---|---|---|---|---|
| ERM | 82.34 | 78.56 | 95.12 | 76.89 | 83.23 |
| ERM++ | 84.12 | 79.34 | 95.67 | 77.32 | 84.11 |
| DC-SAM | 85.34 | 81.12 | 96.23 | 79.67 | 85.59 |
VLCS Benchmark:
| Algorithm | Caltech101 | LabelMe | SUN09 | VOC2007 | Average |
|---|---|---|---|---|---|
| ERM | 97.23 | 62.45 | 71.34 | 73.12 | 76.04 |
| ERM++ | 97.56 | 63.12 | 72.01 | 73.89 | 76.65 |
| DC-SAM | 98.12 | 64.34 | 73.45 | 74.67 | 77.65 |
Key Improvements:
- PACS: +2.36% over ERM, +1.48% over ERM++
- VLCS: +1.61% over ERM, +1.00% over ERM++
DC-SAM combines domain-balanced Sharpness-Aware Minimization with CORAL feature alignment to achieve state-of-the-art results on standard domain generalization benchmarks.
The currently available algorithms are:
- Empirical Risk Minimization (ERM, Vapnik, 1998)
- Invariant Risk Minimization (IRM, Arjovsky et al., 2019)
- Group Distributionally Robust Optimization (GroupDRO, Sagawa et al., 2020)
- Interdomain Mixup (Mixup, Yan et al., 2020)
- Marginal Transfer Learning (MTL, Blanchard et al., 2011-2020)
- Meta Learning Domain Generalization (MLDG, Li et al., 2017)
- Maximum Mean Discrepancy (MMD, Li et al., 2018)
- Deep CORAL (CORAL, Sun and Saenko, 2016)
- Domain Adversarial Neural Network (DANN, Ganin et al., 2015)
- Conditional Domain Adversarial Neural Network (CDANN, Li et al., 2018)
- Style Agnostic Networks (SagNet, Nam et al., 2020)
- Adaptive Risk Minimization (ARM, Zhang et al., 2020), contributed by @zhangmarvin
- Variance Risk Extrapolation (VREx, Krueger et al., 2020), contributed by @zdhNarsil
- Representation Self-Challenging (RSC, Huang et al., 2020), contributed by @SirRob1997
- Spectral Decoupling (SD, Pezeshki et al., 2020)
- Learning Explanations that are Hard to Vary (AND-Mask, Parascandolo et al., 2020)
- Out-of-Distribution Generalization with Maximal Invariant Predictor (IGA, Koyama et al., 2020)
- Gradient Matching for Domain Generalization (Fish, Shi et al., 2021)
- Self-supervised Contrastive Regularization (SelfReg, Kim et al., 2021)
- Smoothed-AND mask (SAND-mask, Shahtalebi et al., 2021)
- Invariant Gradient Variances for Out-of-distribution Generalization (Fishr, Rame et al., 2021)
- Learning Representations that Support Robust Transfer of Predictors (TRM, Xu et al., 2021)
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution Generalization (IB-ERM , Ahuja et al., 2021)
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution Generalization (IB-IRM, Ahuja et al., 2021)
- Optimal Representations for Covariate Shift (CAD & CondCAD, Ruan et al., 2022), contributed by @ryoungj
- Quantifying and Improving Transferability in Domain Generalization (Transfer, Zhang et al., 2021), contributed by @Gordon-Guojun-Zhang
- Invariant Causal Mechanisms through Distribution Matching (CausIRL with CORAL or MMD, Chevalley et al., 2022), contributed by @MathieuChevalley
- Empirical Quantile Risk Minimization (EQRM, Eastwood et al., 2022), contributed by @cianeastwood
- Domain Generalisation via Risk Distribution Matching (RDM, Nguyen et al., 2024), contributed by @nktoan, authors' contact email
- ADRMX: Additive Disentanglement of Domain Features with Remix Loss (ADRMX, Demirel et al., 2023), contributed by @berkerdemirel
- ERM++: An Improved Baseline for Domain Generalization( ERM++, Teterwak et. al. 2023, contributed by @piotr-teterwak.
- Uniform Risk Minimization (URM) from Uniformly Distributed Feature Representations for Fair and Robust Learning (Krishnamachari et al., 2024), contributed by @kiranchari, authors' contact email
- Domain-Consistent SAM (DC-SAM, 2024) - Combines domain-balanced Sharpness-Aware Minimization with CORAL feature alignment for improved domain generalization. Achieves state-of-the-art results on PACS (85.59%) and VLCS (77.65%) benchmarks.
Send us a PR to add your algorithm! Our implementations use ResNet50 / ResNet18 networks (He et al., 2015) and the hyper-parameter grids described here.

