Matbench Discovery Logo Matbench Discovery

Column presets:
Discovery test set:
Training data
require = model's training set must include this dataset, exclude = hide models trained on it requireexclude MPtrj (37) OMat24 (18) sAlex (17) MP 2022 (8) MP Graphs (2) Alex (2) MPF (1) GNoME (1) MatterSim (1) OpenLAM (1) COSMOSDataset (1) MDR-MP PBE ω_q (1)
Openness
Targets
Every model predicts energy (E). require/exclude filter by the other predicted outputs; forces are required by default (hides energy-only models) requireexclude forces (F) (42) stress (S) (41) magmoms (M) (3) Hessian (H) (1) forces/stress via
Presets
#Model CPS Acc F1 DAF Prec MAE R2 κSRME RMSD CMDS CDS Params Targets Date Added Links rcut Training Set Org
1 TECE-OAM-RRA-1.0 0.908 0.978 0.929 6.073 0.928 0.018 0.871 0.093 0.058 0.630 0.270 222M EFSG 2026-07-05 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
2 EquFlashV2 0.907 0.978 0.929 6.069 0.928 0.018 0.873 0.094 0.058 n/a n/a 44.9M EFSG 2026-06-11 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
3 EquiformerV3+DeNS-OAM 0.902 0.978 0.931 6.074 0.928 0.018 0.868 0.118 0.059 n/a n/a 30.3M EFSG 2026-04-07 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
4 GRACE-3L-OAM-L 0.900 0.977 0.925 6.041 0.923 0.018 0.875 0.121 0.058 0.739 0.699 42.1M EFSG 2026-07-02 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
5 PET-OAM-XL 0.898 0.977 0.924 6.075 0.929 0.019 0.864 0.119 0.060 0.634 0.388 730M EFSG 2026-01-10 n/a 6.6M (113M) MPtrj+OMat24+sAlex
6 TACE-OAM-L 0.889 0.972 0.910 5.898 0.902 0.020 0.868 0.126 0.061 0.728 0.761 82.9M EFSG 2026-04-09 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
7 eSEN-30M-OAM 0.888 0.977 0.925 6.069 0.928 0.018 0.866 0.17 0.061 0.625 0.648 30.2M EFSG 2025-03-17 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
8 EquFlash 0.888 0.975 0.919 5.983 0.915 0.019 0.871 0.158 0.060 n/a n/a 28.7M EFSG 2025-06-23 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
9 Nequip-OAM-XL 0.886 0.971 0.906 5.869 0.897 0.020 0.872 0.125 0.063 0.623 0.796 32.1M EFSG 2025-11-30 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
10 MatRIS-10M-OAM 0.877 0.976 0.921 6.039 0.923 0.019 0.871 0.218 0.060 0.581 0.678 10.4M EFSGM 2025-10-29 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
11 SevenNet-Omni-i12* 0.873 0.971 0.906 5.954 0.910 0.021 0.868 0.192 0.062 0.626 0.696 54.9M EFSG 2026-01-12 6 Å 243M COSMOSDataset
12 Nequip-OAM-L 0.870 0.967 0.893 5.823 0.890 0.022 0.865 0.166 0.065 0.670 0.818 9.6M EFSG 2025-09-08 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
13 GRACE-2L-OAM-L 0.865 0.964 0.883 5.840 0.893 0.022 0.862 0.169 0.064 0.759 0.694 26.4M EFSG 2025-09-09 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
14 ORB v3 0.860 0.971 0.905 5.912 0.904 0.024 0.821 0.21 0.075 n/a n/a 25.5M EFSG 2025-04-05 6 Å 6.47M (133M) MPtrj+Alex+OMat24
15 DPA-4.0.1-Pro-MPtrj 0.840 0.956 0.857 5.609 0.857 0.029 0.836 0.211 0.069 0.598 n/a 22.8M EFSG 2026-06-11 6 Å 146k (1.58M) MPtrj
16 Allegro-OAM-L 0.840 0.966 0.895 5.674 0.867 0.022 0.868 0.319 0.065 0.602 0.828 9.7M EFSG 2025-09-08 7 Å 6.6M (113M) MPtrj+OMat24+sAlex
17 GRACE-2L-OAM 0.837 0.963 0.880 5.774 0.883 0.023 0.862 0.294 0.067 0.751 0.527 12.6M EFSG 2025-02-06 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
18 EquiformerV3+DeNS-MP 0.830 0.956 0.863 5.479 0.838 0.029 0.840 0.275 0.070 n/a n/a 30.3M EFSG 2026-04-07 6 Å 146k (1.58M) MPtrj
19 DPA-3.1-3M-FT 0.802 0.963 0.884 5.667 0.866 0.023 0.869 0.469 0.069 0.647 0.533 3.27M EFSG 2025-06-05 6 Å 163M OpenLAM
20 eSEN-30M-MP 0.797 0.946 0.831 5.260 0.804 0.033 0.822 0.34 0.075 0.593 0.417 30.1M EFSG 2025-03-17 6 Å 146k (1.58M) MPtrj
21 MACE-MPA-0 0.795 0.954 0.852 5.582 0.853 0.028 0.842 0.412 0.073 0.659 0.728 9.06M EFSG 2024-12-09 6 Å 3.37M (12M) MPtrj+sAlex
22 MatRIS-10M-MP 0.778 0.951 0.847 5.422 0.829 0.031 0.824 0.489 0.072 0.593 0.264 10.4M EFSGM 2025-10-29 6 Å 146k (1.58M) MPtrj
23 AlphaNet-v1-OAM* 0.769 0.968 0.901 5.747 0.879 0.024 0.831 0.643 0.079 0.653 0.739 4.65M EFSG 2025-05-12 5 Å 6.6M (113M) MPtrj+OMat24+sAlex
24 MatterSim v1 5M 0.767 0.959 0.862 5.852 0.895 0.024 0.863 0.575 0.073 0.659 0.717 4.55M EFSG 2024-12-16 5 Å 17M MatterSim
25 GRACE-1L-OAM 0.761 0.944 0.824 5.255 0.803 0.031 0.842 0.517 0.072 0.749 0.552 3.45M EFSG 2025-02-06 6 Å 6.6M (113M) MPtrj+OMat24+sAlex
26 Eqnorm MPtrj 0.756 0.929 0.786 4.844 0.741 0.040 0.799 0.408 0.084 0.639 0.711 1.31M EFSG 2025-05-26 6 Å 146k (1.58M) MPtrj
27 Nequix MP PFT 0.755 0.914 0.748 4.479 0.685 0.044 0.784 0.307 0.087 0.616 0.565 708k EFSHG 2026-01-08 6 Å 154k (1.59M) MPtrj+MDR-MP PBE ωq
28 Nequip-MP-L 0.730 0.921 0.761 4.704 0.719 0.043 0.791 0.466 0.086 0.629 0.573 9.6M EFSG 2025-09-08 6 Å 146k (1.58M) MPtrj
29 Nequix MP 0.729 0.914 0.751 4.455 0.681 0.044 0.782 0.446 0.085 0.603 0.561 708k EFSG 2025-08-17 6 Å 146k (1.58M) MPtrj
30 Allegro-MP-L 0.720 0.915 0.751 4.516 0.690 0.044 0.778 0.504 0.082 0.556 0.747 18.7M EFSG 2025-09-08 6 Å 146k (1.58M) MPtrj
31 SevenNet-l3i5* 0.714 0.920 0.760 4.629 0.708 0.044 0.776 0.55 0.085 0.651 0.533 1.17M EFSG 2024-12-10 5 Å 146k (1.58M) MPtrj
32 HIENet 0.707 0.929 0.777 4.932 0.754 0.041 0.793 0.642 0.080 0.653 0.649 7.51M EFSG 2025-07-01 5 Å 146k (1.58M) MPtrj
33 GRACE-2L-MPtrj 0.681 0.895 0.691 4.163 0.636 0.052 0.741 0.525 0.090 0.698 0.355 15.3M EFSG 2024-11-21 6 Å 146k (1.58M) MPtrj
34 MACE-MP-0 0.637 0.878 0.669 3.777 0.577 0.057 0.697 0.682 0.091 0.633 0.627 4.69M EFSG 2023-07-14 6 Å 146k (1.58M) MPtrj
35 BAM-MP-core 0.584 0.869 0.623 3.682 0.563 0.060 0.712 0.848 0.086 n/a n/a 17M EFSG 2026-07-16 6 Å 146k (1.58M) MPtrj
36 eqV2 M 0.558 0.975 0.917 6.047 0.924 0.020 0.848 1.771 0.069 0.595 0.283 86.6M EFSD 2024-10-18 12 Å 3.37M (102M) MPtrj+OMat24
37 ORB v2 MPA 0.528 0.965 0.880 6.041 0.924 0.028 0.824 1.734 0.097 0.750 0.678 25.2M EFSD 2024-10-11 10 Å 3.25M (32.1M) MPtrj+Alex
38 eqV2 S DeNS 0.522 0.939 0.815 5.042 0.771 0.036 0.788 1.676 0.076 0.615 0.258 31.2M EFSD 2024-10-18 12 Å 146k (1.58M) MPtrj
39 ORB v2 MPtrj 0.470 0.922 0.765 4.702 0.719 0.045 0.756 1.726 0.101 0.727 0.667 25.2M EFSD 2024-10-14 10 Å 146k (1.58M) MPtrj
40 M3GNet 0.428 0.812 0.569 2.882 0.441 0.075 0.585 1.409 0.112 n/a n/a 228k EFSG 2022-09-20 5 Å 62.8k (188k) MPF
41 CHGNet 0.400 0.851 0.613 3.361 0.514 0.063 0.689 1.717 0.095 0.554 0.297 413k EFSGM 2023-03-03 5 Å 146k (1.58M) MPtrj
42 NequIP-GNoME n/a 0.948 0.829 5.523 0.844 0.035 0.785 n/a n/a n/a n/a 16.2M EFG 2024-02-03 5 Å 6M (89M) GNoME
Download table as  Subscribe via RSS
The CPS (Combined Performance Score) is a metric that weights discovery performance (F1), geometry optimization quality (RMSD), and thermal conductivity prediction accuracy (κSRME). Use the radar chart to adjust the importance of each component.

The training set column shows the number of materials used to train the model. For models trained on DFT relaxations, we show the number of distinct frames in parentheses. In cases where only the number of frames is known, we report the number of frames as the training set size. (N=x) in the Model Params column shows the number of estimators if an ensemble was used. DAF = Discovery Acceleration Factor measures how many more stable materials a model finds compared to random selection from the test set. The unique structure prototypes in the WBM test set have a 15.3% rate of stable crystals, meaning the max possible DAF is (32.9k / 215k)^−1 ≈ 6.54.
CPS
F1 50%κSRME 40%RMSD 10%

CPS Progress Over Time

Each point is a model placed at its benchmark inclusion date; the dashed step line traces the running best ("SOTA frontier") CPS v1, so its jumps mark the models that set a new record when they joined the leaderboard. Use the axis/color/size selectors to compare models across any pair of metrics and parameters. The plot shows the same model cohort as the metrics table above, following the active task preset and table filters.

  • Params 42 models
Log Scale

Matbench Discovery is an interactive leaderboard that ranks ML interatomic potentials across crystal stability prediction, geometry optimization, phonons and thermal conductivity, molecular dynamics, and diatomic potential-energy curves.

We rank 42 models covering multiple methodologies including graph neural network (GNN) interatomic potentials, GNN one-shot predictors, iterative Bayesian optimizers and random forests with shallow-learning structure fingerprints.

EquiformerV3+DeNS-OAM leads the Discovery view with the best F1 of 0.931 .

The benchmark exposes accuracy, robustness, and computational-cost trade-offs across these tasks to help users choose models for static and finite-temperature materials simulations.

📖 Important: In Matbench Discovery, the convex hull used to evaluate stability is constructed from DFT reference energies, not from model predictions. This differs from some other benchmarking approaches and has important implications for metric interpretation. See /tasks/discovery for more information.

To cite Matbench Discovery, use:

Riebesell, J., Goodall, R.E.A., Benner, P. et al. A framework to evaluate machine learning crystal stability predictions. Nat Mach Intell 7, 836–847 (2025). https://doi.org/10.1038/s42256-025-01055-1

Are you submitting a new model? Follow the contributing guide, including its branch-specific PR instructions for loading the model checklist. For other changes, use GitHub’s documented branch selector to open a regular PR with a blank description. Ask support questions via GitHub discussion.

Disclaimer: We evaluate how accurately ML models predict several material properties like thermodynamic stability, thermal conductivity, and atomic positions, in all cases using PBE DFT as reference data. Although these properties are important for high-throughput materials discovery, the ranking cannot give a complete picture of a model’s overall ability to drive materials research. A high ranking does not constitute endorsement by the Materials Project.

🆕 New task — Molecular Dynamics. Matbench Discovery now scores how faithfully MLIPs reproduce finite-temperature observables of ab-initio MD (AIMD): radial distribution functions, vibrational density of states, pressure distributions, and single-point energy/force RMSEs. Explore the new metrics on the MD leaderboard.

⚠️ Interpret with caution. The molecular dynamics task is preliminary. The DynaMat v1.0 reference test set and the metrics are still evolving. Treat the current MD metrics and ranking as indicative only, expect changes as test set and metrics evolve.

The reference set currently holds 17 structures (DynaMat v1.0, spanning pure metals, alloys, high-entropy alloys, transition-metal dichalcogenides, perovskites and molecular crystals at 293–1500 K); an upcoming v2 release will grow this AIMD test set to ~100 structures. Collabs to grow it even further welcome! The public reference data intentionally omits energies and forces. Energy/force RMSEs shown here are maintainer-computed private-label diagnostics and are excluded from CMDS, which ranks trajectory-level observables plus speed. All models currently on the leaderboard were run through a unified script, models/run_md.py. If your model isn’t listed, we invite you to run it and submit your metrics via PR.

For details on the MD modeling task, the DynaMat reference set and the CMDS metric, refer to arXiv:2607.03433.

GitHub Activity

Development activity and community engagement of MLIP GitHub repos. Points are sized by number of contributors and colored by number of commits over the last year.