Matbench Discovery Logo Matbench Discovery

Column presets:
Discovery test set:
Training data
Openness
Targets
Presets
#Model CPS Acc F1 DAF Prec MAE R2 κSRME RMSD CMDS CDS Params Targets Date Added Links rcut Training Set Org
1EquiformerV3+DeNS-OAM0.9020.9780.9316.0740.9280.0180.8680.1180.059n/an/a30.3MEFSG2026-04-07 6 Å6.6M (113M) MPtrj+OMat24+sAlex
2TECE-OAM-RRA-1.00.9080.9780.9296.0730.9280.0180.8710.0930.0580.5990.233222MEFSG2026-07-05 6 Å6.6M (113M) MPtrj+OMat24+sAlex
3EquFlashV20.9070.9780.9296.0690.9280.0180.8730.0940.058n/an/a44.9MEFSG2026-06-11 6 Å6.6M (113M) MPtrj+OMat24+sAlex
4GRACE-3L-OAM-L0.9000.9770.9256.0410.9230.0180.8750.1210.0580.7390.69942.1MEFSG2026-07-02 6 Å6.6M (113M) MPtrj+OMat24+sAlex
5eSEN-30M-OAM0.8880.9770.9256.0690.9280.0180.8660.170.0610.6250.64830.2MEFSG2025-03-17 6 Å6.6M (113M) MPtrj+OMat24+sAlex
6PET-OAM-XL0.8980.9770.9246.0750.9290.0190.8640.1190.0600.6340.388730MEFSG2026-01-10 n/a6.6M (113M) MPtrj+OMat24+sAlex
7MatRIS-10M-OAM0.8770.9760.9216.0390.9230.0190.8710.2180.0600.5810.67810.4MEFSGM2025-10-29 6 Å6.6M (113M) MPtrj+OMat24+sAlex
8EquFlash0.8880.9750.9195.9830.9150.0190.8710.1580.060n/an/a28.7MEFSG2025-06-23 6 Å6.6M (113M) MPtrj+OMat24+sAlex
9eqV2 M0.5580.9750.9176.0470.9240.0200.8481.7710.0690.5950.28386.6MEFSD2024-10-18 12 Å3.37M (102M) MPtrj+OMat24
10TACE-OAM-L0.8890.9720.9105.8980.9020.0200.8680.1260.0610.6250.74982.9MEFSG2026-04-09 6 Å6.6M (113M) MPtrj+OMat24+sAlex
11Nequip-OAM-XL0.8860.9710.9065.8690.8970.0200.8720.1250.0630.6230.79632.1MEFSG2025-11-30 6 Å6.6M (113M) MPtrj+OMat24+sAlex
12SevenNet-Omni-i12*0.8730.9710.9065.9540.9100.0210.8680.1920.0620.6260.69654.9MEFSG2026-01-12 6 Å243M COSMOSDataset
13ORB v30.8600.9710.9055.9120.9040.0240.8210.210.075n/an/a25.5MEFSG2025-04-05 6 Å6.47M (133M) MPtrj+Alex+OMat24
14AlphaNet-v1-OAM*0.7690.9680.9015.7470.8790.0240.8310.6430.0790.6530.7394.65MEFSG2025-05-12 5 Å6.6M (113M) MPtrj+OMat24+sAlex
15Allegro-OAM-L0.8400.9660.8955.6740.8670.0220.8680.3190.0650.6020.8289.7MEFSG2025-09-08 7 Å6.6M (113M) MPtrj+OMat24+sAlex
16Nequip-OAM-L0.8700.9670.8935.8230.8900.0220.8650.1660.0650.6700.8189.6MEFSG2025-09-08 6 Å6.6M (113M) MPtrj+OMat24+sAlex
17DPA-3.1-3M-FT0.8020.9630.8845.6670.8660.0230.8690.4690.0690.6470.5333.27MEFSG2025-06-05 6 Å163M OpenLAM
18GRACE-2L-OAM-L0.8650.9640.8835.8400.8930.0220.8620.1690.0640.7590.69426.4MEFSG2025-09-09 6 Å6.6M (113M) MPtrj+OMat24+sAlex
19GRACE-2L-OAM0.8370.9630.8805.7740.8830.0230.8620.2940.0670.7510.52712.6MEFSG2025-02-06 6 Å6.6M (113M) MPtrj+OMat24+sAlex
20ORB v2 MPA0.5280.9650.8806.0410.9240.0280.8241.7340.0970.7500.67825.2MEFSD2024-10-11 10 Å3.25M (32.1M) MPtrj+Alex
21EquiformerV3+DeNS-MP0.8300.9560.8635.4790.8380.0290.8400.2750.070n/an/a30.3MEFSG2026-04-07 6 Å146k (1.58M) MPtrj
22MatterSim v1 5M0.7670.9590.8625.8520.8950.0240.8630.5750.0730.6590.7174.55MEFSG2024-12-16 5 Å17M MatterSim
23DPA-4.0.1-Pro-MPtrj0.8400.9560.8575.6090.8570.0290.8360.2110.0690.598n/a22.8MEFSG2026-06-11 6 Å146k (1.58M) MPtrj
24MACE-MPA-00.7950.9540.8525.5820.8530.0280.8420.4120.0730.6590.7289.06MEFSG2024-12-09 6 Å3.37M (12M) MPtrj+sAlex
25MatRIS-10M-MP0.7780.9510.8475.4220.8290.0310.8240.4890.0720.5930.26410.4MEFSGM2025-10-29 6 Å146k (1.58M) MPtrj
26eSEN-30M-MP0.7970.9460.8315.2600.8040.0330.8220.340.0750.5930.41730.1MEFSG2025-03-17 6 Å146k (1.58M) MPtrj
27GNoMEn/a0.9480.8295.5230.8440.0350.785n/an/an/an/a16.2MEFG2024-02-03 5 Å6M (89M) GNoME
28GRACE-1L-OAM0.7610.9440.8245.2550.8030.0310.8420.5170.0720.7490.5523.45MEFSG2025-02-06 6 Å6.6M (113M) MPtrj+OMat24+sAlex
29eqV2 S DeNS0.5220.9390.8155.0420.7710.0360.7881.6760.0760.6150.25831.2MEFSD2024-10-18 12 Å146k (1.58M) MPtrj
30Eqnorm MPtrj0.7560.9290.7864.8440.7410.0400.7990.4080.0840.6390.7111.31MEFSG2025-05-26 6 Å146k (1.58M) MPtrj
31HIENet0.7070.9290.7774.9320.7540.0410.7930.6420.0800.6530.6497.51MEFSG2025-07-01 5 Å146k (1.58M) MPtrj
32ORB v2 MPtrj0.4700.9220.7654.7020.7190.0450.7561.7260.1010.7270.66725.2MEFSD2024-10-14 10 Å146k (1.58M) MPtrj
33Nequip-MP-L0.7300.9210.7614.7040.7190.0430.7910.4660.0860.6290.5739.6MEFSG2025-09-08 6 Å146k (1.58M) MPtrj
34SevenNet-l3i5*0.7140.9200.7604.6290.7080.0440.7760.550.0850.6510.5331.17MEFSG2024-12-10 5 Å146k (1.58M) MPtrj
35Nequix MP0.7290.9140.7514.4550.6810.0440.7820.4460.0850.6030.561708kEFSG2025-08-17 6 Å146k (1.58M) MPtrj
36Allegro-MP-L0.7200.9150.7514.5160.6900.0440.7780.5040.0820.5560.74718.7MEFSG2025-09-08 6 Å146k (1.58M) MPtrj
37Nequix MP PFT0.7550.9140.7484.4790.6850.0440.7840.3070.0870.6160.565708kEFSHG2026-01-08 6 Å154k (1.59M) MPtrj+MDR-MP PBE ωq
38GRACE-2L-MPtrj0.6810.8950.6914.1630.6360.0520.7410.5250.0900.6980.35515.3MEFSG2024-11-21 6 Å146k (1.58M) MPtrj
39MACE-MP-00.6370.8780.6693.7770.5770.0570.6970.6820.0910.6330.6274.69MEFSG2023-07-14 6 Å146k (1.58M) MPtrj
40CHGNet0.4000.8510.6133.3610.5140.0630.6891.7170.0950.5540.297413kEFSGM2023-03-03 5 Å146k (1.58M) MPtrj
41M3GNet0.4280.8120.5692.8820.4410.0750.5851.4090.112n/an/a228kEFSG2022-09-20 5 Å62.8k (188k) MPF
Download table as  Subscribe via RSS
The CPS (Combined Performance Score) is a metric that weights discovery performance (F1), geometry optimization quality (RMSD), and thermal conductivity prediction accuracy (κSRME). Use the radar chart to adjust the importance of each component.

The training set column shows the number of materials used to train the model. For models trained on DFT relaxations, we show the number of distinct frames in parentheses. In cases where only the number of frames is known, we report the number of frames as the training set size. (N=x) in the Model Params column shows the number of estimators if an ensemble was used. DAF = Discovery Acceleration Factor measures how many more stable materials a model finds compared to random selection from the test set. The unique structure prototypes in the WBM test set have a 15.3% rate of stable crystals, meaning the max possible DAF is (32.9k / 215k)^−1 ≈ 6.54.
CPS
F1 50%κSRME 40%RMSD 10%

CPS Progress Over Time

Each point is a model placed at its benchmark inclusion date; the dashed step line traces the running best ("SOTA frontier") CPS v1, so its jumps mark the models that set a new record when they joined the leaderboard. Use the axis/color/size selectors to compare models across any pair of metrics and parameters. The plot shows the same model cohort as the metrics table above, following the active task preset and table filters.

  • Params 41 models
Log Scale

Matbench Discovery is an interactive leaderboard that ranks ML interatomic potentials across crystal stability prediction, geometry optimization, phonons and thermal conductivity, molecular dynamics, and diatomic potential-energy curves.

We rank 41 models covering multiple methodologies including graph neural network (GNN) interatomic potentials, GNN one-shot predictors, iterative Bayesian optimizers and random forests with shallow-learning structure fingerprints.

EquiformerV3+DeNS-OAM leads the Discovery view with the best F1 of 0.931 .

The benchmark exposes accuracy, robustness, and computational-cost trade-offs across these tasks to help users choose models for static and finite-temperature materials simulations.

📖 Important: In Matbench Discovery, the convex hull used to evaluate stability is constructed from DFT reference energies, not from model predictions. This differs from some other benchmarking approaches and has important implications for metric interpretation. See /tasks/discovery for more information.

To cite Matbench Discovery, use:

Riebesell, J., Goodall, R.E.A., Benner, P. et al. A framework to evaluate machine learning crystal stability predictions. Nat Mach Intell 7, 836–847 (2025). https://doi.org/10.1038/s42256-025-01055-1

We welcome new models additions to the leaderboard through GitHub PRs. See the contributing guide for details and ask support questions via GitHub discussion.

Disclaimer: We evaluate how accurately ML models predict several material properties like thermodynamic stability, thermal conductivity, and atomic positions, in all cases using PBE DFT as reference data. Although these properties are important for high-throughput materials discovery, the ranking cannot give a complete picture of a model’s overall ability to drive materials research. A high ranking does not constitute endorsement by the Materials Project.

🆕 New task — Molecular Dynamics. Matbench Discovery now scores how faithfully MLIPs reproduce finite-temperature observables of ab-initio MD (AIMD): radial distribution functions, vibrational density of states, pressure distributions, and single-point energy/force RMSEs. Explore the new metrics on the MD leaderboard.

⚠️ Interpret with caution. The molecular dynamics task is preliminary. The DynaMat v1.0 reference test set and the metrics are still evolving. Treat the current MD metrics and ranking as indicative only, expect changes as test set and metrics evolve.

The reference set currently holds 17 structures (DynaMat v1.0, spanning pure metals, alloys, high-entropy alloys, transition-metal dichalcogenides, perovskites and molecular crystals at 293–1500 K); an upcoming v2 release will grow this AIMD test set to ~100 structures. Collabs to grow it even further welcome! The public reference data intentionally omits energies and forces. Energy/force RMSEs shown here are maintainer-computed private-label diagnostics and are excluded from CMDS, which ranks trajectory-level observables plus speed. All models currently on the leaderboard were run through a unified script, models/run_md.py. If your model isn’t listed, we invite you to run it and submit your metrics via PR.

For details on the MD modeling task, the DynaMat reference set and the CMDS metric, refer to arXiv:2607.03433.

GitHub Activity

Development activity and community engagement of MLIP GitHub repos. Points are sized by number of contributors and colored by number of commits over the last year.