Crystal Stability Prediction Metrics

This task measures how effectively a model can triage hypothetical WBM crystals for DFT validation. Models must identify structures that lie on or below a fixed DFT-computed Materials Project convex hull while minimizing costly false positives.

Switch between the full test set, unique prototypes, and each model's 10,000 most-stable predictions to compare overall accuracy, structural diversity, and fixed-budget discovery performance. See Discovery TMI for calibration, element-level, and error-distribution diagnostics.

Methodology: fixed DFT convex hull

Convex Hull Construction in Matbench Discovery

In Matbench Discovery, the convex hull is always constructed from DFT reference energies, not from the ML model’s predicted energies. This is an important methodological choice that differs from some other benchmarking approaches and has several implications. Understanding how the convex hull is constructed is important for correctly interpreting the energy metrics in Matbench Discovery.

What This Means

  • DFT-based hull: When we calculate the distance to the convex hull (Ehull dist) for a material, we compare the model’s predicted formation energy against the DFT-computed convex hull built from Materials Project reference structures.
  • Fixed reference: The hull does not change based on the model’s predictions. All models are evaluated against the same DFT reference hull.
  • Discovery criterion: A material is counted as a “discovery” if the model correctly predicts it to be lower in energy than all known DFT-computed competing phases with the same (reduced) composition in Materials Project. The reference data was pulled on 2023-03-16 (14 GB), database release v2022.10.28.

Why This Matters

This approach means that:

  1. Formation energy MAE = Hull distance MAE: Because both the model’s prediction and the DFT reference are measured on the same energy scale (relative to the same elemental references), the error in formation energy prediction directly equals the error in hull distance prediction. This is a consequence of linear transformations leaving the MAE metric invariant.

  2. Systematic errors are not canceled: If a model has systematic errors (e.g., consistently over- or underpredicting certain elements), these errors will appear in both the formation energy and hull distance metrics. The model cannot “correct” for its own systematic errors by having them affect both the test structures and the reference hull equally.

    • Advantage: Tests absolute accuracy of model predictions against ground truth.
    • Use case: Pre-screening candidates for DFT calculations and evaluating a model’s ability to identify materials below the DFT reference hull.
  3. Different from some literature: Papers like Nature Communications 11:3793 (2020) construct hulls from model predictions, allowing systematic model errors to partially cancel. In that approach, the hull distance MAE can differ from the formation energy MAE.

    • Advantage: Systematic model errors can partially cancel.
    • Use case: Using the model as a complete replacement for DFT.

This distinction is subtle but important for correctly interpreting model performance and making fair comparisons between different benchmarks.

Training data
Openness
Targets
Presets
#Model Acc F1 DAF Prec TNR TPR MAE R2 RMSE Date Added Links
1EquiformerV3+DeNS-OAM0.9780.9316.0740.9280.9870.9330.0180.8680.0672026-04-07
2TECE-OAM-RRA-1.00.9780.9296.0730.9280.9870.9300.0180.8710.0662026-07-05
3EquFlashV20.9780.9296.0690.9280.9870.9300.0180.8730.0662026-06-11
4GRACE-3L-OAM-L0.9770.9256.0410.9230.9860.9260.0180.8750.0652026-07-02
5eSEN-30M-OAM0.9770.9256.0690.9280.9870.9230.0180.8660.0672025-03-17
6PET-OAM-XL0.9770.9246.0750.9290.9870.9200.0190.8640.0682026-01-10
7MatRIS-10M-OAM0.9760.9216.0390.9230.9860.9180.0190.8710.0662025-10-29
8EquFlash0.9750.9195.9830.9150.9840.9220.0190.8710.0662025-06-23
9eqV2 M0.9750.9176.0470.9240.9860.9100.0200.8480.0722024-10-18
10TACE-OAM-L0.9720.9105.8980.9020.9820.9190.0200.8680.0672026-04-09
11Nequip-OAM-XL0.9710.9065.8690.8970.9810.9150.0200.8720.0662025-11-30
12SevenNet-Omni-i12*0.9710.9065.9540.9100.9840.9010.0210.8680.0672026-01-12
13ORB v30.9710.9055.9120.9040.9820.9070.0240.8210.0782025-04-05
14AlphaNet-v1-OAM*0.9680.9015.7470.8790.9770.9240.0240.8310.0762025-05-12
15Allegro-OAM-L0.9660.8955.6740.8670.9740.9230.0220.8680.0672025-09-08
16Nequip-OAM-L0.9670.8935.8230.8900.9800.8950.0220.8650.0682025-09-08
17DPA-3.1-3M-FT0.9630.8845.6670.8660.9740.9030.0230.8690.0672025-06-05
18GRACE-2L-OAM-L0.9640.8835.8400.8930.9810.8740.0220.8620.0682025-09-09
19GRACE-2L-OAM0.9630.8805.7740.8830.9790.8780.0230.8620.0682025-02-06
20ORB v2 MPA0.9650.8806.0410.9240.9870.8410.0280.8240.0772024-10-11
21EquiformerV3+DeNS-MP0.9560.8635.4790.8380.9680.8900.0290.8400.0742026-04-07
22MatterSim v1 5M0.9590.8625.8520.8950.9820.8310.0240.8630.0682024-12-16
23DPA-4.0.1-Pro-MPtrj0.9560.8575.6090.8570.9740.8560.0290.8360.0742026-06-11
24MACE-MPA-00.9540.8525.5820.8530.9730.8510.0280.8420.0732024-12-09
25MatRIS-10M-MP0.9510.8475.4220.8290.9670.8650.0310.8240.0772025-10-29
26eSEN-30M-MP0.9460.8315.2600.8040.9620.8610.0330.8220.0782025-03-17
27GNoME0.9480.8295.5230.8440.9720.8140.0350.7850.0852024-02-03
28GRACE-1L-OAM0.9440.8245.2550.8030.9620.8460.0310.8420.0732025-02-06
29eqV2 S DeNS0.9390.8155.0420.7710.9530.8640.0360.7880.0852024-10-18
30Eqnorm MPtrj0.9290.7864.8440.7410.9460.8380.0400.7990.0832025-05-26
31HIENet0.9290.7774.9320.7540.9520.8010.0410.7930.0842025-07-01
32ORB v2 MPtrj0.9220.7654.7020.7190.9410.8170.0450.7560.0912024-10-14
33Nequip-MP-L0.9210.7614.7040.7190.9420.8090.0430.7910.0842025-09-08
34SevenNet-l3i5*0.9200.7604.6290.7080.9380.8210.0440.7760.0872024-12-10
35Nequix MP0.9140.7514.4550.6810.9280.8360.0440.7820.0862025-08-17
36Allegro-MP-L0.9150.7514.5160.6900.9320.8230.0440.7780.0872025-09-08
37Nequix MP PFT0.9140.7484.4790.6850.9300.8250.0440.7840.0852026-01-08
38GRACE-2L-MPtrj0.8950.6914.1630.6360.9210.7570.0520.7410.0942024-11-21
39MACE-MP-00.8780.6693.7770.5770.8930.7960.0570.6970.1012023-07-14
40CHGNet0.8510.6133.3610.5140.8680.7580.0630.6890.1032023-03-03
41M3GNet0.8120.5692.8820.4410.8130.8030.0750.5850.1182022-09-20

F1 vs Params

The F1 score is the harmonic mean of precision and recall. It is a measure of the model's ability to correctly identify hypothetical crystals in the WBM test set as lying on or below the Materials Project convex hull. Use the axis/color/size selectors to compare models across any pair of metrics and metadata.
  • Params 41 models
Log Scale