Table of Contents
Fetching ...

Identifying the Catalytic Descriptor of Single-Atom Catalysts in Nitrate Reduction Reaction: An Interpretable Machine-Learning Method

Zhen Zhu, Shan Gao, Jing Zhang, Xuxin Kang, Shunfang Li, Xiangmei Duan

TL;DR

This work tackles the lack of quantitative structure–activity rules for nitrate reduction on single-atom catalysts by combining interpretable ML with density functional theory to screen 286 SACs anchored on BC3 divacancies. SHAP-guided analysis identifies three pivotal features—$N_V$, $D_N$, and $C_N$—and leads to the descriptor $\psi$, together with the O–N–H angle $\theta$, that maps activity and yields 16 high-performance, earth-abundant candidates, including Ti–V–1N1 with $U_L=-0.10$ V. The $\psi$–$U_L$ volcano validates the descriptor as a predictive design tool, linking electronic structure and coordination environment to kinetics and stability under reaction conditions. Overall, the framework enables rapid, sustainable discovery of NO3RR catalysts with non-precious metals, advancing practical nitrate remediation and ammonia synthesis.

Abstract

Elucidating the catalytic descriptor that accurately characterizes the structure-activity relationships of typical catalysts for various important heterogeneous catalytic reactions is pivotal for designing high-efficient catalytic systems. Here, an interpretable machine learning technique was employed to identify the key determinants governing the nitrate reduction reaction ($\rm NO_3RR$) performance across 286 single-atom catalysts (SACs) with the active sites anchored on double-vacancy $\rm BC_3$ monolayers. Through Shapley Additive Explanations (SHAP) analysis with reliable predictive accuracy, we quantitatively demonstrated that, favorable $\rm NO_3RR$ activity stems from a delicate balance among three critical factors: low $\rm N_V$, moderate $\rm D_N$, and specific doping patterns. Building upon these insights, we established a descriptor ($ψ$) that integrates the intrinsic catalytic properties and the intermediate O-N-H angle ($θ$), effectively capturing the underlying structure-activity relationship. Guided by this, we further identified 16 promising catalysts with predicted low limiting potential ($U_{\rm L}$). Importantly, these catalysts are composed of cost-effective non-precious metal elements and are predicted to surpass most reported catalysts, with the best-performing Ti-V-1N1 is predicted to have an ultra-low $U_{\rm L}$ of $-0.10$ V.

Identifying the Catalytic Descriptor of Single-Atom Catalysts in Nitrate Reduction Reaction: An Interpretable Machine-Learning Method

TL;DR

This work tackles the lack of quantitative structure–activity rules for nitrate reduction on single-atom catalysts by combining interpretable ML with density functional theory to screen 286 SACs anchored on BC3 divacancies. SHAP-guided analysis identifies three pivotal features—, , and —and leads to the descriptor , together with the O–N–H angle , that maps activity and yields 16 high-performance, earth-abundant candidates, including Ti–V–1N1 with V. The volcano validates the descriptor as a predictive design tool, linking electronic structure and coordination environment to kinetics and stability under reaction conditions. Overall, the framework enables rapid, sustainable discovery of NO3RR catalysts with non-precious metals, advancing practical nitrate remediation and ammonia synthesis.

Abstract

Elucidating the catalytic descriptor that accurately characterizes the structure-activity relationships of typical catalysts for various important heterogeneous catalytic reactions is pivotal for designing high-efficient catalytic systems. Here, an interpretable machine learning technique was employed to identify the key determinants governing the nitrate reduction reaction () performance across 286 single-atom catalysts (SACs) with the active sites anchored on double-vacancy monolayers. Through Shapley Additive Explanations (SHAP) analysis with reliable predictive accuracy, we quantitatively demonstrated that, favorable activity stems from a delicate balance among three critical factors: low , moderate , and specific doping patterns. Building upon these insights, we established a descriptor () that integrates the intrinsic catalytic properties and the intermediate O-N-H angle (), effectively capturing the underlying structure-activity relationship. Guided by this, we further identified 16 promising catalysts with predicted low limiting potential (). Importantly, these catalysts are composed of cost-effective non-precious metal elements and are predicted to surpass most reported catalysts, with the best-performing Ti-V-1N1 is predicted to have an ultra-low of V.
Paper Structure (10 sections, 5 equations, 7 figures)

This paper contains 10 sections, 5 equations, 7 figures.

Figures (7)

  • Figure 1: (a) Schematic illustration of the catalyst space construction for $\rm TM{@}V_{CC}$ and $\rm TM{@}V_{CB}$-$\rm N_{n=0\hbox{-}4}$ configurations. Atomic species are color-coded as follows: transition metal (TM, sky blue), boron (B, green), carbon (C, brown), and nitrogen (N, silvery-white). (b) Systematic workflow of the four-stage high-throughput screening protocol employed for identifying optimal catalyst candidates.
  • Figure 2: (a) Electrochemical stability analysis through Pourbaix diagram construction for nitrogen species. (b) Competitive adsorption profile comparison between hydrogen protons and nitrate ions ($\rm NO_3^-$) at active sites. (c) Scatter plot visualization of catalyst candidates filtered by thermodynamic criteria, specifically Gibbs free energy changes ($\Delta G$) for critical hydrogenation steps: $\rm {^*NO \rightarrow{^*NOH/^*NHO}}$ and $\rm {^*NH_2 \rightarrow{^*NH_3}}$ (threshold: < 0.6 eV.)
  • Figure 3: (a) Systematic feature engineering framework illustrating attribute characterization and corresponding active site segmentation. (b) Integrated high-throughput screening methodology coupled with comprehensive machine learning workflow architecture.
  • Figure 4: (a) Heat map illustrating Pearson coefficients among eight selected features. (b) Comparative performance evaluation of four sampling methodologies on XGBoost model through ten-fold cross-validation, including accuracy, precision, recall, and F1-score metrics. (c) Comprehensive model assessment of seven machine learning algorithms employing ten-fold cross-validation with over-sampling technique, presenting accuracy, precision, recall, F1-score, and AUC values. (d) Statistical distribution analysis of model performance metrics (F1-score and AUC values) across seven algorithms utilizing ten-fold cross-validation with over-sampling approach.
  • Figure 5: (a) Feature importance visualization through SHAP value heatmap analysis across 260 samples based on the XGBoost model, demonstrating classification efficacy between qualified and unqualified catalysts using a decision threshold of $0.04$. (b) Quantitative representation of feature significance through SHAP value distribution for eight critical characteristics. (c-e) Multivariate analysis of key features: (c) valence electron count: $\rm N_V$, (d) nitrogen atom concentration: $\rm D_N$ and (e) nitrogen doping configuration: $\rm C_N$, presented through composite SHAP dependence plots. Density distributions of catalyst performance are color-coded: qualified (red) and unqualified (blue) samples.
  • ...and 2 more figures