Table of Contents
Fetching ...

Large-scale stochastic propagation method beyond the sequential approach

Zhichang Fu, Yunhai Li, Weiqing Zhou, Shengjun Yuan

TL;DR

The paper tackles the bottleneck of sequential time propagation in large-scale quantum simulations by introducing a concurrent stochastic propagation framework based on Chebyshev polynomial expansion. It provides three tailored implementations—state-based, moment-based, and energy-based—along with a time-blocking strategy to balance memory and speed, enabling single long-time propagations to reconstruct all intermediate states. The approach yields up to an order-of-magnitude acceleration for properties like DOS, QE, EC, OC, DP, and CD in billion-atom tight-binding systems, with memory overhead kept modest and numerical accuracy preserved to machine precision. This method broadens the applicability of linear-scaling stochastic propagation to diverse electronic-structure problems and can extend to first-principles DFT calculations with orthogonal bases, offering a significant practical impact for simulating complex quantum materials at ultra-large scales.

Abstract

The $O(N)$ stochastic propagation method, which relies on the numerical solution of the time-dependent Schrödinger equation using random initial states, is widely used in large-scale first-principles calculations. In this work, we eliminate the conventional sequential computation of intermediate states by introducing a concurrent strategy that minimizes information redundancy. The new method, in its state-, moment-, and energy-based implementations, not only surpasses the time step constraint of sequential propagation but also maintains precision within the framework of the Nyquist-Shannon sampling theorem. Systematic benchmarking on one billion atoms within the tight-binding model demonstrates that our new concurrent method achieves up to an order-of-magnitude speedup, enabling the rapid computation of a wide range of electronic, optical, and transport properties. This performance breakthrough offers valuable insights for enhancing other time-propagation algorithms, including those employed in large-scale stochastic density functional theory.

Large-scale stochastic propagation method beyond the sequential approach

TL;DR

The paper tackles the bottleneck of sequential time propagation in large-scale quantum simulations by introducing a concurrent stochastic propagation framework based on Chebyshev polynomial expansion. It provides three tailored implementations—state-based, moment-based, and energy-based—along with a time-blocking strategy to balance memory and speed, enabling single long-time propagations to reconstruct all intermediate states. The approach yields up to an order-of-magnitude acceleration for properties like DOS, QE, EC, OC, DP, and CD in billion-atom tight-binding systems, with memory overhead kept modest and numerical accuracy preserved to machine precision. This method broadens the applicability of linear-scaling stochastic propagation to diverse electronic-structure problems and can extend to first-principles DFT calculations with orthogonal bases, offering a significant practical impact for simulating complex quantum materials at ultra-large scales.

Abstract

The stochastic propagation method, which relies on the numerical solution of the time-dependent Schrödinger equation using random initial states, is widely used in large-scale first-principles calculations. In this work, we eliminate the conventional sequential computation of intermediate states by introducing a concurrent strategy that minimizes information redundancy. The new method, in its state-, moment-, and energy-based implementations, not only surpasses the time step constraint of sequential propagation but also maintains precision within the framework of the Nyquist-Shannon sampling theorem. Systematic benchmarking on one billion atoms within the tight-binding model demonstrates that our new concurrent method achieves up to an order-of-magnitude speedup, enabling the rapid computation of a wide range of electronic, optical, and transport properties. This performance breakthrough offers valuable insights for enhancing other time-propagation algorithms, including those employed in large-scale stochastic density functional theory.
Paper Structure (13 sections, 47 equations, 14 figures, 1 table)

This paper contains 13 sections, 47 equations, 14 figures, 1 table.

Figures (14)

  • Figure 1: Comparison of scalability and computational performance for the new concurrent and the conventional sequential sPM in large-scale systems. (a) Speedup as a function of the number of atoms ($N_a$) for density of states (DOS), electronic conductivity (EC) and quasi-eigenstates (QE). The horizontal dashed lines indicate the average speedup for each corresponding data series. Note that for the electronic conductivity, the data point for the largest system is rescaled from a calculation on $9.6 \times 10^8$ atoms, which was the maximum size feasible due to memory limitations. (b) Wall time as a function of the number of atoms, $N_a$. The plotted lines represent the power-law fits to the corresponding data, with the scaling exponent $c$ provided in the legend in the form of $N_a^c$. These tests utilized nearest-neighbor, single-layer graphene models 1947-Wallace with system sizes spanning from 50 million to 1 billion atoms. The calculations were executed on a computational node of 512 GB of RAM, equipped with two Intel® Xeon® Gold 6326 processors (32 cores total).
  • Figure 2: The Numerical Foundations of Concurrent Stochastic Propagation Method. (a) Expansion load $R(\tilde{\tau}) = N(\tilde{\tau})/\tilde{\tau}$ (the number of polynomial decomposition per unit of time) vs. the rescaled time step $\tilde{\tau}$ for Chebyshev, Legendre, and Taylor expansions of the time-evolution operator $\mathrm{e}^{-\mathrm{i}\tilde{H}\tilde{\tau}}$. (b) $R(\tilde{\tau})$ vs. $\tilde{\tau}$ for the Chebyshev expansion, with $\tilde{\tau} = \pi$ as a reference. Points are marked where $R(\tilde{\tau})$ is reduced to $\tfrac{1}{2}$, $\tfrac{1}{3}$, $\tfrac{1}{4}$, and $\tfrac{1}{5}$ of the reference value.
  • Figure 3: Accuracy validation of the new concurrent method for the density of states calculation. The density of states (DOS) from the DOS-Concurrent-M method (red line) is compared against the baseline DOS-Sequential method (black line). The calculation was performed on a magic-angle twisted bilayer graphene supercell containing 4,763,200 atoms ($D_H=230$, $N_t=4096$). The wall time for each method is provided in parentheses within the legend.
  • Figure 4: Performance comparison of the DOS-Concurrent-M method vs. the baseline DOS-Sequential method. (a) Speedup (line) and relative memory consumption (bars) as a function of the total number of time steps $N_t$ ($D_H=230$). (b) Speedup as a function of the Hamiltonian matrix density $D_H$ ($N_t=4096$).
  • Figure 5: Accuracy validation of the new concurrent method for the quasi-eigenstates calculation. The quasi-eigenstates (QE) calculated by the QE-Concurrent-S (blue line) and QE-Concurrent-E (red line) methods are compared against the baseline QE-Sequential method (black line). For visual clarity, the plot displays the results for a single quasi-eigenstates on the first 100 atoms. The calculation was performed on a magic-angle twisted bilayer graphene supercell containing 4,763,200 atoms ($D_H=120$, $N_t=4096$, $N_E=5$). The wall time for each method is provided in parentheses within the legend.
  • ...and 9 more figures