TeLU Activation Function for Fast and Stable Deep Learning
Alfredo Fernandez, Ankur Mali
TL;DR
TeLU defines the activation $TeLU(x)=x\tanh(e^{x})$, a smooth, analytic, non-monotonic function designed to combine ReLU-like fast convergence with stronger gradient flow in the saturating regime. The paper establishes TeLU as a universal analytic approximator and proves theoretical properties including persistent inactive-region gradients, near-linearity in the active region, computational efficiency, ReLU compatibility, and stability. Extensive experiments across MLPs, CNNs (DenseNet, ResNet), Transformers, VAEs, and RNNs (PTB) on ImageNet, Text8, CIFAR, MNIST, and more demonstrate superior convergence speed, robustness to initialization and depth, and competitive or improved accuracy relative to ReLU and other smooth activations. The work also discusses practical benefits such as easier ReLU replacement, compatibility with second-order optimization, and potential for broader application, including robustness to perturbations and efficient hardware deployment.
Abstract
We propose the Hyperbolic Tangent Exponential Linear Unit (TeLU), a neural network hidden activation function defined as TeLU(x)=xtanh(exp(x)). TeLU's design is grounded in the core principles of key activation functions, achieving strong convergence by closely approximating the identity function in its active region while effectively mitigating the vanishing gradient problem in its saturating region. Its simple formulation enhances computational efficiency, leading to improvements in scalability and convergence speed. Unlike many modern activation functions, TeLU seamlessly combines the simplicity and effectiveness of ReLU with the smoothness and analytic properties essential for learning stability in deep neural networks. TeLU's ability to mimic the behavior and optimal hyperparameter settings of ReLU, while introducing the benefits of smoothness and curvature, makes it an ideal drop-in replacement. Its analytic nature positions TeLU as a powerful universal approximator, enhancing both robustness and generalization across a multitude of experiments. We rigorously validate these claims through theoretical analysis and experimental validation, demonstrating TeLU's performance across challenging benchmarks; including ResNet18 on ImageNet, Dynamic-Pooling Transformers on Text8, and Recurrent Neural Networks (RNNs) on the Penn TreeBank dataset. These results highlight TeLU's potential to set a new standard in activation functions, driving more efficient and stable learning in deep neural networks, thereby accelerating scientific discoveries across various fields.
