Advancing Emerging Optimizers for Accelerated LLM Training with NVIDIA Megatron
22 April 2026 at 20:01
Higher-order optimization algorithms such as Shampoo have been effectively applied in neural network training for at least a decade. These methods have achieved significant success more recently when applied to leading LLMs. In particular, Muon (MomentUm Orthogonalized by Newton-Schulz) was used to train some of todayβs best open source models, including Kimi K2 and GLM-5.