MiniMax's 428B-parameter Mixture-of-Experts model with 23B active parameters, sparse attention, function calling, and strong long-horizon coding and agentic performance.