Hardware
MiniMax M2.7 Advances Scalable Agentic Workflows on NVIDIA Platforms for Complex AI Applications
MiniMax M2.7 is an enhancement of the MiniMax M2.5 model, built as a 230B-parameter Mixture-of-Experts (MoE) model with 10B active parameters per token and a 200K context length, designed for agentic
MiniMax M2.7 is an enhancement of the MiniMax M2.5 model, built as a 230B-parameter Mixture-of-Experts (MoE) model with 10B active parameters per token and a 200K context length, designed for agentic workflows and complex use cases spanning reasoning, ML research, software engineering, and office productivity. To maximize performance, NVIDIA collaborated with the open-source community to integrate high-performance kernels into vLLM and SGLang, with deployment options ranging from data center deployments on NVIDIA Blackwell to fully managed NVIDIA NIM microservices. M2.7 is notably the first model in the M2-series to deeply participate in its own evolution, capable of building complex agent harnesses and completing highly elaborate productivity tasks through Agent Teams, complex Skills, and dynamic tool search.
Related
- glm 5.1 is doing well
- ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
- New course: Efficient Inference with SGLang: Text and Image Generation, built in partnership with LMSys @lmsysorg and RadixArk @radixark, an…
- Nvidia published DWDP (Distributed Weight-Data Parallelism), a new inference parallelism strategy focused on prefill. It sounds slightly ins…
- model page: https://ollama.com/library/minimax-m2.7
Source: NVIDIA Developer | 2026-04-12