MiniMax M2
minimax-m2MiniMax-M2 is a compact, high-efficiency model with 10B active (230B total) parameters, optimized for coding and agentic workflows. It delivers near-frontier reasoning and tool use, excels at multi-file coding tasks and compile-run-fix loops, and performs strongly on benchmarks like SWE-Bench and Terminal-Bench. It also handles long-horizon planning and recovery in agent evaluations, ranking among the top open models across reasoning domains. With fast inference and low cost, it’s ideal for large-scale agents and developer assistants — and works best when reasoning is preserved across turns.
- Context
- 196.6K tokens
- Endpoint
Service Status
Status information temporarily unavailable
Apertis cannot confirm the current service state. This is not a report that the model is down.
Pricing
Quick Start
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="minimax-m2", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="minimax-m2",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsmodelmessagesmax_tokenstemperaturetop_pstreamtoolsreasoning_effortstream_optionsthinkingextra_bodyCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.
Speech 2.8 Turbo
MiniMax Speech 2.8 Turbo text-to-speech. Billed per input character; each Chinese (Han) character counts as 2.
- Context
- N/A
- Input
- — (Not priced per input token)
- Output
- — (Not priced per output token)
Speech 2.8 HD
MiniMax Speech 2.8 HD text-to-speech. Billed per input character; each Chinese (Han) character counts as 2.
- Context
- N/A
- Input
- — (Not priced per input token)
- Output
- — (Not priced per output token)
MiniMax M2.5 (Lightning)
MiniMax-M2.5-Lightning is the high-speed variant of the M2.5 series, optimized for low latency, real-time responsiveness, and high-frequency workloads. It retains the core planning and execution strengths of M2.5 while further improving inference efficiency and response speed, making it ideal for interactive applications, rapid coding assistance, and workflow automation. With enhanced cost efficiency and reduced latency, M2.5-Lightning is particularly well suited for high-throughput, always-on deployments and production environments where speed and scalability are critical.
- Context
- 128K
- Input
- $0.30/M
- Output
- $2.20/M
MiniMax M2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. It incorporates advanced multi-agent collaboration, enabling the model to plan, execute, and iteratively refine complex tasks across dynamic environments. Built for production-grade workflows, M2.7 supports tasks such as live debugging, root cause analysis, financial modeling, and full document generation across Word, Excel, and PowerPoint. With strong benchmark performance—including 56.2% on SWE-Pro, 57.0% on Terminal Bench 2, and 1495 ELO on GDPval-AA—it sets a new standard for multi-agent systems in real-world digital workflows.
- Context
- 204.8K
- Input
- $0.30/M
- Output
- $1.20/M