Skip to content

Kimi K3

kimi-k3

Kimi K3 is Moonshot AI's 2.8T-parameter open-weight multimodal reasoning model, designed for complex coding, knowledge work, and long-horizon agentic workflows. It excels at repository-scale development, tool use, debugging, and iterative problem solving across text, images, logs, tests, and runtime feedback. Built with KDA and Attention Residuals for improved computational efficiency, Kimi K3 delivers strong performance on advanced engineering and multimodal reasoning tasks, making it well suited for autonomous coding agents and large-scale production workflows.

Context
1M tokens
Endpoint
Get API KeyCompare

Pricing

Input$3.00 / 1M
Output$15.00 / 1M
Cache Write (5m)$3.00 / 1M
Cache Write (1h)$3.00 / 1M
Cache Read$3.00 / 1M
Web Search$0 / 1M

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint:
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="kimi-k3",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="kimi-k3",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common7 params
modelmessagesmax_tokenstemperaturetop_pstreamtools
Extended4 params
reasoning_effortstream_optionsthinkingextra_body

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

kimi-k3

Decision guidance

Fit the request contract before production.

Use this model when

  • The workload includes code generation or code-oriented text tasks listed in the current catalog record.
  • The required task matches the listed capabilities: text-to-text, text-to-code, translation.
  • The request depends on listed features such as vision, thinking, function-calling.
  • The input fits within the published 1M-token context record, with output and system overhead budgeted separately.

Check before production

  • Confirm the production client uses one of the listed request surfaces: /v1/chat/completions, /v1/responses, /v1/messages.
  • Estimate a representative request from the current Pay As You Go price fields instead of extrapolating from a tiny prompt.
  • Check the observed-availability card and run your own timeout, retry, and fallback test before relying on the model.