Models
| Provider | Model | Input | Output | Context |
|---|---|---|---|---|
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It combines strong multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across 100+ languages. Designed for long-horizon agentic and coding workflows, Muse Glimmer 30B offers a practical balance of capability and deployment efficiency, making it well suited for local coding assistants, multimodal agents, and production workflows that require sustained autonomous execution. ChatAug 9, 2026 | Input$0.35/1M tokens | Output$1.50/1M tokens | Context131K | |
Muse Spark 1.2 is Meta's multimodal reasoning model designed for complex agentic and software engineering workflows. It supports text, image, video, audio, and PDF inputs with text output, and features a 1M-token context window for sustained reasoning across large, multi-stage tasks. Built for flexible multi-agent execution, Muse Spark 1.2 can serve as either a coordinating main agent or a parallel task-focused subagent. With configurable reasoning effort, structured outputs, parallel function calling, and broad coding-harness compatibility, it is well suited for multi-file refactoring, extended debugging, whole-repository generation, and long-horizon development workflows. ChatAug 4, 2026 | Input$1.25/1M tokens | Output$4.25/1M tokens | Context1M | |
Llama 3.2 1B is a lightweight 1-billion-parameter model built for efficient NLP tasks like summarization, conversation, and multilingual analysis. It runs well in low-resource environments, supports eight core languages (and can be fine-tuned for more), making it a good fit for developers who need capable, multilingual AI without heavy compute costs. ChatSep 24, 2024 | Input$0.0025/1M tokens | Output$0.005/1M tokens | - | |
Llama Guard 4 12B by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.05/1M tokens | Output$0.05/1M tokens | Context164K | |
Llama 4 Scout by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.1/1M tokens | Output$0.3/1M tokens | Context1.0M | |
Llama 4 Maverick by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.2/1M tokens | Output$0.6/1M tokens | Context1.0M | |
Llama Guard 3 8b by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.3/1M tokens | Output$0.3/1M tokens | Context131K | |
Llama 3.3 70B Instruct (Free) by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | InputFreeIncluded | OutputFreeIncluded | Context8K | |
Llama 3.3 70B Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.1/1M tokens | Output$0.25/1M tokens | Context128K | |
Llama 3.2 3B Instruct (Free) by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | InputFreeIncluded | OutputFreeIncluded | Context20K | |
Llama 3.2 3B Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.324/1M tokens | Output$0.324/1M tokens | Context131K | |
Llama 3.2 1B Instruct (Free) by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | InputFreeIncluded | OutputFreeIncluded | Context131K | |
Llama 3.2 90B Vision Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$1.20/1M tokens | Output$1.20/1M tokens | Context131K | |
Llama 3.2 11B Vision Instruct (Free) by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | InputFreeIncluded | OutputFreeIncluded | Context131K | |
Llama 3.2 11B Vision Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.049/1M tokens | Output$0.049/1M tokens | Context131K | |
Llama 3.1 8B Instruct (Free) by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | InputFreeIncluded | OutputFreeIncluded | Context131K | |
Llama 3.1 8B Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.02/1M tokens | Output$0.03/1M tokens | Context16K | |
Llama 3.1 70B Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.1/1M tokens | Output$0.28/1M tokens | Context131K | |
Llama 3.1 405B by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$6/1M tokens | Output$6/1M tokens | Context33K | |
Llama 3.1 405B Instruct by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$2.40/1M tokens | Output$2.40/1M tokens | Context33K | |
Llama 3 70B by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.59/1M tokens | Output$0.79/1M tokens | Context128K | |
Llama 3 8B by Meta. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.05/1M tokens | Output$0.08/1M tokens | Context8K |