Skip to content
MetaChat

Llama 3.2 1B Instruct (Free)

llama-3.2-1b-instruct:free

Llama 3.2 1B Instruct (Free) by Meta.

Context
131K tokens
Endpoint
Get API KeyCompare

Pricing

Input$0 / 1M
Output$0 / 1M
Cache Write$0 / 1M
Cache Read$0 / 1M
Prompt cache writes and reads are included at no additional cost.

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint:
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="llama-3.2-1b-instruct:free",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="llama-3.2-1b-instruct:free",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common7 params
modelmessagesmax_tokenstemperaturetop_pstreamtools
Extended4 params
reasoning_effortstream_optionsthinkingextra_body

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

llama-3.2-1b-instruct:free

Decision guidance

Fit the request contract before production.

Use this model when

  • The workload fits the text or chat tasks shown in this model's current catalog record.
  • The input fits within the published 131K-token context record, with output and system overhead budgeted separately.

Check before production

  • Confirm the production client uses one of the listed request surfaces: /v1/chat/completions, /v1/responses, /v1/messages.
  • Estimate a representative request from the current Free price fields instead of extrapolating from a tiny prompt.
  • Check the observed-availability card and run your own timeout, retry, and fallback test before relying on the model.

Compare with Other Models

See how this model compares to others from the same provider.

Muse Spark 1.3

Muse Spark 1.3 is Meta's multimodal reasoning model designed for long-running agentic, multi-agent, and coding workflows. It maintains context and information across extended tasks, enabling reliable execution in complex, multi-step environments. The model is optimized to resolve conflicting information, seek clarification or confirmation when necessary, and execute concisely, making it well suited for autonomous agents, collaborative multi-agent systems, and long-horizon software engineering workflows.

Context
1M
Input
$1.25/M
Output
$4.25/M

Muse Glimmer 30B

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It combines strong multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across 100+ languages. Designed for long-horizon agentic and coding workflows, Muse Glimmer 30B offers a practical balance of capability and deployment efficiency, making it well suited for local coding assistants, multimodal agents, and production workflows that require sustained autonomous execution.

Context
131K
Input
$0.35/M
Output
$1.50/M

Muse Spark 1.2

Muse Spark 1.2 is Meta's multimodal reasoning model designed for complex agentic and software engineering workflows. It supports text, image, video, audio, and PDF inputs with text output, and features a 1M-token context window for sustained reasoning across large, multi-stage tasks. Built for flexible multi-agent execution, Muse Spark 1.2 can serve as either a coordinating main agent or a parallel task-focused subagent. With configurable reasoning effort, structured outputs, parallel function calling, and broad coding-harness compatibility, it is well suited for multi-file refactoring, extended debugging, whole-repository generation, and long-horizon development workflows.

Context
1M
Input
$1.25/M
Output
$4.25/M

Llama 3.2 1b Instruct

Llama 3.2 1B is a lightweight 1-billion-parameter model built for efficient NLP tasks like summarization, conversation, and multilingual analysis. It runs well in low-resource environments, supports eight core languages (and can be fine-tuned for more), making it a good fit for developers who need capable, multilingual AI without heavy compute costs.

Context
N/A
Input
$0.0025/M
Output
$0.005/M