Skip to content

Molmo2 8B (Free)

molmo-2-8b:free

Molmo2-8B is an open vision-language model from AI2 that supports image, video, and multi-image understanding. Built on Qwen3-8B with a SigLIP 2 vision backbone, it excels at short-video tasks like counting and captioning while remaining competitive on long-video understanding, outperforming other open-weight, open-data models in its class.

Context
128K tokens
Endpoint
Get API KeyCompare

Pricing

Input$0 / 1M
Output$0 / 1M
Cache Write$0 / 1M
Cache Read$0 / 1M
Prompt cache writes and reads are included at no additional cost.

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint:
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="molmo-2-8b:free",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="molmo-2-8b:free",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common7 params
modelmessagesmax_tokenstemperaturetop_pstreamtools
Extended4 params
reasoning_effortstream_optionsthinkingextra_body

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

molmo-2-8b:free

Decision guidance

Fit the request contract before production.

Use this model when

  • The workload fits the text or chat tasks shown in this model's current catalog record.
  • The required task matches the listed capabilities: text-to-text, translation, summarization.
  • The request depends on listed features such as vision, function-calling, json-mode.
  • The input fits within the published 128K-token context record, with output and system overhead budgeted separately.

Check before production

  • Confirm the production client uses one of the listed request surfaces: /v1/chat/completions, /v1/responses, /v1/messages.
  • Estimate a representative request from the current Pay As You Go price fields instead of extrapolating from a tiny prompt.
  • Check the observed-availability card and run your own timeout, retry, and fallback test before relying on the model.

Compare with Other Models

See how this model compares to others from the same provider.

Olmo 3.1 32B Think (Free)

Olmo 3.1 32B Think is a 32B-parameter reasoning model built for complex, multi-step logic and advanced instruction following. It improves on earlier Olmo versions with stronger performance on difficult tasks and more refined reasoning. Released by AI2 under Apache 2.0, it remains fully open, with transparent weights, code, and training process.

Context
65.5K
Input
$0/M
Output
$0/M

Olmo 3.1 32B Instruct

Olmo 3.1 32B Instruct is a 32B-parameter instruction-tuned model optimized for conversational AI and multi-turn dialogue. It focuses on strong instruction following and responsive chat behavior while maintaining solid reasoning and coding performance. Released by AI2 under Apache 2.0, it is fully open and transparent.

Context
65.5K
Input
$0.20/M
Output
$0.60/M

Olmo 3 7B Instruct

Olmo 3 7B Instruct is a 7B instruction-tuned version of the Olmo 3 base model, optimized for Q&A, instruction following, and natural conversation. Trained with high-quality data in an open pipeline, it performs well on everyday NLP tasks and is easy to adopt. Released by AI2 under Apache 2.0, it provides a transparent, community-friendly choice for instruction-driven applications.

Context
65.5K
Input
$0.10/M
Output
$0.20/M

Olmo 3 32B Think

Olmo 3.1 32B Think is a 32B-parameter reasoning model built for complex, multi-step logic and advanced instruction following. It improves on earlier Olmo versions with more refined reasoning and stronger performance on challenging evaluations, and — as part of AI2’s open initiative — it's fully transparent and released under Apache 2.0 with open weights, code, and training details.

Context
65.5K
Input
$0.30/M
Output
$0.45/M