Skip to content
GoogleChat

Gemma 4 31B (Free)

gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model, supporting text and image inputs with text outputs. It features a 256K token context window, configurable thinking/reasoning modes, native function calling, and broad multilingual support across 140+ languages. The model delivers strong performance in coding, reasoning, and document understanding, making it well suited for developer workflows, multilingual applications, and structured knowledge tasks.

Context
262.1K tokens
Endpoint
Get API KeyCompare

Pricing

Input$0 / 1M
Output$0 / 1M
Cache Write$0 / 1M
Cache Read$0 / 1M
Prompt cache writes and reads are included at no additional cost.

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint:
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="gemma-4-31b-it:free",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="gemma-4-31b-it:free",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common7 params
modelmessagesmax_tokenstemperaturetop_pstreamtools
Extended4 params
reasoning_effortstream_optionsthinkingextra_body

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

gemma-4-31b-it:free

Decision guidance

Fit the request contract before production.

Coding-agent gateway guide

Use this model when

  • The workload includes code generation or code-oriented text tasks listed in the current catalog record.
  • The required task matches the listed capabilities: text-to-text, text-to-code, translation.
  • The request depends on listed features such as vision, thinking, function-calling.
  • The input fits within the published 262.1K-token context record, with output and system overhead budgeted separately.

Check before production

  • Confirm the production client uses one of the listed request surfaces: /v1/chat/completions, /v1/responses, /v1/messages.
  • Estimate a representative request from the current Free price fields instead of extrapolating from a tiny prompt.
  • Check the observed-availability card and run your own timeout, retry, and fallback test before relying on the model.

Compare with Other Models

See how this model compares to others from the same provider.

Gemini 3.7 Flash

Gemini 3.7 Flash is Google's fast multimodal model designed for agentic workflows, coding, and complex multi-step reasoning. It combines responsive inference with reliable problem-solving capabilities, making it well suited for interactive and production-scale applications. Optimized for speed and dependable multi-step execution, Gemini 3.7 Flash is a strong choice for coding assistants, autonomous agents, and high-throughput workflows that require both low latency and capable reasoning.

Context
1M
Input
$0.375/M
Output
$1.88/M

Gemini 2.5 Flash Preview 05-20 (thinking)

Gemini 2.5 Flash Preview (May 2025) is Google's high-performance general model built for advanced reasoning, coding, math, and science. It includes built-in “thinking” features to deliver more accurate, context-aware answers.

Context
1.0M
Input
$0.075/M
Output
$1.75/M

Gemini 2.5 Flash Lite Preview 09-2025

Gemini 2.5 Flash-Lite is a lightweight, low-latency model focused on speed and cost efficiency. It generates tokens quickly and outperforms earlier Flash models on common benchmarks. “Thinking” (multi-pass reasoning) is off by default for maximum speed, but can be turned on through the Reasoning API when deeper reasoning is needed.

Context
1.0M
Input
$0.05/M
Output
$0.20/M

Gemini 2.5 Flash Lite Preview 06-17

Gemini 2.5 Flash-Lite is a smaller, low-latency model focused on speed and cost efficiency. It delivers faster generation and better benchmark performance than earlier Flash models. Thinking mode is off by default for maximum speed, but developers can enable it when they want deeper reasoning at a higher cost.

Context
1.0M
Input
$0.025/M
Output
$0.10/M