/v1/audio/transcriptionsSpeech to text| Rate | USD per unit |
|---|---|
| InputAudio | $0.000002 USD per token |
| OutputText | $0.000012 USD per token |
Cached input tokens are charged at the full configured customer input rate.
gemini-3.5-transcribeGoogle Gemini 3.5 Transcribe speech-to-text. Billed per input and output token.
Status information temporarily unavailable
Apertis cannot confirm the current service state. This is not a report that the model is down.
Base rates shown at GroupRatio=1. Your account's group multiplier may change the final charge.
/v1/audio/transcriptionsSpeech to text| Rate | USD per unit |
|---|---|
| InputAudio | $0.000002 USD per token |
| OutputText | $0.000012 USD per token |
Cached input tokens are charged at the full configured customer input rate.
Actual settled usage is deducted from your Apertis credit quota.
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="gemini-3.5-transcribe", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="gemini-3.5-transcribe",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )modelfilelanguagepromptresponse_formattemperaturetimestamp_granularitiesUse these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
See how this model compares to others from the same provider.
Gemini 2.5 Flash Preview (May 2025) is Google's high-performance general model built for advanced reasoning, coding, math, and science. It includes built-in “thinking” features to deliver more accurate, context-aware answers.
Gemini 2.5 Flash-Lite is a lightweight, low-latency model focused on speed and cost efficiency. It generates tokens quickly and outperforms earlier Flash models on common benchmarks. “Thinking” (multi-pass reasoning) is off by default for maximum speed, but can be turned on through the Reasoning API when deeper reasoning is needed.
Gemini 2.5 Flash-Lite is a smaller, low-latency model focused on speed and cost efficiency. It delivers faster generation and better benchmark performance than earlier Flash models. Thinking mode is off by default for maximum speed, but developers can enable it when they want deeper reasoning at a higher cost.
Gemini 2.5 Flash Preview (May 2025) is Google's high-performance general model built for advanced reasoning, coding, math, and science. It includes built-in “thinking” features to deliver more accurate, context-aware answers.