Veo 3.1 4K (Fast)
veo3.1-fast-4kVeo 3.1 is a state-of-the-art generative AI video model developed by Google DeepMind (part of the broader Gemini/Flow ecosystem). It builds on the earlier Veo models to make AI-generated video creation more realistic, expressive, and controllable.
- Context
- N/A
- Endpoint
Pricing
Pay Per RequestQuick Start
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="veo3.1-fast-4k", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="veo3.1-fast-4k",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsmodelpromptsizedurationresponse_formatseedwatermarkidcontentremixCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.
Gemini 3.8 Flash
Gemini 3.8 Flash is Google's most intelligent Flash-class model, delivering significant improvements over Gemini 3.7 Flash across software engineering, agentic workflows, and complex multi-step reasoning. Designed to combine strong capability with Flash-tier efficiency, it is well suited for coding assistants, autonomous agents, and high-throughput production workflows that require responsive performance without sacrificing reasoning quality.
- Context
- 1M
- Input
- $0.75/M
- Output
- $3.75/M
Gemini 3.7 Flash
Gemini 3.7 Flash is Google's fast multimodal model designed for agentic workflows, coding, and complex multi-step reasoning. It combines responsive inference with reliable problem-solving capabilities, making it well suited for interactive and production-scale applications. Optimized for speed and dependable multi-step execution, Gemini 3.7 Flash is a strong choice for coding assistants, autonomous agents, and high-throughput workflows that require both low latency and capable reasoning.
- Context
- 1M
- Input
- $0.375/M
- Output
- $1.88/M
Gemini 2.5 Flash Preview 05-20 (thinking)
Gemini 2.5 Flash Preview (May 2025) is Google's high-performance general model built for advanced reasoning, coding, math, and science. It includes built-in “thinking” features to deliver more accurate, context-aware answers.
- Context
- 1.0M
- Input
- $0.075/M
- Output
- $1.75/M
Gemini 2.5 Flash Lite Preview 09-2025
Gemini 2.5 Flash-Lite is a lightweight, low-latency model focused on speed and cost efficiency. It generates tokens quickly and outperforms earlier Flash models on common benchmarks. “Thinking” (multi-pass reasoning) is off by default for maximum speed, but can be turned on through the Reasoning API when deeper reasoning is needed.
- Context
- 1.0M
- Input
- $0.05/M
- Output
- $0.20/M