Compare models
Put up to 4 models beside each other — token prices, context windows, capabilities and provider, from the same catalogue the model pages read.
- S1Fish AudioRemove
- Grok 4.7SpaceXAIRemove
- Gemini 3.5 TranscribeGoogleRemove
- Voxtral Mini 3B 2507Mistral AIRemove
4 is the maximum. Remove one to add another.
| Attribute | S1s1 | Grok 4.7grok-4.7 | Gemini 3.5 Transcribegemini-3.5-transcribe | Voxtral Mini 3B 2507voxtral-mini-3b-2507 |
|---|---|---|---|---|
| Pricing | ||||
| Input | $0 / 1M | $1.60 / 1M | $0 / 1M | $0 / 1M |
| Output | $0 / 1M | $4.80 / 1M | $0 / 1M | $0 / 1M |
| Cache Write (5m) | Not applicable | $1.60 / 1M | Not applicable | Not applicable |
| Cache Write (1h) | Not applicable | $1.60 / 1M | Not applicable | Not applicable |
| Cache Read | Not applicable | $1.60 / 1M | Not applicable | Not applicable |
| Web Search | $0 / 1M | $0 / 1M | $0 / 1M | $0 / 1M |
| Context | ||||
| Max context | N/A | 500K | 98.3K | N/A |
| Max output | N/A | N/A | N/A | N/A |
| Capabilities | ||||
| Vision | No | Yes | Yes | No |
| Function Calling | No | Yes | Yes | No |
| JSON Mode | No | Yes | No | No |
| Streaming | No | Yes | No | No |
| Catalogue | ||||
| Provider | Fish Audio | SpaceXAI | Mistral AI | |
| Category | voice | chat | voice | voice |
| Charge type | Pay As You Go | Pay As You Go | Pay As You Go | Pay As You Go |
| Released | 2026-07-29 | — | 2026-09-25 | 2026-08-13 |
| Description | ||||
| Summary | Fish Audio S1 text-to-speech. Billed per UTF-8 byte of input text. | Grok 4.7 is SpaceXAI's flagship model for coding, agentic workflows, and professional knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering, self-verification, and long-context execution, while improving capabilities in document drafting, presentations, and other professional tasks. Trained with extended reinforcement learning focused on multi-hour problems, Grok 4.7 is optimized for sustained, complex task execution and natively supports the Grok Bot harness for conversational workflows. It also introduces an enhanced safeguard stack designed to combine strong jailbreak resistance with low refusal rates for legitimate technical work. Reported benchmark results use xhigh reasoning effort. | Google Gemini 3.5 Transcribe speech-to-text. Billed per input and output token. | Mistral AI Voxtral Mini (3B) speech-to-text. Billed per second of audio. |