Skip to content

Compare models

Put up to 4 models beside each other — token prices, context windows, capabilities and provider, from the same catalogue the model pages read.

  1. Transcribe 1 ProFish AudioRemove
  2. Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)GoogleRemove
  3. Perceptron Mk1.5PerceptronRemove
transcribe-1-pro vs gemini-3.1-flash-lite-image vs perceptron-mk1.5
AttributeTranscribe 1 Protranscribe-1-proNano Banana 2 Lite (Gemini 3.1 Flash Lite Image)gemini-3.1-flash-lite-imagePerceptron Mk1.5perceptron-mk1.5
Pricing
Input$0 / 1M—$0.15 / 1M
Output$0 / 1M—$1.50 / 1M
Cache Write (5m)Not applicableNot applicable$0.15 / 1M
Cache Write (1h)Not applicableNot applicable$0.15 / 1M
Cache ReadNot applicableNot applicable$0.15 / 1M
Web Search$0 / 1M—$0 / 1M
Request—$0.7754 / request—
Billing—Pay Per Request—
Context
Max contextN/A66K36.9K
Max outputN/AN/AN/A
Capabilities
VisionNoYesNo
Function CallingNoYesNo
JSON ModeNoYesNo
StreamingNoNoYes
Catalogue
ProviderFish AudioGooglePerceptron
Categoryvoiceimagechat
Charge typePay As You GoPay Per RequestPay As You Go
Released2026-09-24—2026-09-25
Description
SummaryFish Audio Transcribe 1 Pro speech-to-text with speaker labels; transcripts include speaker tags such as <|speaker:0|>. Billed per second of audio.Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest and most cost-efficient multimodal image generation model, designed for high-throughput visual workflows and real-time applications. It supports text-to-image generation, image editing, and multi-image composition through a unified API, while also producing text outputs alongside images. Delivering image generation in approximately 4 seconds, it combines fast inference with strong character consistency, precise editing, and real-world knowledge. The model generates 1K-resolution images across 14 aspect ratios and embeds an invisible SynthID watermark in all outputs. Optimized for the best balance of quality, speed, and cost, Nano Banana 2 Lite is ideal for prototyping, developer pipelines, and large-scale visual content generation.Perceptron Mk1.5 chat model that accepts audio input. Audio and text input are billed per token.