Models
| Provider | Model | Input | Output | Context |
|---|---|---|---|---|
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end professional work, designed for advanced analysis, software engineering, deep research, scientific tasks, and document creation. It is particularly strong in long-horizon agentic workflows, including tasks that require sustained reasoning, tool orchestration, and computer and browser use, making it well suited for complex autonomous workflows and production-grade knowledge work. ChatSep 4, 2026 | Input$10/1M tokens | Output$50/1M tokens | Context1M | |
GPT-6 Astra Pro uses the same underlying model as GPT-6 Astra, but runs with reasoning.mode set to pro for higher-quality responses on complex tasks. Optimized for deeper reasoning, greater accuracy, and more reliable multi-step execution, it is well suited for demanding coding, analysis, and agentic workflows where solution quality takes priority over speed and cost. ChatSep 4, 2026 | Input$10/1M tokens | Output$50/1M tokens | Context1M | |
Muse Spark 1.3 is Meta's multimodal reasoning model designed for long-running agentic, multi-agent, and coding workflows. It maintains context and information across extended tasks, enabling reliable execution in complex, multi-step environments. The model is optimized to resolve conflicting information, seek clarification or confirmation when necessary, and execute concisely, making it well suited for autonomous agents, collaborative multi-agent systems, and long-horizon software engineering workflows. ChatSep 2, 2026 | Input$1.25/1M tokens | Output$4.25/1M tokens | Context1M | |
Gemini 3.8 Flash is Google's most intelligent Flash-class model, delivering significant improvements over Gemini 3.7 Flash across software engineering, agentic workflows, and complex multi-step reasoning. Designed to combine strong capability with Flash-tier efficiency, it is well suited for coding assistants, autonomous agents, and high-throughput production workflows that require responsive performance without sacrificing reasoning quality. ChatSep 1, 2026 | Input$0.75/1M tokens | Output$3.75/1M tokens | Context1M | |
Claude Fable 5.1 is an upgraded version of Fable 5, delivering broad improvements with particularly strong gains in agentic coding, long-running workflows, and professional knowledge work. It excels at large code refactors, front-end and visual code generation, financial analysis, and complex analytical tasks. Compared with Fable 5, it also produces more concise plans and summaries while maintaining strong performance across extended tasks, making it a natural upgrade for existing Fable workflows and a strong option alongside Opus 5 for reasoning-intensive applications. ChatSep 1, 2026 | Input$10/1M tokens | Output$50/1M tokens | Context1M | |
Tencent Hy4 Preview is a Mixture-of-Experts (MoE) model from Tencent, featuring 770B total parameters with 49B activated per token. It is designed for coding agents, complex tool-driven workflows, and professional productivity tasks that require strong planning and reliable execution. Optimized for context continuity and sustained multi-step work, Hy4 Preview is well suited for long-horizon coding, agentic automation, tool orchestration, and complex real-world workflows. ChatAug 27, 2026 | Input$0.834/1M tokens | Output$2.50/1M tokens | Context1M | |
GLM-5.3-Flash is Z.AI's efficient native multimodal model, designed for coding and long-horizon agentic workflows. It combines strong multimodal capabilities with an architecture optimized for responsive, cost-efficient task execution. Built on a hybrid sparse and linear attention architecture, GLM-5.3-Flash maintains accurate long-context behavior while reducing computational overhead, making it well suited for coding agents, extended multi-step tasks, and scalable production workloads. ChatAug 25, 2026 | Input$0.075/1M tokens | Output$0.25/1M tokens | Context1M | |
GLM-5.3 is Z.ai's large-scale reasoning model designed for complex software engineering and long-horizon agentic workflows. It supports text input and output with a 1M-token context window, enabling sustained reasoning across large codebases and extended multi-step tasks. Building on GLM-5.2, it delivers stronger coding performance while improving the balance between capability and token efficiency, making it well suited for autonomous coding agents, large-scale engineering workflows, and complex task execution. ChatAug 18, 2026 | Input$1.40/1M tokens | Output$4.40/1M tokens | Context1M | |
Qwen3.8 27B is an open-weight dense vision-language model from Qwen, designed for coding, professional knowledge work, research, and multimodal interaction. It combines strong text and visual understanding with capabilities optimized for sustained, real-world agentic tasks. The model supports flexible thinking modes that can be enabled for deeper reasoning or disabled for faster execution, making it well suited for long-running agents, multimodal workflows, coding assistants, and cost-conscious self-hosted deployments. ChatAug 13, 2026 | Input$0.45/1M tokens | Output$3.20/1M tokens | Context262K | |
Gemini 3.7 Flash is Google's fast multimodal model designed for agentic workflows, coding, and complex multi-step reasoning. It combines responsive inference with reliable problem-solving capabilities, making it well suited for interactive and production-scale applications. Optimized for speed and dependable multi-step execution, Gemini 3.7 Flash is a strong choice for coding assistants, autonomous agents, and high-throughput workflows that require both low latency and capable reasoning. ChatAug 13, 2026 | Input$0.375/1M tokens | Output$1.88/1M tokens | Context1M | |
Qwen3.8 2.4T A95B is Qwen's open-weight sparse Mixture-of-Experts (MoE) model and the open-weight counterpart to Qwen3.8 Max. It features 2.4T total parameters with 95B activated per token, combining frontier-scale capacity with efficient sparse inference. Designed for coding, research, complex reasoning, and agentic workflows, the model is well suited for demanding long-horizon tasks and advanced autonomous systems while providing the flexibility and customization benefits of open weights. ChatAug 12, 2026 | Input$1.80/1M tokens | Output$5.40/1M tokens | Context262K | |
Grok 4.6 is SpaceXAI's smartest frontier model, delivering top-tier performance across coding, knowledge work, and STEM reasoning. It is designed for demanding technical and professional workloads that require strong problem solving, accurate instruction following, and reliable execution. Optimized for software engineering, scientific analysis, and complex knowledge tasks, Grok 4.6 is well suited for advanced coding, research, and agentic workflows where high capability and reasoning quality are critical. ChatAug 12, 2026 | Input$2/1M tokens | Output$6/1M tokens | Context500K | |
DeepSeek V4 Pro 0813 is DeepSeek's large-scale Mixture-of-Experts (MoE) model and the general availability (GA) release of DeepSeek V4 Pro. It is designed for high-capability workloads requiring advanced reasoning, coding, and agentic task execution. As the production-ready V4 Pro release, it is well suited for complex software engineering, long-horizon agent workflows, and demanding reasoning tasks where reliability and model capability are critical. ChatAug 11, 2026 | Input$0.435/1M tokens | Output$0.87/1M tokens | Context1M | |
NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference. Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key. ChatAug 10, 2026 | Input$0/1M tokens | Output$0/1M tokens | Context1M | |
NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference. Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key. ChatAug 10, 2026 | InputFreeIncluded | OutputFreeIncluded | Context1M | |
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It combines strong multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across 100+ languages. Designed for long-horizon agentic and coding workflows, Muse Glimmer 30B offers a practical balance of capability and deployment efficiency, making it well suited for local coding assistants, multimodal agents, and production workflows that require sustained autonomous execution. ChatAug 9, 2026 | Input$0.35/1M tokens | Output$1.50/1M tokens | Context131K | |
Muse Spark 1.2 is Meta's multimodal reasoning model designed for complex agentic and software engineering workflows. It supports text, image, video, audio, and PDF inputs with text output, and features a 1M-token context window for sustained reasoning across large, multi-stage tasks. Built for flexible multi-agent execution, Muse Spark 1.2 can serve as either a coordinating main agent or a parallel task-focused subagent. With configurable reasoning effort, structured outputs, parallel function calling, and broad coding-harness compatibility, it is well suited for multi-file refactoring, extended debugging, whole-repository generation, and long-horizon development workflows. ChatAug 4, 2026 | Input$1.25/1M tokens | Output$4.25/1M tokens | Context1M | |
Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to Qwen3.8 Max Preview. It is a multimodal reasoning model designed for complex tasks across reasoning, visual understanding, coding, and agentic workflows. As the production-ready top tier of the Qwen3.8 family, it is well suited for advanced problem solving, multimodal analysis, software engineering, and long-running tool-driven applications. ChatAug 2, 2026 | Input$2/1M tokens | Output$6/1M tokens | Context1M | |
Claude Opus 5 (Fast) is the fast version of Opus 5 model. ChatJul 27, 2026 | Input$10/1M tokens | Output$50/1M tokens | Context1M | |
Qwen3.7 Flash is Alibaba's vision-language reasoning model, designed for multimodal agents, visual coding, search, and computer interaction. It combines fast inference with strong visual understanding, including object recognition, spatial reasoning, and real-world scene perception. Optimized for interactive and agentic workflows, Qwen3.7 Flash is well suited for GUI understanding, visual question answering, multimodal search, and computer-use applications that require responsive reasoning across text and images. ChatJul 27, 2026 | Input$0.03/1M tokens | Output$0.13/1M tokens | Context1M | |
Claude Opus 5 is Anthropic's flagship model for advanced reasoning, coding, and long-horizon agentic workflows. It excels at end-to-end software engineering, code review, bug detection, visual analysis of charts and documents, complex office deliverables, and parallel subagent coordination. The model maintains reliable instruction following and tool use across extended tasks, while remaining effective at lower reasoning-effort settings for workloads that prioritize latency and token efficiency. ChatJul 24, 2026 | Input$5/1M tokens | Output$25/1M tokens | Context1M | |
Gemini 3.5 Flash-Lite is Google's high-efficiency model with enhanced agentic capabilities, optimized for fast, cost-effective inference. It is designed to handle focused tasks with low latency while maintaining strong reasoning and execution quality. Well suited for subagents in complex multi-agent systems, Gemini 3.5 Flash-Lite excels at executing specialized tasks within larger workflows, making it ideal for scalable agent orchestration and high-throughput production environments. ChatJul 20, 2026 | Input$0.3/1M tokens | Output$2.50/1M tokens | Context1M | |
Gemini 3.6 Flash is Google's high-efficiency model for coding, agentic workflows, and web and application development. It is optimized to produce polished, production-ready outputs with fewer unnecessary revisions, less hedging, and more direct task execution. By reducing both token usage and the number of model calls required to complete complex tasks, Gemini 3.6 Flash is well suited for high-throughput development, scalable agent systems, and cost-sensitive production workflows. ChatJul 20, 2026 | Input$1.50/1M tokens | Output$7.50/1M tokens | Context1M | |
Kimi K3 is Moonshot AI's 2.8T-parameter open-weight multimodal reasoning model, designed for complex coding, knowledge work, and long-horizon agentic workflows. It excels at repository-scale development, tool use, debugging, and iterative problem solving across text, images, logs, tests, and runtime feedback. Built with KDA and Attention Residuals for improved computational efficiency, Kimi K3 delivers strong performance on advanced engineering and multimodal reasoning tasks, making it well suited for autonomous coding agents and large-scale production workflows. ChatJul 15, 2026 | Input$3/1M tokens | Output$15/1M tokens | Context1M | |
Inkling is a general-purpose multimodal autoregressive transformer model from Thinking Machines Lab, supporting text, image, and audio inputs with text output. It is designed for a wide range of applications, including reasoning, coding, multilingual understanding, and multimodal interaction. Optimized for agentic workflows, tool use, coding assistants, chatbots, retrieval-augmented generation (RAG), and instruction following, Inkling provides a versatile foundation for production AI systems requiring robust language and multimodal capabilities. ChatJul 14, 2026 | InputFreeIncluded | OutputFreeIncluded | Context1M | |
GPT-5.6 Terra Pro uses the same underlying model as GPT-5.6 Terra, but runs with reasoning.mode set to pro to deliver higher-quality responses on complex tasks. Optimized for deeper reasoning and greater reliability, it is well suited for advanced coding, multi-step reasoning, and agentic workflows where improved accuracy and solution quality are more important than maximizing speed or minimizing cost. ChatJul 8, 2026 | Input$2/1M tokens | Output$12/1M tokens | Context1M | |
GPT-5.6 Luna Pro uses the same underlying model as GPT-5.6 Luna, but runs with reasoning.mode set to pro to deliver higher-quality responses on complex tasks. It is optimized for deeper reasoning, advanced coding, and multi-step agentic workflows, offering improved accuracy and solution quality while retaining the efficiency and scalability of the Luna tier. ChatJul 8, 2026 | Input$0.2/1M tokens | Output$1.20/1M tokens | Context1M | |
GPT-5.6 Sol Pro uses the same underlying model as GPT-5.6 Sol, but runs with reasoning.mode set to pro for higher-quality responses on complex tasks. Optimized for deeper reasoning and more reliable execution, it is particularly well suited for advanced coding, long-horizon problem solving, and agentic workflows where accuracy and solution quality take priority over speed and cost. ChatJul 8, 2026 | Input$5/1M tokens | Output$30/1M tokens | Context1M | |
GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is designed for everyday coding, reasoning, and agentic workflows, delivering strong performance while balancing capability and cost. Offering near-flagship quality at approximately half the cost of Sol, GPT-5.6 Terra is well suited for production applications that require reliable reasoning, software development, and scalable agent execution. ChatJul 8, 2026 | Input$2/1M tokens | Output$12/1M tokens | Context1M | |
GPT-5.6 Luna is the fast, cost-efficient model in OpenAI's GPT-5.6 series, optimized for high-volume, latency-sensitive workloads. It delivers capable reasoning at an affordable price point, making it ideal for chat applications, classification, and lightweight agentic workflows. Designed for scalable production deployments, GPT-5.6 Luna balances speed, cost, and reliability, providing efficient performance for real-time applications and large-scale automation tasks. ChatJul 8, 2026 | Input$0.2/1M tokens | Output$1.20/1M tokens | Context1M | |
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series, designed for complex reasoning, coding, and agentic workflows. It delivers strong performance on multi-step software engineering tasks, command-line workflows, and long-horizon problem solving, making it well suited for advanced development and autonomous execution. Optimized for high-reliability reasoning and end-to-end task completion, GPT-5.6 Sol excels in coding, tool-driven automation, and large-scale engineering workflows that require sustained context and precise execution. ChatJul 8, 2026 | Input$5/1M tokens | Output$30/1M tokens | Context1M | |
Grok 4.5 is SpaceXAI’s flagship frontier model, delivering top-tier performance across coding, knowledge work, and STEM reasoning. Designed for demanding professional and technical workloads, it combines strong reasoning, accurate instruction following, and robust problem-solving capabilities. Optimized for software engineering, scientific analysis, and complex knowledge tasks, Grok 4.5 is well suited for advanced coding, research, and agent-driven workflows that require high accuracy and reliable long-horizon reasoning. ChatJul 7, 2026 | Input$2/1M tokens | Output$6/1M tokens | Context500K | |
Hy3 is Tencent's 295B-parameter Mixture-of-Experts (MoE) model, activating 21B parameters per token across 192 experts, and designed for reasoning, agentic workflows, and production-scale applications. It supports a 256K-token context window and configurable reasoning modes, including no-think, low, and high reasoning effort to balance speed and problem-solving depth. Optimized for long-horizon tasks, coding, and tool-driven execution, Hy3 delivers strong performance in multi-turn reasoning, constraint tracking, and stable tool calling. With an emphasis on grounded responses and reduced hallucinations, it is well suited for software development, document processing, financial analysis, game development, and enterprise agent workflows. ChatJul 5, 2026 | InputFreeIncluded | OutputFreeIncluded | Context262K | |
Hy3 is Tencent's 295B-parameter Mixture-of-Experts (MoE) model, activating 21B parameters per token across 192 experts, and designed for reasoning, agentic workflows, and production-scale applications. It supports a 256K-token context window and configurable reasoning modes, including no-think, low, and high reasoning effort to balance speed and problem-solving depth. Optimized for long-horizon tasks, coding, and tool-driven execution, Hy3 delivers strong performance in multi-turn reasoning, constraint tracking, and stable tool calling. With an emphasis on grounded responses and reduced hallucinations, it is well suited for software development, document processing, financial analysis, game development, and enterprise agent workflows. ChatJul 4, 2026 | Input$0.14/1M tokens | Output$0.58/1M tokens | Context262K | |
Sonnet 5 is Anthropic's most capable Sonnet-class model, delivering frontier-level performance across coding, agentic workflows, and professional knowledge tasks. It supports text, image, and file inputs, features a 1M-token context window, and offers adaptive thinking with configurable reasoning levels (low, medium, high, max, and x-high) to balance speed, cost, and reasoning depth. Optimized for complex coding, long-horizon agent execution, and professional workflows, Sonnet 5 combines strong reasoning, robust instruction following, and enhanced safety features, including an updated tokenizer and real-time cyber safeguards for high-risk dual-use scenarios. ChatJun 30, 2026 | Input$2/1M tokens | Output$10/1M tokens | Context1M | |
Fugu Ultra is the high-performance model in Sakana AI's Fugu family, built as a learned multi-agent orchestration system rather than a single monolithic model. It intelligently routes tasks across a pool of underlying models and can recursively invoke itself to solve complex problems more effectively. Optimized for multi-step reasoning, coding, and agentic workflows, Fugu Ultra supports configurable reasoning effort, native tool calling, and built-in web search. Its orchestration-based design makes it well suited for advanced autonomous agents and complex task execution requiring adaptive model coordination. ChatJun 23, 2026 | Input$5/1M tokens | Output$30/1M tokens | Context1M | |
GLM-5.2 is Z.AI's flagship model for long-horizon task execution, designed to handle complex, project-scale workflows with high reliability. Featuring a 1M-token context window, it can maintain and reason over extensive engineering context, enabling consistent execution across large, multi-stage tasks. Optimized for end-to-end software development, GLM-5.2 follows engineering standards reliably and can manage the full workflow from requirements analysis and implementation to testing and multi-platform deployment, making it well suited for advanced coding agents and large-scale autonomous engineering projects. ChatJun 16, 2026 | Input$1.40/1M tokens | Output$4.40/1M tokens | Context1M | |
Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, designed for long-horizon software engineering and agentic development workflows. Built on a native multimodal Mixture-of-Experts (MoE) architecture, it supports text, image, and video inputs and operates exclusively in thinking mode, preserving reasoning across multi-turn interactions. With approximately 1T total parameters and 32B activated per token, plus a 256K-token context window, K2.7 Code excels at end-to-end programming tasks, agentic task decomposition, repository-scale reasoning, and extended coding conversations, making it well suited for advanced coding agents and long-context development workflows. ChatJun 12, 2026 | Input$0.95/1M tokens | Output$4/1M tokens | Context262K | |
Claude Fable 5 is Anthropic's Mythos-class model, designed for autonomous knowledge work, coding, and long-running agentic workflows. It supports text, image, and file inputs with text output, includes reasoning capabilities, and features a 1M-token context window for handling complex, high-context tasks. Optimized for asynchronous and long-horizon execution, Claude Fable 5 excels at end-to-end tasks that would typically require hours, days, or weeks of human effort. It combines strong reasoning, autonomous verification and self-correction loops, and robust safeguards, making it well suited for complex research, software engineering, and large-scale knowledge work. ChatJun 8, 2026 | Input$10/1M tokens | Output$50/1M tokens | Context1M | |
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, designed for content moderation, safety classification, and AI policy enforcement. Supporting text and image inputs with text output, it evaluates both user prompts and model responses, providing safe/unsafe classifications, safety category labels, and optional reasoning traces. Fine-tuned from Gemma-3-4B and supporting 12 languages with a 128K-token context window, the model is well suited for prompt moderation, response filtering, content classification, and enterprise safety pipelines. As part of the NVIDIA Nemotron family, it offers a configurable reasoning mode and integrates easily into agentic AI systems requiring robust guardrails and compliance controls. ChatJun 3, 2026 | InputFreeIncluded | OutputFreeIncluded | Context128K | |
NVIDIA Nemotron 3 Ultra is an open frontier reasoning and orchestration model featuring a 550B-parameter Mixture-of-Experts (MoE) architecture with 55B active parameters per token. Built on a hybrid Transformer–Mamba design, it supports text input and output with a 1M-token context window, enabling large-scale reasoning and long-horizon task execution. Optimized for agent orchestration, coding agents, deep research, and complex enterprise workflows, the model excels at multi-step reasoning, planning, and sustained execution. With high-throughput inference designed for large-scale agent pipelines, Nemotron 3 Ultra serves as a powerful foundation for advanced agentic AI systems. ChatJun 3, 2026 | InputFreeIncluded | OutputFreeIncluded | Context1M | |
NVIDIA Nemotron 3 Ultra is an open frontier reasoning and orchestration model featuring a 550B-parameter Mixture-of-Experts (MoE) architecture with 55B active parameters per token. Built on a hybrid Transformer–Mamba design, it supports text input and output with a 1M-token context window, enabling large-scale reasoning and long-horizon task execution. Optimized for agent orchestration, coding agents, deep research, and complex enterprise workflows, the model excels at multi-step reasoning, planning, and sustained execution. With high-throughput inference designed for large-scale agent pipelines, Nemotron 3 Ultra serves as a powerful foundation for advanced agentic AI systems. ChatJun 3, 2026 | Input$0.5/1M tokens | Output$2.50/1M tokens | Context1M | |
Qwen3.7-Plus is a cost-effective multimodal model in Alibaba's Qwen3.7 series, supporting text and image inputs with text output. It combines the series' strong language capabilities with significantly enhanced vision-language understanding, while retaining full-stack agent-level intelligence for coding, tool use, and productivity workflows. Its standout capability is multimodal interactive agency—the ability to perceive real-world scenes, understand screens and graphical interfaces, generate code from visual references, and perform end-to-end navigation within applications. This makes Qwen3.7-Plus well suited for GUI automation, visual coding, productivity agents, and multimodal task execution. ChatJun 2, 2026 | Input$0.4/1M tokens | Output$1.60/1M tokens | Context1M | |
MiniMax-M3 is a multimodal foundation model from MiniMax, supporting text, image, and video inputs with text output and a 1M-token context window. It is designed for long-horizon agentic workflows, coding, and tool-driven task execution, enabling sustained reasoning across complex tasks. Built on MiniMax Sparse Attention (MSA), the model dramatically improves long-context efficiency by replacing full attention with KV-block selection, reducing compute costs at 1M-token contexts while maintaining strong performance. Trained as a native multimodal model and optimized for multi-turn, production-style collaboration, MiniMax-M3 excels at extended, multi-step workflows rather than single-turn interactions. ChatMay 31, 2026 | Input$0.3/1M tokens | Output$1.20/1M tokens | Context1M | |
This is the fast version of Opus 4.8 ChatMay 27, 2026 | Input$10/1M tokens | Output$50/1M tokens | Context1M | |
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family, designed for highly autonomous agents, long-horizon workflows, and advanced knowledge work. It supports text, image, and file inputs with text output, includes reasoning capabilities, and features a 1M-token context window for maintaining coherence across extended tasks and sessions. The model excels at multi-step reasoning, complex coding, and end-to-end project orchestration, including large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond software engineering, it is highly effective for document drafting, presentation creation, data analysis, and memory-driven workflows, delivering consistent quality across very long outputs and complex projects. ChatMay 27, 2026 | Input$4/1M tokens | Output$20/1M tokens | Context1M | |
Grok Build 0.1 is xAI's fast coding model designed specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding agents, tool use, and multi-step development tasks. Powering the Grok Build CLI, the model features a 256K token context window with effectively no text output limit, making it well suited for long-horizon coding, automation, and continuous development workflows. Currently available in early access. ChatMay 20, 2026 | Input$1/1M tokens | Output$2/1M tokens | Context256K | |
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series, designed for agent-centric workloads with strong performance in coding, productivity, and long-horizon autonomous execution. It supports text input and output and delivers notable improvements in coding and agentic capabilities over previous Qwen generations. Optimized for real-world workflows, the model also supports explicit prompt caching for efficient reuse of repeated context, making it well suited for scalable development, office automation, and advanced agent systems. ChatMay 20, 2026 | Input$2.50/1M tokens | Output$7.50/1M tokens | Context1M | |
Gemini 3.5 Flash is Google's high-efficiency multimodal model, delivering near-Pro level performance in coding and reasoning at Flash-tier speed and cost. It supports text, image, video, audio, and PDF inputs, making it well suited for diverse multimodal workflows. Optimized for coding proficiency and parallel agentic execution, the model defaults to medium thinking effort for faster, cost-efficient responses while supporting configurable thinking levels (minimal, low, medium, high) for fine-grained cost–performance control. ChatMay 18, 2026 | Input$1.50/1M tokens | Output$9/1M tokens | Context1M | |
This is the fast version of Opus 4.7 ChatMay 12, 2026 | Input$30/1M tokens | Output$150/1M tokens | Context1M |