本文へ移動

モデル

Sort
比較
プロバイダーモデル入力出力コンテキスト

Gemini 3.6 Flash is Google's high-efficiency model for coding, agentic workflows, and web and application development. It is optimized to produce polished, production-ready outputs with fewer unnecessary revisions, less hedging, and more direct task execution. By reducing both token usage and the number of model calls required to complete complex tasks, Gemini 3.6 Flash is well suited for high-throughput development, scalable agent systems, and cost-sensitive production workflows.

Chat2026年7月20日
入力$1.50/1M tokens出力$7.50/1M tokens
コンテキスト1M

Gemini 3.5 Flash-Lite is Google's high-efficiency model with enhanced agentic capabilities, optimized for fast, cost-effective inference. It is designed to handle focused tasks with low latency while maintaining strong reasoning and execution quality. Well suited for subagents in complex multi-agent systems, Gemini 3.5 Flash-Lite excels at executing specialized tasks within larger workflows, making it ideal for scalable agent orchestration and high-throughput production environments.

Chat2026年7月20日
入力$0.3/1M tokens出力$2.50/1M tokens
コンテキスト1M

Kimi K3 is Moonshot AI's 2.8T-parameter open-weight multimodal reasoning model, designed for complex coding, knowledge work, and long-horizon agentic workflows. It excels at repository-scale development, tool use, debugging, and iterative problem solving across text, images, logs, tests, and runtime feedback. Built with KDA and Attention Residuals for improved computational efficiency, Kimi K3 delivers strong performance on advanced engineering and multimodal reasoning tasks, making it well suited for autonomous coding agents and large-scale production workflows.

Chat2026年7月15日
入力$3/1M tokens出力$15/1M tokens
コンテキスト1M

Inkling is a general-purpose multimodal autoregressive transformer model from Thinking Machines Lab, supporting text, image, and audio inputs with text output. It is designed for a wide range of applications, including reasoning, coding, multilingual understanding, and multimodal interaction. Optimized for agentic workflows, tool use, coding assistants, chatbots, retrieval-augmented generation (RAG), and instruction following, Inkling provides a versatile foundation for production AI systems requiring robust language and multimodal capabilities.

Chat2026年7月14日
入力FreeIncluded出力FreeIncluded
コンテキスト1M

GPT-5.6 Luna Pro uses the same underlying model as GPT-5.6 Luna, but runs with reasoning.mode set to pro to deliver higher-quality responses on complex tasks. It is optimized for deeper reasoning, advanced coding, and multi-step agentic workflows, offering improved accuracy and solution quality while retaining the efficiency and scalability of the Luna tier.

Chat2026年7月8日
入力$1/1M tokens出力$6/1M tokens
コンテキスト1M

GPT-5.6 Luna is the fast, cost-efficient model in OpenAI's GPT-5.6 series, optimized for high-volume, latency-sensitive workloads. It delivers capable reasoning at an affordable price point, making it ideal for chat applications, classification, and lightweight agentic workflows. Designed for scalable production deployments, GPT-5.6 Luna balances speed, cost, and reliability, providing efficient performance for real-time applications and large-scale automation tasks.

Chat2026年7月8日
入力$1/1M tokens出力$6/1M tokens
コンテキスト1M

GPT-5.6 Terra Pro uses the same underlying model as GPT-5.6 Terra, but runs with reasoning.mode set to pro to deliver higher-quality responses on complex tasks. Optimized for deeper reasoning and greater reliability, it is well suited for advanced coding, multi-step reasoning, and agentic workflows where improved accuracy and solution quality are more important than maximizing speed or minimizing cost.

Chat2026年7月8日
入力$2.50/1M tokens出力$15/1M tokens
コンテキスト1M

GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is designed for everyday coding, reasoning, and agentic workflows, delivering strong performance while balancing capability and cost. Offering near-flagship quality at approximately half the cost of Sol, GPT-5.6 Terra is well suited for production applications that require reliable reasoning, software development, and scalable agent execution.

Chat2026年7月8日
入力$2.50/1M tokens出力$15/1M tokens
コンテキスト1M

GPT-5.6 Sol Pro uses the same underlying model as GPT-5.6 Sol, but runs with reasoning.mode set to pro for higher-quality responses on complex tasks. Optimized for deeper reasoning and more reliable execution, it is particularly well suited for advanced coding, long-horizon problem solving, and agentic workflows where accuracy and solution quality take priority over speed and cost.

Chat2026年7月8日
入力$5/1M tokens出力$30/1M tokens
コンテキスト1M

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series, designed for complex reasoning, coding, and agentic workflows. It delivers strong performance on multi-step software engineering tasks, command-line workflows, and long-horizon problem solving, making it well suited for advanced development and autonomous execution. Optimized for high-reliability reasoning and end-to-end task completion, GPT-5.6 Sol excels in coding, tool-driven automation, and large-scale engineering workflows that require sustained context and precise execution.

Chat2026年7月8日
入力$5/1M tokens出力$30/1M tokens
コンテキスト1M

Grok 4.5 is SpaceXAI’s flagship frontier model, delivering top-tier performance across coding, knowledge work, and STEM reasoning. Designed for demanding professional and technical workloads, it combines strong reasoning, accurate instruction following, and robust problem-solving capabilities. Optimized for software engineering, scientific analysis, and complex knowledge tasks, Grok 4.5 is well suited for advanced coding, research, and agent-driven workflows that require high accuracy and reliable long-horizon reasoning.

Chat2026年7月7日
入力$2/1M tokens出力$6/1M tokens
コンテキスト500K

Hy3 is Tencent's 295B-parameter Mixture-of-Experts (MoE) model, activating 21B parameters per token across 192 experts, and designed for reasoning, agentic workflows, and production-scale applications. It supports a 256K-token context window and configurable reasoning modes, including no-think, low, and high reasoning effort to balance speed and problem-solving depth. Optimized for long-horizon tasks, coding, and tool-driven execution, Hy3 delivers strong performance in multi-turn reasoning, constraint tracking, and stable tool calling. With an emphasis on grounded responses and reduced hallucinations, it is well suited for software development, document processing, financial analysis, game development, and enterprise agent workflows.

Chat2026年7月5日
入力FreeIncluded出力FreeIncluded
コンテキスト262K

Hy3 is Tencent's 295B-parameter Mixture-of-Experts (MoE) model, activating 21B parameters per token across 192 experts, and designed for reasoning, agentic workflows, and production-scale applications. It supports a 256K-token context window and configurable reasoning modes, including no-think, low, and high reasoning effort to balance speed and problem-solving depth. Optimized for long-horizon tasks, coding, and tool-driven execution, Hy3 delivers strong performance in multi-turn reasoning, constraint tracking, and stable tool calling. With an emphasis on grounded responses and reduced hallucinations, it is well suited for software development, document processing, financial analysis, game development, and enterprise agent workflows.

Chat2026年7月4日
入力$0.14/1M tokens出力$0.58/1M tokens
コンテキスト262K

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest and most cost-efficient multimodal image generation model, designed for high-throughput visual workflows and real-time applications. It supports text-to-image generation, image editing, and multi-image composition through a unified API, while also producing text outputs alongside images. Delivering image generation in approximately 4 seconds, it combines fast inference with strong character consistency, precise editing, and real-world knowledge. The model generates 1K-resolution images across 14 aspect ratios and embeds an invisible SynthID watermark in all outputs. Optimized for the best balance of quality, speed, and cost, Nano Banana 2 Lite is ideal for prototyping, developer pipelines, and large-scale visual content generation.

Image2026年6月30日
入力$0/1M tokens出力$0/1M tokens
コンテキスト66K

Sonnet 5 is Anthropic's most capable Sonnet-class model, delivering frontier-level performance across coding, agentic workflows, and professional knowledge tasks. It supports text, image, and file inputs, features a 1M-token context window, and offers adaptive thinking with configurable reasoning levels (low, medium, high, max, and x-high) to balance speed, cost, and reasoning depth. Optimized for complex coding, long-horizon agent execution, and professional workflows, Sonnet 5 combines strong reasoning, robust instruction following, and enhanced safety features, including an updated tokenizer and real-time cyber safeguards for high-risk dual-use scenarios.

Chat2026年6月30日
入力$2/1M tokens出力$10/1M tokens
コンテキスト1M

Fugu Ultra is the high-performance model in Sakana AI's Fugu family, built as a learned multi-agent orchestration system rather than a single monolithic model. It intelligently routes tasks across a pool of underlying models and can recursively invoke itself to solve complex problems more effectively. Optimized for multi-step reasoning, coding, and agentic workflows, Fugu Ultra supports configurable reasoning effort, native tool calling, and built-in web search. Its orchestration-based design makes it well suited for advanced autonomous agents and complex task execution requiring adaptive model coordination.

Chat2026年6月23日
入力$5/1M tokens出力$30/1M tokens
コンテキスト1M

GLM-5.2 is Z.AI's flagship model for long-horizon task execution, designed to handle complex, project-scale workflows with high reliability. Featuring a 1M-token context window, it can maintain and reason over extensive engineering context, enabling consistent execution across large, multi-stage tasks. Optimized for end-to-end software development, GLM-5.2 follows engineering standards reliably and can manage the full workflow from requirements analysis and implementation to testing and multi-platform deployment, making it well suited for advanced coding agents and large-scale autonomous engineering projects.

Chat2026年6月16日
入力$1.40/1M tokens出力$4.40/1M tokens
コンテキスト1M

Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, designed for long-horizon software engineering and agentic development workflows. Built on a native multimodal Mixture-of-Experts (MoE) architecture, it supports text, image, and video inputs and operates exclusively in thinking mode, preserving reasoning across multi-turn interactions. With approximately 1T total parameters and 32B activated per token, plus a 256K-token context window, K2.7 Code excels at end-to-end programming tasks, agentic task decomposition, repository-scale reasoning, and extended coding conversations, making it well suited for advanced coding agents and long-context development workflows.

Chat2026年6月12日
入力$0.95/1M tokens出力$4/1M tokens
コンテキスト262K

Claude Fable 5 is Anthropic's Mythos-class model, designed for autonomous knowledge work, coding, and long-running agentic workflows. It supports text, image, and file inputs with text output, includes reasoning capabilities, and features a 1M-token context window for handling complex, high-context tasks. Optimized for asynchronous and long-horizon execution, Claude Fable 5 excels at end-to-end tasks that would typically require hours, days, or weeks of human effort. It combines strong reasoning, autonomous verification and self-correction loops, and robust safeguards, making it well suited for complex research, software engineering, and large-scale knowledge work.

Chat2026年6月8日
入力$10/1M tokens出力$50/1M tokens
コンテキスト1M

NVIDIA Nemotron 3 Ultra is an open frontier reasoning and orchestration model featuring a 550B-parameter Mixture-of-Experts (MoE) architecture with 55B active parameters per token. Built on a hybrid Transformer–Mamba design, it supports text input and output with a 1M-token context window, enabling large-scale reasoning and long-horizon task execution. Optimized for agent orchestration, coding agents, deep research, and complex enterprise workflows, the model excels at multi-step reasoning, planning, and sustained execution. With high-throughput inference designed for large-scale agent pipelines, Nemotron 3 Ultra serves as a powerful foundation for advanced agentic AI systems.

Chat2026年6月3日
入力FreeIncluded出力FreeIncluded
コンテキスト1M

NVIDIA Nemotron 3 Ultra is an open frontier reasoning and orchestration model featuring a 550B-parameter Mixture-of-Experts (MoE) architecture with 55B active parameters per token. Built on a hybrid Transformer–Mamba design, it supports text input and output with a 1M-token context window, enabling large-scale reasoning and long-horizon task execution. Optimized for agent orchestration, coding agents, deep research, and complex enterprise workflows, the model excels at multi-step reasoning, planning, and sustained execution. With high-throughput inference designed for large-scale agent pipelines, Nemotron 3 Ultra serves as a powerful foundation for advanced agentic AI systems.

Chat2026年6月3日
入力$0.5/1M tokens出力$2.50/1M tokens
コンテキスト1M

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, designed for content moderation, safety classification, and AI policy enforcement. Supporting text and image inputs with text output, it evaluates both user prompts and model responses, providing safe/unsafe classifications, safety category labels, and optional reasoning traces. Fine-tuned from Gemma-3-4B and supporting 12 languages with a 128K-token context window, the model is well suited for prompt moderation, response filtering, content classification, and enterprise safety pipelines. As part of the NVIDIA Nemotron family, it offers a configurable reasoning mode and integrates easily into agentic AI systems requiring robust guardrails and compliance controls.

Chat2026年6月3日
入力FreeIncluded出力FreeIncluded
コンテキスト128K

Qwen3.7-Plus is a cost-effective multimodal model in Alibaba's Qwen3.7 series, supporting text and image inputs with text output. It combines the series' strong language capabilities with significantly enhanced vision-language understanding, while retaining full-stack agent-level intelligence for coding, tool use, and productivity workflows. Its standout capability is multimodal interactive agency—the ability to perceive real-world scenes, understand screens and graphical interfaces, generate code from visual references, and perform end-to-end navigation within applications. This makes Qwen3.7-Plus well suited for GUI automation, visual coding, productivity agents, and multimodal task execution.

Chat2026年6月2日
入力$0.4/1M tokens出力$1.60/1M tokens
コンテキスト1M

MiniMax-M3 is a multimodal foundation model from MiniMax, supporting text, image, and video inputs with text output and a 1M-token context window. It is designed for long-horizon agentic workflows, coding, and tool-driven task execution, enabling sustained reasoning across complex tasks. Built on MiniMax Sparse Attention (MSA), the model dramatically improves long-context efficiency by replacing full attention with KV-block selection, reducing compute costs at 1M-token contexts while maintaining strong performance. Trained as a native multimodal model and optimized for multi-turn, production-style collaboration, MiniMax-M3 excels at extended, multi-step workflows rather than single-turn interactions.

Chat2026年5月31日
入力$0.3/1M tokens出力$1.20/1M tokens
コンテキスト1M

This is the fast version of Opus 4.8

Chat2026年5月27日
入力$10/1M tokens出力$50/1M tokens
コンテキスト1M

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family, designed for highly autonomous agents, long-horizon workflows, and advanced knowledge work. It supports text, image, and file inputs with text output, includes reasoning capabilities, and features a 1M-token context window for maintaining coherence across extended tasks and sessions. The model excels at multi-step reasoning, complex coding, and end-to-end project orchestration, including large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond software engineering, it is highly effective for document drafting, presentation creation, data analysis, and memory-driven workflows, delivering consistent quality across very long outputs and complex projects.

Chat2026年5月27日
入力$4/1M tokens出力$20/1M tokens
コンテキスト1M

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series, designed for agent-centric workloads with strong performance in coding, productivity, and long-horizon autonomous execution. It supports text input and output and delivers notable improvements in coding and agentic capabilities over previous Qwen generations. Optimized for real-world workflows, the model also supports explicit prompt caching for efficient reuse of repeated context, making it well suited for scalable development, office automation, and advanced agent systems.

Chat2026年5月20日
入力$2.50/1M tokens出力$7.50/1M tokens
コンテキスト1M

Grok Build 0.1 is xAI's fast coding model designed specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding agents, tool use, and multi-step development tasks. Powering the Grok Build CLI, the model features a 256K token context window with effectively no text output limit, making it well suited for long-horizon coding, automation, and continuous development workflows. Currently available in early access.

Chat2026年5月20日
入力$1/1M tokens出力$2/1M tokens
コンテキスト256K

Gemini 3.5 Flash is Google's high-efficiency multimodal model, delivering near-Pro level performance in coding and reasoning at Flash-tier speed and cost. It supports text, image, video, audio, and PDF inputs, making it well suited for diverse multimodal workflows. Optimized for coding proficiency and parallel agentic execution, the model defaults to medium thinking effort for faster, cost-efficient responses while supporting configurable thinking levels (minimal, low, medium, high) for fine-grained cost–performance control.

Chat2026年5月18日
入力$1.50/1M tokens出力$9/1M tokens
コンテキスト1M

This is the fast version of Opus 4.7

Chat2026年5月12日
入力$30/1M tokens出力$150/1M tokens
コンテキスト1M

Gemini 3.1 Flash TTS Preview is Google's next-generation text-to-speech model, delivering a major upgrade over Gemini 2.5 Flash TTS. It converts text into natural audio across 70+ languages, with significantly expanded language coverage and improved quality. The model introduces 200+ inline audio control tags (e.g., [whispers], [laughs], [excited]) for fine-grained control over emotion, tone, and pacing, along with support for two speakers with independent voice and style settings. It outputs 24 kHz / 16-bit PCM audio, includes SynthID watermarking, and supports a 32K token context window. Designed for expressive and controllable voice generation, it is well suited for dialogue systems, storytelling, character-driven content, and advanced audio production workflows.

Voice2026年4月30日
入力$27.50/1M tokens出力$0/1M tokens
コンテキスト8K

GPT-4o Mini TTS is OpenAI's cost-efficient text-to-speech model, designed to convert text into natural-sounding audio output. It supports a variety of voices and tones, enabling flexible and expressive speech generation. Optimized for scalability and low cost, it is well suited for real-time voice applications, content narration, and high-volume audio generation workflows.

Voice2026年4月30日
入力$0.3/1M tokens出力$0/1M tokens
コンテキスト4K

GPT-4o Mini Transcribe is a smaller, cost-efficient speech-to-text model built on GPT-4o Mini's audio capabilities. It is designed for high-volume transcription workloads, delivering reliable performance with lower cost and latency. Priced per token (input and output), it provides transparent, fine-grained billing, making it well suited for scalable transcription pipelines, real-time applications, and cost-sensitive deployments.

Voice2026年4月30日
入力$0.625/1M tokens出力$0.625/1M tokens
コンテキスト128K

GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o's audio capabilities. It delivers accurate transcription with strong language understanding, making it suitable for a wide range of audio processing tasks. Priced per token (input and output), it offers transparent, fine-grained billing, making it well suited for workflows that require scalable transcription, integration with LLM pipelines, and cost-aware processing.

Voice2026年4月30日
入力$1.25/1M tokens出力$0/1M tokens
コンテキスト128K

Whisper Large V3 Turbo is an optimized version of OpenAI's Whisper Large V3 speech recognition model, designed for high-speed and cost-efficient transcription. It supports 99+ languages and accepts common audio formats including mp3, mp4, wav, webm, flac, and ogg. With a ~12% word error rate and real-time speed factors up to 216×, it delivers fast, scalable performance for latency-sensitive and high-throughput transcription workloads, making it ideal for real-time and large-scale speech processing applications.

Voice2026年4月30日
入力$3.33/1M tokens出力$0/1M tokens
-

Whisper Large V3 is OpenAI's advanced open-source automatic speech recognition (ASR) model, supporting both audio transcription and translation across 99+ languages. It accepts common audio formats including mp3, mp4, wav, webm, flac, and ogg, and delivers strong performance in noisy, real-world conditions. With 1.55B parameters and a low 10.3% word error rate, it provides accurate, multilingual transcription with support for word- and segment-level timestamps, making it well suited for high-quality, noise-robust speech processing applications.

Voice2026年4月30日
入力$9.25/1M tokens出力$0/1M tokens
-

Whisper (whisper-1) is OpenAI's open-source automatic speech recognition (ASR) model, designed for audio transcription and translation. It supports 50+ languages and processes audio files up to 25 MB, accepting formats such as mp3, mp4, wav, and webm. Optimized for reliable speech-to-text conversion across diverse audio inputs, Whisper is priced per minute of audio, billed to the nearest second, making it well suited for transcription, localization, and voice-driven applications.

Voice2026年4月30日
入力$75/1M tokens出力$75/1M tokens
-

Mistral Medium 3.5 is a 128B dense instruction-following model from Mistral AI, supporting text and image inputs with text output. It is designed for agentic workflows, coding, and complex multi-step reasoning, with strong reliability in multi-tool orchestration and long-horizon tasks. The model features a 256K token context window, configurable reasoning effort per request, and a custom vision encoder that handles variable image sizes and aspect ratios. With support for self-hosting on as few as four GPUs and availability under open weights, it is well suited for scalable, production-grade deployments.

Chat2026年4月29日
入力$1.50/1M tokens出力$7.50/1M tokens
コンテキスト262K

Grok 4.3 is a reasoning-focused model from xAI designed for agentic workflows, instruction following, and high factual accuracy tasks. It supports text and image inputs with text output, with reasoning always active and not configurable by effort level. The model features a 1M-token context window with effectively no output token limit, making it well suited for long-document analysis, deep research, and multi-step agentic workflows. It uses tiered pricing, with higher rates applied to requests exceeding 200K total tokens.

Chat2026年4月29日
入力$1.25/1M tokens出力$2.50/1M tokens
コンテキスト1M

NVIDIA Nemotron 3 Nano Omni is an open 30B-A3B multimodal model designed as a perception and context sub-agent for enterprise agent systems. It supports text, image, video, and audio inputs with text output, enabling unified multimodal reasoning within a single inference loop. Built on a hybrid MoE Transformer–Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS), it delivers significantly improved efficiency for video reasoning—achieving ~2× higher throughput and 2.5× lower compute compared to separate pipelines. With up to 300K context length and extended thinking support, it is well suited for scalable, multimodal agent workflows.

Chat2026年4月27日
入力FreeIncluded出力FreeIncluded
コンテキスト256K

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse Mixture-of-Experts (MoE) architecture with approximately 1 trillion parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window. The model includes an integrated thinking mode that preserves reasoning across multi-turn interactions, along with support for structured outputs and function calling. Available exclusively via Alibaba Cloud Model Studio and Qwen Studio APIs, it is designed for high-performance, production-grade agent workflows.

Chat2026年4月26日
入力$1.30/1M tokens出力$7.80/1M tokens
コンテキスト262K

Qwen3.6 Flash is a fast and efficient model from Alibaba's Qwen 3.6 series, supporting text, image, and video inputs with a 1M-token context window for high-context multimodal workflows. Optimized for performance and cost efficiency, it features tiered pricing beyond 256K tokens and supports prompt caching with both cache creation and read pricing, making it well suited for large-scale, high-throughput applications.

Chat2026年4月26日
入力$0.25/1M tokens出力$1.50/1M tokens
コンテキスト1M

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba, supporting text, image, and video inputs with text output. It features a 1M-token context window, enabling large-scale reasoning and multimodal workflows within a single interaction. This updated version of Qwen3.5 Plus introduces tiered pricing beyond 256K tokens, making it suitable for high-context applications while maintaining flexibility for cost optimization in long-input scenarios.

Chat2026年4月26日
入力$0.4/1M tokens出力$2.40/1M tokens
コンテキスト1M

GPT-5.5 is OpenAI's frontier model for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on challenging tasks. It supports text and image inputs and features a 1M+ token context window (≈922K input, 128K output) for large-scale, high-context workflows. Designed for advanced applications, GPT-5.5 excels in reasoning, coding, and multimodal workflows, enabling efficient execution of complex, multi-step tasks within a single system.

Chat2026年4月23日
入力$4/1M tokens出力$24/1M tokens
コンテキスト1.1M

GPT-5.5 Pro is OpenAI's high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It supports text and image inputs and features a 1M+ token context window (≈922K input, 128K output) for handling large-scale, long-context tasks. Designed for long-horizon problem solving, agentic coding, and precise multi-step execution, GPT-5.5 Pro delivers strong reliability and performance across advanced engineering, research, and complex workflow scenarios.

Chat2026年4月23日
入力$30/1M tokens出力$180/1M tokens
コンテキスト1.1M

DeepSeek V4 Pro is a large-scale Mixture-of-Experts (MoE) model with 1.6T total parameters and 49B activated per token, supporting a 1M-token context window for advanced reasoning and long-horizon workflows. It delivers strong performance across knowledge, mathematics, and software engineering tasks, making it suitable for complex, real-world applications. Built on a hybrid attention architecture for efficient long-context processing, the model supports configurable reasoning modes to balance speed and depth. It is well suited for full codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are essential.https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro

Chat2026年4月23日
入力$0.435/1M tokens出力$0.87/1M tokens
コンテキスト1.0M

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts (MoE) model with 284B total parameters and 13B activated per token, designed for fast inference and high-throughput workloads. It supports a 1M-token context window, enabling large-scale reasoning and long-context processing. Built with hybrid attention for efficiency, the model maintains strong performance in reasoning and coding while offering configurable reasoning modes. It is well suited for coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are critical.

Chat2026年4月23日
入力$0.14/1M tokens出力$0.28/1M tokens
コンテキスト1.0M

Qwen3.6-35B-A3B is an open-weight Mixture-of-Experts (MoE) multimodal model designed for agentic coding and long-horizon workflows. It features ~35–36B total parameters with ~3B activated per token, enabling strong performance with high inference efficiency. The model supports text and image inputs with a ~260K token context window, and is optimized for repository-level reasoning, multi-step development, and tool-driven workflows. With strong benchmark performance and improved coherence across extended tasks, Qwen3.6-35B-A3B is well suited for developer tools, coding agents, and real-world engineering applications that require both reasoning depth and efficiency.

Chat2026年4月22日
入力$5.40/1M tokens出力$32.40/1M tokens
コンテキスト262K

Qwen3.6-27B is an open-weight 27B-parameter dense multimodal model from the Qwen3.6 series, designed to deliver flagship-level coding and agentic performance at a practical deployment scale. It supports both text and image inputs and introduces improvements in agentic coding, repository-level reasoning, and iterative development workflows. Despite its relatively compact size, it achieves state-of-the-art results on coding benchmarks, outperforming much larger models in tasks such as SWE-bench and terminal-based workflows. It also provides strong reasoning and multimodal capabilities, along with features like thinking preservation to maintain context across interactions, making it well suited for developer tools, coding agents, and real-world engineering tasks.

Chat2026年4月22日
入力$0.195/1M tokens出力$1.56/1M tokens
コンテキスト262K

MiMo-V2.5-Pro is Xiaomi's flagship model, delivering top-tier performance in agentic capabilities, complex software engineering, and long-horizon tasks. It ranks highly on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro, demonstrating strong real-world reliability. The model can autonomously complete professional tasks that would take human experts days or weeks, executing thousands of tool calls within a single workflow. With a 1M-token context window, it is well suited for integration into advanced agent frameworks and large-scale task orchestration systems.

Chat2026年4月21日
入力$1/1M tokens出力$3/1M tokens
コンテキスト1.0M
表示中 50 / 508