qwen-plus-2025-07-28:thinkingQwen Plus 0728 is a hybrid reasoning model built on the Qwen3 foundation, featuring a 1M-token context window and a balanced trade-off between performance, speed, and cost.
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="qwen-plus-2025-07-28:thinking", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="qwen-plus-2025-07-28:thinking",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )modelmessagesmax_tokenstemperaturetop_pstreamtoolsreasoning_effortstream_optionsthinkingextra_bodyUse these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
See how this model compares to others from the same provider.
Qwen3.7-Plus is a cost-effective multimodal model in Alibaba's Qwen3.7 series, supporting text and image inputs with text output. It combines the series' strong language capabilities with significantly enhanced vision-language understanding, while retaining full-stack agent-level intelligence for coding, tool use, and productivity workflows. Its standout capability is multimodal interactive agency—the ability to perceive real-world scenes, understand screens and graphical interfaces, generate code from visual references, and perform end-to-end navigation within applications. This makes Qwen3.7-Plus well suited for GUI automation, visual coding, productivity agents, and multimodal task execution.
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series, designed for agent-centric workloads with strong performance in coding, productivity, and long-horizon autonomous execution. It supports text input and output and delivers notable improvements in coding and agentic capabilities over previous Qwen generations. Optimized for real-world workflows, the model also supports explicit prompt caching for efficient reuse of repeated context, making it well suited for scalable development, office automation, and advanced agent systems.
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse Mixture-of-Experts (MoE) architecture with approximately 1 trillion parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window. The model includes an integrated thinking mode that preserves reasoning across multi-turn interactions, along with support for structured outputs and function calling. Available exclusively via Alibaba Cloud Model Studio and Qwen Studio APIs, it is designed for high-performance, production-grade agent workflows.
See how this model compares to others from the same provider.
Qwen3-30B-A3B-Thinking-2507 is a 30B-parameter Mixture-of-Experts (MoE) reasoning model optimized for complex tasks that require extended, multi-step reasoning. It is purpose-built for thinking mode, where internal reasoning traces are explicitly separated from final outputs, enabling more structured and reliable problem solving. Compared to earlier Qwen3-30B variants, this release delivers notable gains across logical reasoning, mathematics, science, coding, and multilingual benchmarks, while also strengthening instruction following, tool usage, and alignment with human preferences. With improved reasoning efficiency and larger output budgets, Qwen3-30B-A3B-Thinking-2507 is well suited for advanced research, competitive problem solving, and agentic applications that demand robust long-context and structured reasoning capabilities.
Qwen3.6 Flash is a fast and efficient model from Alibaba's Qwen 3.6 series, supporting text, image, and video inputs with a 1M-token context window for high-context multimodal workflows. Optimized for performance and cost efficiency, it features tiered pricing beyond 256K tokens and supports prompt caching with both cache creation and read pricing, making it well suited for large-scale, high-throughput applications.
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it achieves substantial improvements in factual accuracy, complex reasoning, instruction following, alignment with human preferences, and agentic behavior. Optimized for advanced problem solving and long-horizon tasks, Qwen3-Max-Thinking is well suited for research, complex analysis, and agentic applications where reliability and structured reasoning are critical.
Qwen3 Coder Flash is Alibaba's fast, cost-efficient coding agent model — a lighter version of Qwen3 Coder Plus — built for autonomous programming through tool calling and environment interaction, while still retaining strong general-purpose abilities.
Initialized observational baseline with no recorded failures