Models
| Provider | Model | Input | Output | Context |
|---|---|---|---|---|
ERNIE-4.5-21B-A3B-Thinking is Baidu's upgraded lightweight MoE model focused on deeper, higher-quality reasoning. It's tuned to perform strongly on logic, math, science, coding, text generation, and advanced academic benchmarks, delivering top-tier capability with efficient compute usage. ChatOct 8, 2025 | Input$0.035/1M tokens | Output$0.14/1M tokens | Context131K | |
This is a multimodal Mixture-of-Experts chat model with 28B total parameters and 3B active per token, designed for strong text and vision understanding. It uses modality-isolated expert routing for efficiency, supports a long 131K-token context window, and delivers high-throughput inference. Advanced post-training (SFT, DPO, UPO) and RLVR alignment further enhance cross-modal reasoning and generation quality. ChatAug 11, 2025 | Input$0.375/1M tokens | Output$1.50/1M tokens | Context30K | |
This is a text-based Mixture-of-Experts (MoE) model with 21B total parameters and 3B active per token, designed for efficient, high-quality generation and understanding. It supports a long 131K-token context window and achieves fast inference through parallel expert routing and quantization. Advanced post-training methods (SFT, DPO, UPO) and specialized routing/balancing techniques optimize performance across diverse tasks. ChatAug 11, 2025 | Input$0.375/1M tokens | Output$1.50/1M tokens | Context120K | |
ERNIE X1 32K is a large language model designed with an extended 32K-token context window, enabling deeper understanding and reasoning over long documents and multi-turn interactions. It balances strong performance across general reasoning, comprehension, and instruction following, making it well suited for tasks that require handling extended contexts such as document analysis, summarization, and detailed multi-step workflows. ChatMay 31, 2025 | Input$0.75/1M tokens | Output$3/1M tokens | Context32K | |
ERNIE X1 32K is a large language model designed with an extended 32K-token context window, enabling deeper understanding and reasoning over long documents and multi-turn interactions. It balances strong performance across general reasoning, comprehension, and instruction following, making it well suited for tasks that require handling extended contexts such as document analysis, summarization, and detailed multi-step workflows. ChatMay 31, 2025 | Input$1.50/1M tokens | Output$6/1M tokens | Context32K | |
ERNIE X1 32K is a large language model designed with an extended 32K-token context window, enabling deeper understanding and reasoning over long documents and multi-turn interactions. It balances strong performance across general reasoning, comprehension, and instruction following, making it well suited for tasks that require handling extended contexts such as document analysis, summarization, and detailed multi-step workflows. ChatMay 31, 2025 | Input$0.75/1M tokens | Output$3/1M tokens | Context32K | |
ERNIE X1 32K is a large language model designed with an extended 32K-token context window, enabling deeper understanding and reasoning over long documents and multi-turn interactions. It balances strong performance across general reasoning, comprehension, and instruction following, making it well suited for tasks that require handling extended contexts such as document analysis, summarization, and detailed multi-step workflows. ChatMay 31, 2025 | Input$1.50/1M tokens | Output$6/1M tokens | Context32K | |
ERNIE 4.5 300B A47B by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$0.3/1M tokens | Output$1/1M tokens | Context123K | |
ERNIE 4.0 8K by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$49.72/1M tokens | Output$49.72/1M tokens | Context8K | |
ERNIE Lite 8K 0308 by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$112.50/1M tokens | Output$225/1M tokens | Context8K | |
ERNIE Lite 8K 0922 by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$112.50/1M tokens | Output$225/1M tokens | Context8K | |
ERNIE Speed 128K by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$1.46/1M tokens | Output$2.92/1M tokens | Context131K | |
ERNIE Speed 8K by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$1.46/1M tokens | Output$2.92/1M tokens | Context8K | |
ERNIE 3.5 8K by Baidu. Use it from Apertis SDKs, provider-compatible SDKs, or direct HTTP requests. Chat | Input$5.57/1M tokens | Output$5.57/1M tokens | Context8K |