PerplexityChat
Llama 3.1 Sonar 8B Online
llama-3.1-sonar-small-128k-onlineLlama 3.1 Sonar 8B Online by Perplexity.
- Context
- 127.1K tokens
- Endpoint
Pricing
Input$0.20 / 1M
Output$0.20 / 1M
Cache Write (5m)$0.20 / 1M
Cache Write (1h)$0.20 / 1M
Cache Read$0.20 / 1M
Web Search$0 / 1M
Quick Start
Select an endpoint and copy a working example for this model.
Endpoint
python
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="llama-3.1-sonar-small-128k-online", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="llama-3.1-sonar-small-128k-online",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsCommon7 params
modelmessagesmax_tokenstemperaturetop_pstreamtoolsExtended4 params
reasoning_effortstream_optionsthinkingextra_bodyCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.
Sonar Pro Search
Sonar Pro Search is Perplexity's most advanced agentic search system, available only via OpenRouter. It powers the Pro Search mode on Perplexity, adding autonomous multi-step research workflows instead of single-query answers. Pricing combines token costs with a per-request fee.
- Context
- 200K
- Input
- $3.75/M
- Output
- $18.75/M
Sonar Reasoning Pro
- Context
- 128K
- Input
- $2.00/M
- Output
- $8.00/M
Sonar Pro
- Context
- 200K
- Input
- $3.00/M
- Output
- $15.00/M
Sonar Deep Research
- Context
- 128K
- Input
- $2.00/M
- Output
- $8.00/M